Pith. sign in

Paper Citation Record · LEDGER

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

As of 11 August 2026, this Paper Citation Record lists 100 of 154 outbound references and 100 inbound Pith citation observations for arXiv:2504.10479.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.10479 v3

Coverage vector

measured 100 of 154 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T13:41:07.991012Z

measured 200 of 200 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 100 of 745 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T20:17:26.566560Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

100 of 154 outbound references displayed

  • verified exact53
  • verified fuzzy43
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

6
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 56017071-d3c3-44f2-9f57-4525753d64e0 · outbound

This paper cites GPT-4 Technical Report.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models GPT-4 Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-10T13:41:08.208175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:de7f6622585413f9e7e8f06114690fbe0db53d755fe248f1ae70be76b6cfc34e

Observation 2e17a19a-faef-4fb3-89d2-93dccb5fc7e7 · outbound

This paper cites CG-bench: Clue-grounded question answering benchmark for long video understanding.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models CG-bench: Clue-grounded question answering benchmark for long video understanding

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T13:41:08.383503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:1e0435ae184dc26adb38c2fbcdfc40298978d6177b98911919573f3a6c6cac8f

Observation 4cd3784e-e2b9-4c1d-973c-a9cb153bf5dc · outbound

This paper cites The claude 3 model family: Opus, sonnet, haiku.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models The claude 3 model family: Opus, sonnet, haiku

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T13:41:08.385656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:9d0a5b17d64317f9d959fdf0e1548bdcf3c0d1b69ceef084ab3b4c0daf0ccd43

Observation 63863f0e-fa66-4640-bfa0-d32424c8477e · outbound

This paper cites Program Synthesis with Large Language Models.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Program Synthesis with Large Language Models

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-10T13:41:08.210622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:aa575bf481463efa4f42f23e72357ef4e2aaf3febfeba3f95dda0acd34a0349a

Observation dacd29b5-663f-401d-ad1b-1e58422fd67b · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-10T13:41:08.213086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:b665b488fe4ea9b4262c768ae563e30284b8fbc0806ab5ce4631cd8c678e1a3e

Observation 0fa9c2e0-f66a-476b-84f0-a2fde5a65dbe · outbound

This paper cites Qwen2.5-VL Technical Report.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Qwen2.5-VL Technical Report

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-10T13:41:08.215773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:d9fe66f97f17237aa67deb2e877434fed57f1e7e0d0c723f4e9aa0b25b1718df

Observation 56e9debf-ff99-48e2-8b42-93b9f328610b · outbound

This paper cites Smollm-corpus.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Smollm-corpus

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T13:41:08.393133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:9eff4d5329d7d68e8525ea2f436dbd83e2cd15836ea4d2014da07e5ede9cfab6

Observation 16bddcd9-47ee-4ec1-8499-4d73fc809b7e · outbound

This paper cites Scene text visual question answering.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Scene text visual question answering

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T13:41:08.398362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:b61f740d3d6853e70809cf143393c95e313af95b285fd6bda9df2b58a0a40666

Observation 3e75801c-f429-4383-8e31-44b2df9bf678 · outbound

This paper cites An augmented benchmark dataset for geometric question answering through dual parallel text encoding.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models An augmented benchmark dataset for geometric question answering through dual parallel text encoding

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T13:41:08.400351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:90831ed99d3fa94dc8af0bf107222a204dbc54446b836db66b66c4add4fb89b1

Observation 13f3b8a6-0ebf-4a73-8942-754f92418f5f · outbound

This paper cites MapQA: A Dataset for Question Answering on Choropleth Maps.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models MapQA: A Dataset for Question Answering on Choropleth Maps

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:41:08.218987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:a126c2eda8a905bd77606e0f2dfc6bb7a84f389d6afc1f884b0522bf717b9250

Observation 429dc09c-1502-4571-9bb8-5aeaac25600d · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-12T15:52:36.167809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:9add7eed9dd82b204c16102bf19734b9216f4d19d736addbbf675a54bb44817f

Observation af2f885b-c983-414d-aaa1-0758eb7f8dc2 · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-12T19:41:44.612219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:8c0711aceac12ce655a62ef92948040b4348e374cab5251440b9e01bb60d6b56

Observation 8ad0c483-aaf4-418f-ab69-77dcef1b97ab · outbound

This paper cites Evaluating Large Language Models Trained on Code.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Evaluating Large Language Models Trained on Code

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-10T13:41:08.227090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:32c42cbe84e512912df9e2fb9ff0d1088328f2cf211d81e659b4eddc831624d4

Observation c777bb48-c43a-44be-9197-3c0d73e0aa55 · outbound

This paper cites InternEvo: Efficient Long-sequence Large Language Model Training via Hybrid Parallelism and Redundant Sharding.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models InternEvo: Efficient Long-sequence Large Language Model Training via Hybrid Parallelism and Redundant Sharding

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T13:41:08.230480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:24791fec4261262132d11380e26d0c987fa438b55a563f498abeadd491cc2c37

Observation 3b253375-e068-4098-b7f3-71a665c88a01 · outbound

This paper cites M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:41:08.233339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:d40558a7dfe5e709752a0db18a28fa2af10686c5ed90b37ed1e35e16f927d5dd

Observation 56c4f3d4-7eeb-46d2-9c5e-0224e90f59ad · outbound

This paper cites Theoremqa: A theorem-driven question answering dataset.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Theoremqa: A theorem-driven question answering dataset

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T13:41:08.412722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:498c8b2aeb545c233e886e2ce384c3245c5a1861bba1423034a694e0cacc797d

Observation dca35e67-146b-42ea-926a-304bbc25abee · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-10T13:41:08.239742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:31291794b4faa0178550de3a09a764591e7514ae2dc0e5b021f4e504f5670d09

Observation e18683e8-78b7-4cc3-bf15-7280ace80e77 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-12T20:58:59.653229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:63683a5b042e2ce07b3b90d04c78c0dfbb74476fff6e2cd73d8cab23caf3ad7e

Observation 64555317-3d21-4729-84f6-87b8917bc516 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T13:41:08.418301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:2e134ee2b0b3da6b8acb7af7e723c70b9ed82a93b88d6004500896110e41dd34

Observation 961d9827-e406-43d2-8d20-f8de8d333c19 · outbound

This paper cites SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-17T10:09:46.778093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:5e053c58043dd6ca70c54d4b523271e4df701e3455ed189a00dffe2d59cb7d80

Observation 227cab50-7277-48a6-be3a-21050c0ba8bc · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-11T02:44:53.738794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:c91a1e05b774509d7bd7b13dd30b370bc552de9786c93d9e49c5a5c86627b191

Observation d93645c5-e167-4955-b7ae-aeebb07c2b9d · outbound

This paper cites Simple and effective multi-paragraph reading comprehension.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Simple and effective multi-paragraph reading comprehension

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T13:41:08.423799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:f2754212dbe8a51f2b32c636d7553cef9776587529b34ce4a1d1ab7aa7da45c8

Observation 0269d5fd-ab27-485f-92ec-8843919b3837 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Training Verifiers to Solve Math Word Problems

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-10T13:41:08.255963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:7d70060eb5b8c0a5bafd86dae791c27c420588a96992748588e76ee56ca92a41

Observation 63360a63-bc28-49cd-a554-19ad2935a695 · outbound

This paper cites Opencompass: A universal evaluation platform for foundation models.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Opencompass: A universal evaluation platform for foundation models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T13:41:08.427459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:33e1781396d6c3e28ecfe7d8a8242114753c498c45cf684f23872622e8e11d07

Observation aed4d604-7f67-4fb5-acd5-19d54a362ef9 · outbound

This paper cites Grok-1.5 vision preview: Connecting the digital and physical worlds with our first multimodal model.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Grok-1.5 vision preview: Connecting the digital and physical worlds with our first multimodal model

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T13:41:08.429410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:600f53b53a2482c3bc36187b6a4c610cfe3aaf1b8a7d880c1702d8431665c5eb

Observation f6e4e7ae-6dc9-4717-8785-6343b01fb55a · outbound

This paper cites NVLM: Open Frontier-Class Multimodal LLMs.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models NVLM: Open Frontier-Class Multimodal LLMs

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:41:08.258635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:2d372d6bec657bc8e6350935e08ce382e011f8ba3ccab7077e512fa7da53bfa5

Observation e404920f-6e33-477a-ad91-4c1664591cb8 · outbound

This paper cites Gemini 2.0 is now available to everyone.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Gemini 2.0 is now available to everyone

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T13:41:08.436338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:849bf4b47fd573341922cf7038d5f18e2f9edbda2755f49b02c33b5828a43801

Observation 2ee03a1c-19f5-41c9-9997-a3b38fd9f75b · outbound

This paper cites Introducing gemini 2.0: our new ai model for the agentic era.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Introducing gemini 2.0: our new ai model for the agentic era

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T13:41:08.438018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:42a3e4bd32934557bea58da1a05e049c3becae7b61f2d9dcd4adef8fb2372f72

Observation e00a2882-d65a-4442-a090-1dfebde268d7 · outbound

This paper cites Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-15T01:55:12.888993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:e6e48062586dc2491ea65f9226a219a26d6d849d299670817d199d18ae0e5457

Observation 35a7f493-0b75-4099-bfd8-01d8a6304fee · outbound

This paper cites InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:41:08.267570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:42effbd8beb4b5b2b1a6563b160e8557ff97e19b9f6dbbfacff3a0eeab97d143

Observation 47cad76e-e666-46ea-ba99-6a4137ee179a · outbound

This paper cites Vlmevalkit: An open-source toolkit for evaluating large multi-modality models.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Vlmevalkit: An open-source toolkit for evaluating large multi-modality models

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T13:41:08.443073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:f54bceddf3bb09a068a1109b236b407d7381524740ae9377462a7902a6678ae3

Observation 7b947b91-e3f4-487e-9520-5283c96336b4 · outbound

This paper cites The Llama 3 Herd of Models.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models The Llama 3 Herd of Models

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-10T13:41:08.270190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:e2a29d5bb56c9ac5fdb996f80ab66e2c8aa86fc334f61bab8d3ba95e47e5d65e

Observation a6b6f22f-c849-40be-aee7-4c23dfdc372a · outbound

This paper cites MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:41:08.272679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:a29d917a1584ad42bb707b5899398844513bbd7729ecdb4bc1127d17cbca3801

Observation 1f651706-54cc-4dd6-907f-48668b377f05 · outbound

This paper cites Learning generative visual models from few training examples: An incremental bayesian approach tested on 101 object categories.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Learning generative visual models from few training examples: An incremental bayesian approach tested on 101 object categories

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T13:41:08.448276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:75dcab23e2a64125b86a16aa1c8e9b3ff62e0add72143327c24488352718c345

Observation e1a23374-18ff-42bc-945b-9af51eaeeb19 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-10T20:25:34.870358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:266407242b7bab722332b56969da3afc13ded4cf2604cccd9718c3630326112b

Observation 39e2b465-27e0-4d0a-8d4a-4112ab8b9bf3 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:58:42.298802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:eb03778127a6fa2dcd9d476b1a34151f2d836d4d3cf1cdecdac2b45a1d3154fe

Observation 2a242639-a304-471e-8c5f-5b5ad78f44e1 · outbound

This paper cites BLINK: Multimodal Large Language Models Can See but Not Perceive.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:18:15.850579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:2e4833590dc9497ed41df16d61ffec1e167cc42d431119ae8d21f7e2ad71951a

Observation 99c40049-2da8-48ca-908d-407fabca8222 · outbound

This paper cites G-LLaVA: Solving Geometric Problem with Multi-Modal Large Language Model.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models G-LLaVA: Solving Geometric Problem with Multi-Modal Large Language Model

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:41:08.303127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:abc5d9c373b4086176cdcfa7a33c34e004d982016ac2fd62b2055d2f0344909b

Observation 74a09442-949e-43fa-b0e9-8dcd2b66185e · outbound

This paper cites Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:41:08.305987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:922020cd8a5aaa13425ff25dae7f8af3919f9281210c44e33bdce0d3ae78b9a8

Observation 7567a43a-81e0-4548-bf55-f8e0092efe2b · outbound

This paper cites V2PE: Improving Multimodal Long-Context Capability of Vision-Language Models with Variable Visual Position Encoding.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models V2PE: Improving Multimodal Long-Context Capability of Vision-Language Models with Variable Visual Position Encoding

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:41:08.309183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:87053d6a4bea0e6aee326f588faf7ca50e22f3e2d2070d7e0ba1889eb5f720d0

Observation 6e854932-c0a3-48ab-bd1f-687d4ebbd5ea · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answering.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Making the v in vqa matter: Elevating the role of image understanding in visual question answering

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T13:41:08.461898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:8cd6b04a3c0dc0f174fe23b726cef5ee50b7adbb1e4edcf29e1c404a431ca695

Observation 62830ee8-01c9-4fcd-a88b-38ec23fbcc33 · outbound

This paper cites Infinity-MM: Scaling Multimodal Performance with Large-Scale and High-Quality Instruction Data.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Infinity-MM: Scaling Multimodal Performance with Large-Scale and High-Quality Instruction Data

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:41:08.311987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:e464506fe299d07d2664dc520efe0fdefa11be1515eaa721401662d4d70addde

Observation 6f218f3e-0620-4daf-b312-46e0ebea4e4b · outbound

This paper cites HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-17T01:22:04.360724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:c008b567a77c071fc968f44e502650405feec0e247ffd381bd9296d82a81a3cb

Observation 279fd556-6349-42b0-959c-6fb4965345ff · outbound

This paper cites Measuring massive multitask language understanding.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Measuring massive multitask language understanding

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T13:41:08.467485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:63f3e1034cb797de5b0743036ad276038c58c6dafd6ce7ed25fdac4971a55406

Observation 34e69bd0-2ebd-4d68-9c92-ce33ac0843c3 · outbound

This paper cites Measuring mathematical problem solving with the MATH dataset.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Measuring mathematical problem solving with the MATH dataset

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T13:41:08.469249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:91e595eda9c716d83108a0542877f016c0cd4389bd0979271836e8f39bb2e333

Observation 6c1e602a-50df-4a1e-954e-f2cfd5bdf537 · outbound

This paper cites C-eval: A multi-level multi-discipline chinese evaluation suite for foundation models.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models C-eval: A multi-level multi-discipline chinese evaluation suite for foundation models

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T13:41:08.470959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:9b417684326590f8c15b933e55c018e8b6cef807853bc5be51b756807345655a

Observation ad463b37-61ec-4d5f-82f7-14b4a3592bb1 · outbound

This paper cites Icdar2019 competition on scanned receipt ocr and information extraction.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Icdar2019 competition on scanned receipt ocr and information extraction

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T13:41:08.472622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:40282f9587ec9740e543fe0ba0e48970a44886833e5968eca9322789b4c4f804

Observation 97185d93-673d-4468-ad97-ab643a72e62d · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T13:41:08.474271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:d5b02aa525c25e1ab18a5a1c10ab4bd23d98bcf5761dada6d37616bcbf4e789d

Observation c5e544a3-4cd7-4dc9-a46a-cdf8e342dd3a · outbound

This paper cites MANTIS: Interleaved Multi-Image Instruction Tuning.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models MANTIS: Interleaved Multi-Image Instruction Tuning

Reference 51

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T13:41:08.322697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:ae0be1e7552fa8c8716f3ab3172ad291e36659b7230e44ce1c4795a7e8908f8f

Observation ba68e2b8-5aaa-4396-8d27-03fa089ed5d5 · outbound

This paper cites TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:00:28.883814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:8cd4cba4493b71e26e529a3569ea48a26ab2d61a54198d90bdf1e62b3f9e7d27

Observation 197dbeae-0714-4cb4-b632-28d1edc8729f · outbound

This paper cites Binary Classifier Optimization for Large Language Model Alignment.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Binary Classifier Optimization for Large Language Model Alignment

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:41:08.328045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:5b94d6a13479d0f85873c92dc872241abd3ac61393dd667eb3b6285f2190492c

Observation ca1c5b0f-7d92-4288-b8ad-0a6c1657a248 · outbound

This paper cites Dvqa: Understanding data visualizations via question answering.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Dvqa: Understanding data visualizations via question answering

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T13:41:08.480948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:44591978e494068b1ffda796e2f2f4121b7426dc6160acfb2a15b1076060d1d9

Observation 9c4e5890-e323-4baa-91b5-86b83b4d3d67 · outbound

This paper cites GeomVerse: A Systematic Evaluation of Large Models for Geometric Reasoning.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models GeomVerse: A Systematic Evaluation of Large Models for Geometric Reasoning

Reference 55

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T13:41:08.330629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:835ec7cbca77faa1eb4e06e11e863f2803f074c295499d74470c212acb5df057

Observation af44c27e-5484-4633-b6b8-e00174edd9bc · outbound

This paper cites Referitgame: Referring to objects in photographs of natural scenes.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Referitgame: Referring to objects in photographs of natural scenes

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T13:41:08.488031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:d27e897a629ea2cc7beda39697e61c15aa88cc7c80acd9357f3271ff327fea54

Observation a0f678d5-20ab-4f89-95ec-4107f9bdf393 · outbound

This paper cites A diagram is worth a dozen images.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models A diagram is worth a dozen images

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T13:41:08.489624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:190866861b36311fc7f5a9b5c511848ddbc7779266e97b2292f76179763a72ad

Observation cbe6585c-29e2-4f76-98d2-f600653cae00 · outbound

This paper cites Natural questions: a benchmark for question answering research.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Natural questions: a benchmark for question answering research

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T13:41:08.498152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:fa1f163424da4c41489dd465079e2e0f76e3b55eac2e826c214da814ff7e724e

Observation efac3ba4-3f1d-4ee6-8e09-c65a3426cbcf · outbound

This paper cites RACE: Large-scale ReAding Comprehension Dataset From Examinations.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models RACE: Large-scale ReAding Comprehension Dataset From Examinations

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:41:08.333091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:cbc1c5d6e362c7af6da2367b5840779e740a1dd50cb798379f68c4605125fe89

Observation 68f46997-cb03-46d6-9992-693bf6a94c74 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models LLaVA-OneVision: Easy Visual Task Transfer

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:41:08.335417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:b07a05871e311d7303b9c25c5911b3f7e8b5b0fe50442091810b33661160b328

Observation 6ffa1d7f-aaee-47f1-9efa-8d2f643abf52 · outbound

This paper cites SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:41:08.337965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:8f3752f5424b05c8daa1ecfda2e73d5e582d1d09300ea54e693e3bbb44421df1

Observation cde53698-b607-4154-bbb3-3f3c3d1803a7 · outbound

This paper cites R-Bench: Are your Large Multimodal Model Robust to Real-world Corruptions?.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models R-Bench: Are your Large Multimodal Model Robust to Real-world Corruptions?

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:41:08.340531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:9f077904011c88da0409c0068da62cdb9301ff5b2fc8b11be8cf59ddcb307a2f

Observation bc7aac1e-221a-4981-9a6c-2daedea49371 · outbound

This paper cites CMMLU: Measuring massive multitask language understanding in Chinese.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models CMMLU: Measuring massive multitask language understanding in Chinese

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-16T21:01:11.049549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:001d1b84a6cd74d6f265e7320abc57e7cd2db6aa8bd8982a880261f359391af9

Observation 96dde0f3-38d0-406b-b44d-5595cf56c437 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models VideoChat: Chat-Centric Video Understanding

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:30:00.879549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:2e81537aab0663f0059feeb513369b8ec6badaa46e98620a420c6e4d2db2c894

Observation ea763f14-65cd-43f4-8f5d-e3f4e8522369 · outbound

This paper cites Mvbench: A comprehensive multi-modal video understanding benchmark.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Mvbench: A comprehensive multi-modal video understanding benchmark

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T13:41:08.407622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:8629ebfdd23c81b734570d3a52929681d7f6fbc0648c9300959b8f8a3e9d5b2d

Observation e9259743-9772-4012-a169-a6436c8abd83 · outbound

This paper cites Mvitv2: Improved multiscale vision transformers for classification and detection.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Mvitv2: Improved multiscale vision transformers for classification and detection

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T13:41:08.409297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:c809fa4e7d0d70630ddeacc34238fd2f44d9e53c7a16183116d5b8e6d8119d10

Observation a4326a14-ba66-4440-b8f4-2269e7f0dee1 · outbound

This paper cites Evaluating object hallucination in large vision-language models.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Evaluating object hallucination in large vision-language models

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T13:41:08.410923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:886ccf20c9268f0254d5fc25b3ff1961d92406e74cf7c015566e030685803b3e

Observation 3984b339-064a-43db-9912-503af55cc6e9 · outbound

This paper cites Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:41:08.362333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:3f146bc8ab2579ba6e62b5cfaf0a05f715caa8084229acbfdd93fc280e563431

Observation 5c37d1f1-2132-4333-af88-222f777f6598 · outbound

This paper cites Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:41:08.369391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:1f6b731bb91588250ba76a68ef7a3ea61aa6a8ba306fcdcce152cae0cdf2a0f5

Observation 7e3c5189-a7ae-4b70-aecc-3c8789e7dcf6 · outbound

This paper cites Let’s verify step by step.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Let’s verify step by step

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T13:41:08.420005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:285a0923dd1c77e1c2ee305dc10e706ece6d14f34d7991aae6369ca0af277e71

Observation 112be398-657f-4b46-a2cc-051d2e6efc37 · outbound

This paper cites Vila: On pre-training for visual language models.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Vila: On pre-training for visual language models

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T13:41:08.421886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:847cf6d0f7318a775fec78f8c253914bba99cdcbf269baf4e63e87dcd6dc4c0e

Observation 027b7536-7492-4f42-9d0e-d2ceb2612068 · outbound

This paper cites CLEVR-Math: A Dataset for Compositional Language, Visual and Mathematical Reasoning.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models CLEVR-Math: A Dataset for Compositional Language, Visual and Mathematical Reasoning

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:41:08.372728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:6f3e28a5216eda68a6641fff00ef494a23b5602bf95d1d02b0da6711c1d8157a

Observation 0e44c233-ef2e-47ed-8809-536391a18559 · outbound

This paper cites Visual instruction tuning.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Visual instruction tuning

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T13:41:08.430979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:78c3e2ded77dd4de81c08337c88a12df0185c08ecfa9f04332fe3d50a92f1828

Observation cacf4376-3d67-4178-b69f-3dd2c259453a · outbound

This paper cites Grounding dino: Marrying dino with grounded pre-training for open-set object detection.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Grounding dino: Marrying dino with grounded pre-training for open-set object detection

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T13:41:08.439770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:cd911b3711d359cd6caccd86158419272912a9992d92959dc1d86b40ad60e079

Observation 4c8570eb-9042-47be-b7d3-e4241449b233 · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models MMBench: Is Your Multi-modal Model an All-around Player?

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-05-12T17:20:54.147488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:993710de33589bcabd57c182083588cbc91b833c04f645b0b7faab3439600584

Observation cbc8a068-f9e8-4fda-8e1c-d49fe18adf14 · outbound

This paper cites OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-05-17T09:55:36.035986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:8d89008606bed8fd02c23aad1873134e321a1628afbf1b3d81a607b278cbbfd2

Observation 16e18c3b-2b72-4467-b02b-d442d328f963 · outbound

This paper cites Acemath: Advancing frontier math reasoning with post-training and reward modeling.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Acemath: Advancing frontier math reasoning with post-training and reward modeling

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T13:41:08.446557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:a73915de0b00867967d17a0e8d5ead4ae9320b91c35e0ca7e1ac99c60b39fd1d

Observation 78546f28-51b5-41e9-b238-f48203b0cc7f · outbound

This paper cites Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 78

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:41:08.381147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:a019c0bdfd86ad7c7f41dc0322bde5f9430022e9b661e17c7e747e590243cc53

Observation c6f1bbd5-7f6b-4ba7-ac2d-974bb218370d · outbound

This paper cites Scp-116k: A high-quality problem-solution dataset and a generalized pipeline for automated extraction in the higher education science domain.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Scp-116k: A high-quality problem-solution dataset and a generalized pipeline for automated extraction in the higher education science domain

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T13:41:08.451773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:4d7a049b9f9ff65d1465b7998f20e6cf412dd21f8d41f50f33b0f53bd15dedbc

Observation 4aa5ed8b-6816-413f-b710-1e86b5f59be5 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 80

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:30:15.750849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:9f00f45a34c10aa8eaeedbafff1788b8f7b5bbca194e8f6ad537aa939cecf4f7

Observation edf6f1d4-dce2-49f9-a15c-43dbd9a766a7 · outbound

This paper cites Inter-GPS: Interpretable Geometry Problem Solving with Formal Language and Symbolic Reasoning.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Inter-GPS: Interpretable Geometry Problem Solving with Formal Language and Symbolic Reasoning

Reference 81

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:41:08.040558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:4df6ef77ec6a1e891b75bfa7fbfddbd2e7d28078edd9b39b7d52bce52e824c1c

Observation 9a5922a2-f8e7-4a2f-973d-420edea72035 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T13:41:08.458109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:3a8c6652d733ee46c54287853dc8e571d622bb7b96fcf4a0e12b5d5dd623374d

Observation 16e0f5ce-ae4c-47e4-a1fb-bc94f4fbe4aa · outbound

This paper cites IconQA: A New Benchmark for Abstract Diagram Understanding and Visual Language Reasoning.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models IconQA: A New Benchmark for Abstract Diagram Understanding and Visual Language Reasoning

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:41:08.043710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:329c3382806f1ff07eddcd26325ff3e8dd729ed4bc226255c276965cde1c629e

Observation f4566433-df2b-4547-be0c-63e97c43124e · outbound

This paper cites Ovis: Structural Embedding Alignment for Multimodal Large Language Model.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 84

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:41:08.046852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:3b41618c1c2f2c798e51b236c4dbd731ae3186ec74c6b16ed96ddb035ffd7a15

Observation 50d1bb8a-1bda-4b21-8f31-ac3177e5c7f9 · outbound

This paper cites BlueLM-V-3B: Algorithm and System Co-Design for Multimodal Large Language Models on Mobile Devices.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models BlueLM-V-3B: Algorithm and System Co-Design for Multimodal Large Language Models on Mobile Devices

Reference 85

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:41:08.049836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:77a3d5adea38985d988dfc4d329aada2d34f5d53ca36d55f371d90043ddecfa7

Observation 2b254b4a-fbad-4a33-8c1f-dc89cd8c24f3 · outbound

This paper cites WildVision: Evaluating Vision-Language Models in the Wild with Human Preferences.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models WildVision: Evaluating Vision-Language Models in the Wild with Human Preferences

Reference 86

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:41:08.056502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:c057a88d20860a4d2e9092ae676943b602c6f6e744c17fb2de230746082012c0

Observation 3296f90c-8885-472d-a340-987a6d7f00a5 · outbound

This paper cites Improve Mathematical Reasoning in Language Models by Automated Process Supervision.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Improve Mathematical Reasoning in Language Models by Automated Process Supervision

Reference 87

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:53:45.995574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:34a8db6764848e8404e827024fc743cd986dfe6196ac2ebde6e182bac2b2ccdc

Observation 6f9fd756-dd20-4ea3-bf3f-de1e663be830 · outbound

This paper cites Generation and comprehension of unambiguous object descriptions.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Generation and comprehension of unambiguous object descriptions

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T13:41:08.479426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:c0659b9ca86ce021eb6fb9b67fefc40840950b27c6dc88fb9fba2b8525f04e2c

Observation 7c025a5f-c007-4400-a17c-1a43417b5062 · outbound

This paper cites SmolVLM: Redefining small and efficient multimodal models.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models SmolVLM: Redefining small and efficient multimodal models

Reference 89

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:23:51.804797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:f59927dbd0b93824c04e0cd2d7060bbe753d2c8b0cd285313e74454c4c179d0e

Observation e2b99258-9630-4cc9-98c8-43525eddabcb · outbound

This paper cites Ok-vqa: A visual question answering benchmark requiring external knowledge.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Ok-vqa: A visual question answering benchmark requiring external knowledge

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T13:41:08.387631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:56e65a445fa5146ae2cde5615298d671537a025735e654da4a37c6f27aa44ee5

Observation 7a9a8293-78e7-4bf0-ab5e-7b079d1fadb9 · outbound

This paper cites Chartqa: A benchmark for question answering about charts with visual and logical reasoning.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Chartqa: A benchmark for question answering about charts with visual and logical reasoning

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T13:41:08.389598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:c44b35c12d32e4de27f4aeb2eb91faf089703a5f3fb9b5bedf23b1d85a4d18bf

Observation 1a202c3a-cd35-4fe1-aaa8-f723909d9df5 · outbound

This paper cites Infographicvqa.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Infographicvqa

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T13:41:08.391383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:4e1c565139d12eb611b46cb93eecadc26b103fc0c90ce4470f1d90871fd4e88e

Observation fc41b8a2-3aa1-4bda-8acc-14efe6d738f8 · outbound

This paper cites Docvqa: A dataset for vqa on document images.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Docvqa: A dataset for vqa on document images

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T13:41:08.402281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:4ef0ec486cc5ddec3444844b888bf962fad796869b673fb9147f3acf19783e64

Observation ee956037-fb1a-4883-9db4-5f2edfa415a8 · outbound

This paper cites LLM Critics Help Catch LLM Bugs.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models LLM Critics Help Catch LLM Bugs

Reference 94

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:41:08.064965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:c2a2059684e4e9c5d84281592c4c44deff0503e44084a85e23b463c53535c574

Observation 9ff29946-03b1-4140-bb99-fcbb4ebea566 · outbound

This paper cites MMIU: Multimodal Multi-image Understanding for Evaluating Large Vision-Language Models.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models MMIU: Multimodal Multi-image Understanding for Evaluating Large Vision-Language Models

Reference 95

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:41:08.067886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:46dcc0c051b1c5681c22d36b07d2b9e19119c444ccb11c218aaed617fa2c78c3

Observation 53ffb316-dd3d-4bcd-9465-b5f05e34d038 · outbound

This paper cites Ocr-vqa: Visual question answering by reading text in images.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Ocr-vqa: Visual question answering by reading text in images

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T13:41:08.414450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:82a00dae312d788b157a75efb5d717ed8540f6c3e8c68a98953ef338f785a194

Observation aa8fb495-d1be-44c2-9a5b-44455e210b93 · outbound

This paper cites Gpt-4v(ision) system card.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Gpt-4v(ision) system card

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T13:41:08.416484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:7d52398a86cb9c0522bb141fc434de81837481048cd373c8fc7b8b4c38e77e02

Observation 2965058a-3ca8-4656-ad6a-724504d790ef · outbound

This paper cites an unresolved cited work.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Unresolved cited work

Reference 98

Resolution
unresolved
raw_fallback, observed 2026-05-10T13:41:08.425663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:0fe12cd16bfeb48b0618039a8c86e16af3be872cc49f806cd98ee45a80c867e5

Observation 23cea178-7685-4b9a-b84e-6bc3ed66b0a6 · outbound

This paper cites Gpt-4o system card.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Gpt-4o system card

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T13:41:08.441272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:a9c5009ba6b35c52bdac6294d6770f77edca7234af94992193224940d3ada8d0

Observation 91ef0e03-403d-4559-b0fb-82f687bd6a8f · outbound

This paper cites We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?

Reference 100

Resolution
verified exact
arxiv_id, observed 2026-05-16T21:55:41.398838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:00731a4e4388612ff1fb41631830261c8268a0f864f459a4fefe2f2d5d2a764c

Observation eafb055f-075c-42e3-9559-2f6650fb124f · outbound

This paper cites UI-TARS: Pioneering Automated GUI Interaction with Native Agents.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models UI-TARS: Pioneering Automated GUI Interaction with Native Agents

Reference 101

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:41:08.074041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:66be4b196bb0dd72026f2651054640a40d5a065fba6ddad861229ebcd054e0cc

Observation 26ecea0e-3d97-473e-863f-5c16f2e232c2 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.Advances in Neural Information Processing Systems, 36.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Direct preference optimization: Your language model is secretly a reward model.Advances in Neural Information Processing Systems, 36

Reference 102

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T13:41:08.453953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:f2380886def1b3229e977e29760e122282bfe7c8e682b86c0e298fda1d020dd9

Pith citing papers

Observation 05f338c6-f496-4ed5-a922-455affbcbc31 · inbound

Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos cites this paper.

Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 124

Resolution
verified exact
local_arxiv, observed 2026-05-16T11:39:22.581047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T11:39:22.340737Z digest=sha256:524d6197dfe612e587079e98cca28ef4432a422af3f2067e8f0c026edcbc5f21

Observation 03d07df6-49f5-4699-b8b3-d34a45cb49ec · inbound

Compositional Generative Model of Unbounded 4D Cities cites this paper.

Compositional Generative Model of Unbounded 4D Cities InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 123

Resolution
unresolved
no resolver link, observed 2026-08-10T20:17:26.566560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:17:26.566560Z digest=sha256:4aea0ad50144f087cbd6536b24ca0bcda3c327932f560e14c574808be91f01b6

Observation b17fde1e-7c7c-4d82-bd6b-44a9f5314801 · inbound

Perception Encoder: The best visual embeddings are not at the output of the network cites this paper.

Perception Encoder: The best visual embeddings are not at the output of the network InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 168

Resolution
verified exact
local_arxiv, observed 2026-05-13T22:21:15.829279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T22:21:15.681336Z digest=sha256:27e63522c4ab4c729e49a487fe6e08b53eee0d9576320f32e99d0bf08d533361

Observation f5318b83-0017-4777-817a-fd88113da0cd · inbound

Benchmarking Large Vision-Language Models on Fine-Grained Image Tasks: A Comprehensive Evaluation cites this paper.

Benchmarking Large Vision-Language Models on Fine-Grained Image Tasks: A Comprehensive Evaluation InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-05-22T18:26:55.270297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-22T18:26:12.597756Z digest=sha256:55caf11ff7f17b922160f198a6e8b580a0d89923052afc26ea8ce27379707c03

Observation 1a098fb1-db94-4d62-9439-b3f194309c96 · inbound

DeepEyes: Incentivizing "Thinking with Images" via Reinforcement Learning cites this paper.

DeepEyes: Incentivizing "Thinking with Images" via Reinforcement Learning InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-11T14:42:56.781289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-11T14:42:56.565621Z digest=sha256:7e9c6a0c93a67a64a9e8d00a1894e22d2d523c7976c9a8353639b314e6eeeee2

Observation cda1f626-b060-455b-89a8-a02f3d16b44e · inbound

KORGym: A Dynamic Game Platform for LLM Reasoning Evaluation cites this paper.

KORGym: A Dynamic Game Platform for LLM Reasoning Evaluation InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:28.095327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:28.095327Z digest=sha256:52bb23b508af140fc285f5056f7a1529f41197cb053ce6ad2a573fb0d318d851

Observation e2f50624-c532-4090-9107-576c2a735bad · inbound

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation cites this paper.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:44.591263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:44.591263Z digest=sha256:1af3eceddcc1359b7a267ab5814e6343737bac0a4e92a602c7aec26de4ed32b3

Observation 0e74cc5c-5dd6-4746-92a7-02b407ac0979 · inbound

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation cites this paper.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:55.921584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:55.921584Z digest=sha256:228b0fd75db276368307f226be96e234ce78bcf1d0c927f1731f1effd22f99cc

Observation b9921659-b815-4ea3-a041-8092867677ee · inbound

LENS: Multi-level Evaluation of Multimodal Reasoning with Large Language Models cites this paper.

LENS: Multi-level Evaluation of Multimodal Reasoning with Large Language Models InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-22T13:51:37.695934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-22T13:47:51.436258Z digest=sha256:7f92148e64d0114f22305350db822a4091c61eb3471c0fbc06a22e8029eb9072

Observation f67027c9-4458-48db-9a08-8c0e0be4206f · inbound

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs cites this paper.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:42.305127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:42.305127Z digest=sha256:f4bbef362dc89d8f72fb59d25927f8553838b16bbe083a0a3c41b3ea9a5bfb95

Observation 0994b7bf-7d19-4bf0-8e6a-0ce65e4607dc · inbound

GRIT: Teaching MLLMs to Think with Images cites this paper.

GRIT: Teaching MLLMs to Think with Images InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-22T13:31:36.139144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-22T13:29:47.529564Z digest=sha256:4f274eeb75bd32eb8d9020d1daab92320490cd0d7d402df5e072528f29a36589

Observation 3ab6f159-b65f-4550-b531-c09f3494b5ba · inbound

PhyX: Does Your Model Have the "Wits" for Physical Reasoning? cites this paper.

PhyX: Does Your Model Have the "Wits" for Physical Reasoning? InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T15:14:55.105060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:14:55.105060Z digest=sha256:862140515908112221ecf789c9d284069fbb55dd080d813c7bcffc543d1ae949

Observation c3b24cbf-3b98-4fd6-9d57-9e382a28dd35 · inbound

QuickVideo: Real-Time Long Video Understanding with System Algorithm Co-Design cites this paper.

QuickVideo: Real-Time Long Video Understanding with System Algorithm Co-Design InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:02.863981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:02.863981Z digest=sha256:c3c1960409e0169b325f65cdba7192f643b842c6dc7d4d747cb52d2925e91869

Observation 75822fd1-1674-4244-a18f-b958e6bb4628 · inbound

Bridging the Dynamic Perception Gap: Training-Free Draft Chain-of-Thought for Dynamic Multimodal Spatial Reasoning cites this paper.

Bridging the Dynamic Perception Gap: Training-Free Draft Chain-of-Thought for Dynamic Multimodal Spatial Reasoning InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:27.349532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:02:27.349532Z digest=sha256:304f64bd7694bd6391572c1414ceb49c9e77207eb9cf799fec64ddf0e5cc2e6e

Observation eb45cb7a-5c9c-48e7-83b4-dc1a04123dcb · inbound

RBench-V: A Primary Assessment for Visual Reasoning Models with Multi-modal Outputs cites this paper.

RBench-V: A Primary Assessment for Visual Reasoning Models with Multi-modal Outputs InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:58:44.144189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:58:44.144189Z digest=sha256:7dc6142db7633acaade15de5813ea45a8a7309e86ce6dd29da4370500535bb55

Observation 04e31361-0eda-4b64-b7f2-0025f048c250 · inbound

LaViDa: A Large Diffusion Language Model for Multimodal Understanding cites this paper.

LaViDa: A Large Diffusion Language Model for Multimodal Understanding InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-07T14:59:40.415254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:59:40.415254Z digest=sha256:56f5449c3e8f8262507ea0e871d8b3a70cf20db0e69f4d86c52ce62439cd5136

Observation 9886c18f-904d-4ce6-b7fa-422291e00bf2 · inbound

DetailMaster: Can Your Text-to-Image Model Handle Long Prompts? cites this paper.

DetailMaster: Can Your Text-to-Image Model Handle Long Prompts? InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:31.877954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:31.877954Z digest=sha256:cc65c6abb6e0ddf5accf86832d73443501c375b79742df810f2a951abbcd232d

Observation 8563c017-fdd0-409c-9824-a7e4b58df3d9 · inbound

SpatialScore: Towards Comprehensive Evaluation for Spatial Intelligence cites this paper.

SpatialScore: Towards Comprehensive Evaluation for Spatial Intelligence InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 114

Resolution
verified exact
local_arxiv, observed 2026-05-22T13:11:35.695264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-22T13:07:11.548885Z digest=sha256:09f3b2caf87b2e927585d9fe1b67cae50b514b5474b1eacdae667eebae7fc165

Observation a11578a7-f8a2-47b0-86d1-823c5a417955 · inbound

OCR-Reasoning Benchmark: Unveiling the True Capabilities of MLLMs in Complex Text-Rich Image Reasoning cites this paper.

OCR-Reasoning Benchmark: Unveiling the True Capabilities of MLLMs in Complex Text-Rich Image Reasoning InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T14:57:21.933270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:57:21.933270Z digest=sha256:c05afd0c48fca64daee2f334b6818c12215b627aef41351394de21cfed974303

Observation 239bc833-9097-451a-a98d-d3e16827e969 · inbound

FullFront: Benchmarking MLLMs Across the Full Front-End Engineering Workflow cites this paper.

FullFront: Benchmarking MLLMs Across the Full Front-End Engineering Workflow InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T14:51:18.223323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:51:18.223323Z digest=sha256:3a6c6b7c1d15d1796dec656456df689f5196a40ece48d86609d74e5a9e946fdb

Observation 316e5b31-197b-41f0-a4a1-8706a1e0a564 · inbound

VIBE: Annotation-Free Video-to-Text Information Bottleneck Evaluation for TL;DR cites this paper.

VIBE: Annotation-Free Video-to-Text Information Bottleneck Evaluation for TL;DR InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 48

Resolution
malformed identifier
no resolver link, observed 2026-08-07T14:51:44.389472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:51:44.389472Z digest=sha256:7eb034544694230e52be4f75d557369e44ba8d72727521fa17f195d5f11745b3

Observation 904ae12a-c11c-4182-9fbe-504249297311 · inbound

EVADE-Bench: Multimodal Benchmark for Evaluating and Enhancing Evasive Content Detection cites this paper.

EVADE-Bench: Multimodal Benchmark for Evaluating and Enhancing Evasive Content Detection InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T14:46:06.041945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:46:06.041945Z digest=sha256:31b1c18c7a32cf5ce8ed684178d10d58ae0bb21a60cb1a070a9f07e9938d9650

Observation f8f34cbe-3bb1-4204-aef1-f334d06d5fc7 · inbound

Generative RLHF-V: Learning Principles from Multi-modal Human Preference cites this paper.

Generative RLHF-V: Learning Principles from Multi-modal Human Preference InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:52.342099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:52.342099Z digest=sha256:43a69ef7cefd87ba452304df2fadcbb199a36836a42eec5852f4d6042845bcfc

Observation 1f977bd7-d0b0-4f54-af21-f70a07258c6f · inbound

MLLMs are Deeply Affected by Modality Bias cites this paper.

MLLMs are Deeply Affected by Modality Bias InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:21.513272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:21.513272Z digest=sha256:99b71898297bebcac7b1565eafb3695c92c06ab547bebe3ee62c849b677c9b81

Observation 68370baa-f6fd-4626-9c4e-34af8b7f9cf4 · inbound

Benchmarking Multimodal Knowledge Conflict for Large Multimodal Models cites this paper.

Benchmarking Multimodal Knowledge Conflict for Large Multimodal Models InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:54.741327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:54.741327Z digest=sha256:2862daa315c29fbc0ac2e14366bee5d6aef625048a1a4c9ea9485d878930a2b2

Observation 7d815d66-67ce-4078-a33f-b49ae13ce16b · inbound

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval cites this paper.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 118

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:47.630465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:47.630465Z digest=sha256:409dcefeb470999e41fc9fa8b7377535b63387d06f914f9b344f46a7f92bed53

Observation f059adeb-1584-42cc-8b4d-52f63e67b90f · inbound

Vad-R1: Towards Video Anomaly Reasoning via Perception-to-Cognition Chain-of-Thought cites this paper.

Vad-R1: Towards Video Anomaly Reasoning via Perception-to-Cognition Chain-of-Thought InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 92

Resolution
malformed identifier
no resolver link, observed 2026-08-07T14:09:16.421330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:09:16.421330Z digest=sha256:847318b259d0b71974f62a11644c0c4281349967d7378f8802a684b1d62c7a21

Observation bbe0ee83-2b7a-4ec4-b01a-5d9eb00677f7 · inbound

Seeing is Believing, but How Much? A Comprehensive Analysis of Verbalized Calibration in Vision-Language Models cites this paper.

Seeing is Believing, but How Much? A Comprehensive Analysis of Verbalized Calibration in Vision-Language Models InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:25.854052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:01:25.854052Z digest=sha256:37b29cffaa8e1207fd9ad6f83bb130cebf689d5bf8c82c268284b216088467c3

Observation d9652d4f-5803-404f-bfbe-04cbc5d6017a · inbound

Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models cites this paper.

Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T13:50:34.081389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:50:34.081389Z digest=sha256:9548478bf635bd689649a9135fe4250c9ae97314c5b58dd181bfbf4818799175

Observation d5188b38-781a-443d-9b52-42126a12fce0 · inbound

MME-Reasoning: A Comprehensive Benchmark for Logical Reasoning in MLLMs cites this paper.

MME-Reasoning: A Comprehensive Benchmark for Logical Reasoning in MLLMs InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T13:36:03.052140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:36:03.052140Z digest=sha256:b588f846475ae5b8d21e70c7481eb23733412591fcf09bf61cf4e8fbe9166508

Observation 779aaedc-eb90-4265-a5f1-bf46bf209fca · inbound

Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning? cites this paper.

Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning? InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-17T05:40:56.011410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T05:40:55.944288Z digest=sha256:98aeab355f4ee0cadf06c4cc68195fef36851a43d9728381b6ef0d725465f0cb

Observation 281c2a70-5c47-4125-9504-d3bffeb7b64b · inbound

ID-Align: RoPE-Conscious Position Remapping for Dynamic High-Resolution Adaptation in Vision-Language Models cites this paper.

ID-Align: RoPE-Conscious Position Remapping for Dynamic High-Resolution Adaptation in Vision-Language Models InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:58.230319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:58.230319Z digest=sha256:7d387fd925d446d088e68e0d6ce637f95804a3dd6f929a60d1f6e5e97135556f

Observation c0c5c59e-0ee5-4ba8-94f0-7349e1a74c8a · inbound

UI-Genie: A Self-Improving Approach for Iteratively Boosting MLLM-based Mobile GUI Agents cites this paper.

UI-Genie: A Self-Improving Approach for Iteratively Boosting MLLM-based Mobile GUI Agents InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T13:34:16.540169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:34:16.540169Z digest=sha256:bdd1aac5fade31e5c07de7ffcbe485c7e6f2fca3512e5c680d8dec83015f74f9

Observation 6e800659-f0a3-4137-b139-2382ac133c49 · inbound

Training Free Stylized Abstraction cites this paper.

Training Free Stylized Abstraction InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T13:06:22.456572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:06:22.456572Z digest=sha256:40ebe8b5ae620ef71c579843925d9abcecb683df01e2e6be9df5a5b1dfc8ad48

Observation 8046bc62-217f-40d9-83c9-4fe1946366d6 · inbound

VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos cites this paper.

VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:46.117617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:46.117617Z digest=sha256:00c9d05c43c2e3f833b779c5b5fd7405616be2866a48ee266eaec5f70f9b4165

Observation 7c83732d-0799-4372-a423-e1939376ee40 · inbound

ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding cites this paper.

ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T12:40:47.180213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:40:47.180213Z digest=sha256:f4e24810679256d3f784b7e79ce8df059c16f76d1de46bb32713e28e85917f5f

Observation d0c1c0d0-f2ce-4a7e-8433-21199a2041fd · inbound

mRAG: Elucidating the Design Space of Multi-modal Retrieval-Augmented Generation cites this paper.

mRAG: Elucidating the Design Space of Multi-modal Retrieval-Augmented Generation InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:28.211056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:42:28.211056Z digest=sha256:7bc4deaa92fa6b32297508cfb264d909c69462102e3091802dd396ee54df5bed

Observation ea093f35-2b5b-40ce-9134-c4dd340489b8 · inbound

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models cites this paper.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:50.749824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:50.749824Z digest=sha256:6af9cd8e4751d24c8fcbef59b748a22c12d042abf748fa214aed4a693730ff7e

Observation 01b07d1d-4d82-4309-8863-116f46653402 · inbound

Seeing is Not Reasoning: MVPBench for Graph-based Evaluation of Multi-path Visual Physical CoT cites this paper.

Seeing is Not Reasoning: MVPBench for Graph-based Evaluation of Multi-path Visual Physical CoT InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:36:18.458818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:36:18.458818Z digest=sha256:d8dd6ba3ddd7c42b09f25325ba119b437a4d348b9546b6561ea019ff0f1f1c49

Observation 393bbb5a-41ba-444d-872b-ce64e739b817 · inbound

Benchmarking Foundation Models for Zero-Shot Biometric Tasks cites this paper.

Benchmarking Foundation Models for Zero-Shot Biometric Tasks InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T12:33:56.192143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:33:56.192143Z digest=sha256:ce8f38a09e1d57281b7fb7ee2f76c58aa0bfd634728074505a9236c471c6d886

Observation b3ee151d-61fb-4225-b0fd-cf459e23e674 · inbound

Draw ALL Your Imagine: A Holistic Benchmark and Agent Framework for Complex Instruction-based Image Generation cites this paper.

Draw ALL Your Imagine: A Holistic Benchmark and Agent Framework for Complex Instruction-based Image Generation InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T12:18:35.370202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:18:35.370202Z digest=sha256:40c1339b31ebad0e8f6fc5c957dee3d245dcafef0cdc95b20564884df2f4b1b3

Observation fe697fdf-c8d0-45cd-88c1-95da879b32b7 · inbound

Agent-X: Evaluating Deep Multimodal Reasoning in Vision-Centric Agentic Tasks cites this paper.

Agent-X: Evaluating Deep Multimodal Reasoning in Vision-Centric Agentic Tasks InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T12:18:57.509715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:18:57.509715Z digest=sha256:11c6f8c747ddff4512cbfd66eff2086036bef9504bd19f366b4390d31cd0c67a

Observation 2e645ceb-1ac1-4c26-88ae-c40d0630f781 · inbound

MedBookVQA: A Systematic and Comprehensive Medical Benchmark Derived from Open-Access Book cites this paper.

MedBookVQA: A Systematic and Comprehensive Medical Benchmark Derived from Open-Access Book InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:05.706419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:05.706419Z digest=sha256:e2eb03b60aba8385346c64e6872112cce6448f1ad21f93974438b926e71ffb28

Observation dabf967d-32cf-406a-95d9-51cdf832db31 · inbound

Affordance Benchmark for MLLMs cites this paper.

Affordance Benchmark for MLLMs InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:56.340485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:59:56.340485Z digest=sha256:d0ef930ad3da078833eca523dbaff9d88978c9bd74e457e4ac51288255cf7934

Observation ab1a015a-c659-4ae2-979b-d101bc05bf37 · inbound

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking cites this paper.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:14.847482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:14.847482Z digest=sha256:2117dc68bf5b5271cab79fc0656e2adfda68dc1bae2b240af4b58a3af8ea897d

Observation 6242eb91-4f6f-497f-9b41-262e710bdd93 · inbound

ReFoCUS: Reinforcement-guided Frame Optimization for Contextual Understanding cites this paper.

ReFoCUS: Reinforcement-guided Frame Optimization for Contextual Understanding InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T11:51:30.225766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:51:30.225766Z digest=sha256:c197d85f12ceba709f4b5ea0ca7f7f050dbce505983afc16d58b3684baa97252

Observation 24a627c5-91d6-4e12-86c9-66262efea867 · inbound

Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement cites this paper.

Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:57.963146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:57.963146Z digest=sha256:191b98e54144b76dc0108f4aac30647370b6e127397e7368ccdd5ee006ec5341

Observation 56948934-2ff3-42cf-864d-e24574e7f022 · inbound

Improve Multi-Modal Embedding Learning via Explicit Hard Negative Gradient Amplifying cites this paper.

Improve Multi-Modal Embedding Learning via Explicit Hard Negative Gradient Amplifying InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T13:18:48.190114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:18:48.190114Z digest=sha256:7fa4c2a4d75acd3c193ce4a734a31687839664a4b4443860b3c8ee6ca9d54700

Observation 8c76ed4c-6170-46ba-b043-abac5c72b6c3 · inbound

VS-Bench: Evaluating VLMs for Strategic Abilities in Multi-Agent Environments cites this paper.

VS-Bench: Evaluating VLMs for Strategic Abilities in Multi-Agent Environments InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 89

Resolution
verified exact
local_arxiv, observed 2026-05-19T11:57:16.295831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T11:57:08.314088Z digest=sha256:1f459e51d983334b03110d152cc8a0a25670c51f5eb4ccb62350cddebfd9aa38

Observation ee514e3a-760c-49f7-8dfb-3f8be28bc8ec · inbound

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence cites this paper.

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:43.461381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:43.461381Z digest=sha256:290eb1177d19e894ebdcb9850eb19cfcbe1591a01cc4d617640937e7f907900f

Observation c394f00e-7816-4bb2-b57b-b74c3e7e0d02 · inbound

SVGenius: Benchmarking LLMs in SVG Understanding, Editing and Generation cites this paper.

SVGenius: Benchmarking LLMs in SVG Understanding, Editing and Generation InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T11:11:54.781412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:11:54.781412Z digest=sha256:5485f5f2ccd1d75302c85d5ebc9732e7487bc40a7b6d1e12a28d640d1476c6df

Observation b4727461-a149-4bd3-a45b-6ea9595f4598 · inbound

VLMs Can Aggregate Scattered Training Patches cites this paper.

VLMs Can Aggregate Scattered Training Patches InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:05.852456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:05.852456Z digest=sha256:25655e7f8323a00c784a6f3cc26866ce44bc9b27121eb39be79c1f95ed51999b

Observation b59dcb72-876f-4160-82fa-5a30bde78ca9 · inbound

Rex-Thinker: Grounded Object Referring via Chain-of-Thought Reasoning cites this paper.

Rex-Thinker: Grounded Object Referring via Chain-of-Thought Reasoning InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 76

Resolution
malformed identifier
no resolver link, observed 2026-08-07T10:55:14.911563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:14.911563Z digest=sha256:b53a47538c2c8989078445355775ab0d3b09b4aed564dfd2c6d03112fe5bad7a

Observation a93883e8-83a9-4a11-ac5c-a00642dd05e3 · inbound

MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos cites this paper.

MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T10:52:35.668975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:52:35.668975Z digest=sha256:0eb9e5977325cdba190af924bcbf7ea611fb91dd50d0cd08c4e42f261c059769

Observation 242fbcc3-9161-4688-add2-97f2000fa578 · inbound

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark cites this paper.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.821194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.821194Z digest=sha256:e1a0ca16319f157c1bbfe25881414a372189fef4db2272f52fa58c5b8f10dd97

Observation 6be6bd9a-0430-44a7-9f86-3859b73daa87 · inbound

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs cites this paper.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.468595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.468595Z digest=sha256:de32215fbe99609202f3dfa039005a835953c20fceb9622004a594d7b126c848

Observation 493fb17e-3258-499d-b49a-9c284528201d · inbound

VideoMolmo: Spatio-Temporal Grounding Meets Pointing cites this paper.

VideoMolmo: Spatio-Temporal Grounding Meets Pointing InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:31.011248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:31.011248Z digest=sha256:f8da0985bb009c1623f67a9d64f33221d388dc33026ea1d37ac2250dd90ba622

Observation 15d0e53f-db6a-4af0-b7c5-10e219a037f9 · inbound

VideoMathQA: Benchmarking Mathematical Reasoning via Multimodal Understanding in Videos cites this paper.

VideoMathQA: Benchmarking Mathematical Reasoning via Multimodal Understanding in Videos InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T10:25:36.358934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:25:36.358934Z digest=sha256:d14499919e033d62ec40726bfeec56ddfb233ace945ea81ce251dd56b33862f6

Observation 797883c6-4fa8-489b-90fa-97838cf7c6db · inbound

SIV-Bench: A Video Benchmark for Social Interaction Understanding and Reasoning cites this paper.

SIV-Bench: A Video Benchmark for Social Interaction Understanding and Reasoning InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 64

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T11:37:15.720291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T11:36:36.687324Z digest=sha256:dc760bf3973ec8c1e394602eeace22ae8b8a0bcc5ba8a6874c2ca1ab40b39843

Observation 54b1c1b9-7d54-497d-8e8f-d898cebd1c34 · inbound

Degradation-Aware Image Enhancement via Vision-Language Classification cites this paper.

Degradation-Aware Image Enhancement via Vision-Language Classification InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:23.558902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:23.558902Z digest=sha256:771993ce0c0c1ecc43cca90328ca9c5563ee0aeb98ae2d7c2b6ad470377218a5

Observation 71324b7e-4868-4669-9169-e59b8f1dc13f · inbound

PuzzleWorld: A Benchmark for Multimodal, Open-Ended Reasoning in Puzzlehunts cites this paper.

PuzzleWorld: A Benchmark for Multimodal, Open-Ended Reasoning in Puzzlehunts InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-05-19T10:47:15.135941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T10:43:02.601014Z digest=sha256:ccfbf368943478173c75d5492975127d7bd96747dd8ace0ba6b44fef38361513

Observation ef4cd2f1-7031-42bb-99f4-518c23f15f8a · inbound

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding cites this paper.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:57.032403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:57.032403Z digest=sha256:633101b879f9a6c3b350ddab996ee512d661667ee420adfee5c714849e407bc5

Observation 482ffa6f-a7bc-4d60-aff8-32ef23cb0176 · inbound

CyberV: Cybernetics for Test-time Scaling in Video Understanding cites this paper.

CyberV: Cybernetics for Test-time Scaling in Video Understanding InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:47.320450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:26:47.320450Z digest=sha256:3416a5b76d05238770bb2021638ea9ef0d3f5b42c5e52f4232f3c80792429ef9

Observation 49cb0657-4e38-4676-b54e-3acf8100c50d · inbound

GUI-Reflection: Empowering Multimodal GUI Models with Self-Reflection Behavior cites this paper.

GUI-Reflection: Empowering Multimodal GUI Models with Self-Reflection Behavior InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:59.359646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:59.359646Z digest=sha256:389bb24bb3d6dd360ae64df520e13f598e93b51ff192b30af89fc0c698f3985d

Observation a00704df-5a5f-47aa-b894-baf2e328cd43 · inbound

ReCogDrive: A Reinforced Cognitive Framework for End-to-End Autonomous Driving cites this paper.

ReCogDrive: A Reinforced Cognitive Framework for End-to-End Autonomous Driving InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-15T07:36:24.492946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T07:36:24.319361Z digest=sha256:6a5f3e5624c5f54ddd46b9cee0100d3d84470f8e3a14ad4757597b48dd1bd895

Observation 44f0f79c-8ad7-4331-afd8-bbf13190d480 · inbound

AD^2-Bench: A Hierarchical CoT Benchmark for MLLM in Autonomous Driving under Adverse Conditions cites this paper.

AD^2-Bench: A Hierarchical CoT Benchmark for MLLM in Autonomous Driving under Adverse Conditions InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T04:49:30.197911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:49:30.197911Z digest=sha256:fe95da1846fc8abeb776eb04cdfe25a87b8ab9063957877834bf8abeaa29e0c4

Observation 34043248-d0ab-47ce-a149-911caeab7419 · inbound

EditInspector: A Benchmark for Evaluation of Text-Guided Image Edits cites this paper.

EditInspector: A Benchmark for Evaluation of Text-Guided Image Edits InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:37.973647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:37.973647Z digest=sha256:fc2640bc2480dc31810cee52254853ce7a0d1b3222f969a1752e2af8c4fd5828

Observation 6d1d6a56-dde6-489d-8637-f9f54565b541 · inbound

Breaking Bad Molecules: Are MLLMs Ready for Structure-Level Molecular Detoxification? cites this paper.

Breaking Bad Molecules: Are MLLMs Ready for Structure-Level Molecular Detoxification? InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 155

Resolution
unresolved
no resolver link, observed 2026-08-07T04:17:59.889670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:17:59.889670Z digest=sha256:197e223afaabac2ff27c4961151f3a7cfa2bf5cc90812c05e65e7d294f9756a9

Observation a8344eeb-4409-476e-8514-35f8472891fe · inbound

Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs cites this paper.

Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T04:20:33.229755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:20:33.229755Z digest=sha256:7985638173ab5773d5a3f8f9ab61bba7f5d06db34346f5c71e72ac65c845757e

Observation 3cbee560-f79c-4c3e-850a-a752a3295f19 · inbound

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? cites this paper.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:34.864725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:07:34.864725Z digest=sha256:a92ead1c74c41f4c849bc1295d0a0c834380f977f11c0523512e4917e8c44bfa

Observation 2193e504-0396-4d27-8390-f69c30c0412f · inbound

Dynamic Mixture of Curriculum LoRA Experts for Continual Multimodal Instruction Tuning cites this paper.

Dynamic Mixture of Curriculum LoRA Experts for Continual Multimodal Instruction Tuning InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T04:09:30.755791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:09:30.755791Z digest=sha256:7813332ef087355ce868c592455a107d4d02f102b1d5f5bd9d57ac0e1f63da73

Observation ee80fc58-473b-4dee-ad2c-8da8c5ef9b52 · inbound

VGR: Visual Grounded Reasoning cites this paper.

VGR: Visual Grounded Reasoning InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-05-19T09:12:14.511231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T09:11:00.295700Z digest=sha256:5c0e995d650ba2f19584c58c343355b9e006b97417cbb0ec424ce4245ecb0601

Observation 6054d47b-14a4-44d6-a8ec-76c98143f6de · inbound

DinoCompanion: An Attachment-Theory Informed Multimodal Robot for Emotionally Responsive Child-AI Interaction cites this paper.

DinoCompanion: An Attachment-Theory Informed Multimodal Robot for Emotionally Responsive Child-AI Interaction InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T00:55:30.570553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:55:30.570553Z digest=sha256:bb376186e13de886fa1af8be1da7bc1e1fecde690a4f1e413f5a9ce80adfa50d

Observation fce9c156-8fc7-47df-8f60-e76731845313 · inbound

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models cites this paper.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:34.911797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:34.911797Z digest=sha256:192d0a6a878133cfc4160b3dba2d6b62d69c50cdc25e65e9ace9668add7c9365

Observation 6cbfbc00-bc4e-4090-88cc-c476612276a2 · inbound

HKD4VLM: A Progressive Hybrid Knowledge Distillation Framework for Robust Multimodal Hallucination and Factuality Detection in VLMs cites this paper.

HKD4VLM: A Progressive Hybrid Knowledge Distillation Framework for Robust Multimodal Hallucination and Factuality Detection in VLMs InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T00:41:41.321355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:41:41.321355Z digest=sha256:5820398f56cb772cf98a6f26028a88b8dcbda62c372d4841fe0e025ab6aa3aeb

Observation 39ae18ee-d0a8-474b-9222-8cddd87ad256 · inbound

AceReason-Nemotron 1.1: Advancing Math and Code Reasoning through SFT and RL Synergy cites this paper.

AceReason-Nemotron 1.1: Advancing Math and Code Reasoning through SFT and RL Synergy InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T00:41:36.479613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:41:36.479613Z digest=sha256:ee1d5ac365910336a1225bae6c22870727edc59e3f75a7288cdea110e8b9635b

Observation 12a90f74-4402-4da3-8371-388257d89423 · inbound

Dense360: Dense Understanding from Omnidirectional Panoramas cites this paper.

Dense360: Dense Understanding from Omnidirectional Panoramas InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:56.322623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:56.322623Z digest=sha256:d568100b8150abafdca1407e0740d70bb6fc195a7303bccd03bc60b4884aad9c

Observation cba8af1b-81fd-4ced-9f49-ae5ffed51561 · inbound

CLGRPO: Reasoning Ability Enhancement for Small VLMs cites this paper.

CLGRPO: Reasoning Ability Enhancement for Small VLMs InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T23:27:45.989262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:27:45.989262Z digest=sha256:ca19327d64f8740440be22c6a3a48775240093077c68225cea878c7747a7e6f2

Observation 349009e5-9f3b-4f1c-9ff6-80ab18e6d2ae · inbound

Task-Aware KV Compression For Cost-Effective Long Video Understanding cites this paper.

Task-Aware KV Compression For Cost-Effective Long Video Understanding InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:39.894521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:39.894521Z digest=sha256:b685c0c73321e587eacd75303e2c2780da528aee9220fcb27ebe4280df9bdb08

Observation 9c23b3d1-a3c8-4ad6-8de2-7dfd2f22f349 · inbound

LLaVA-Pose: Enhancing Human Pose and Action Understanding via Keypoint-Integrated Instruction Tuning cites this paper.

LLaVA-Pose: Enhancing Human Pose and Action Understanding via Keypoint-Integrated Instruction Tuning InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T22:32:45.281287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:32:45.281287Z digest=sha256:a39ad4e6e23697064f3704c36ede03b106c9b8ffa76a86297d062b2e2b6e863f

Observation d9dcd1a2-1e1c-42ab-934b-5e83343b41f0 · inbound

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling cites this paper.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:38.978885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:38.978885Z digest=sha256:7fb554a0f1ab6d769212c531e302030c5c985dbe68966bab81ab265686a735cc

Observation 853d58cd-65fb-4489-8823-a2d88a1469ae · inbound

R1-Track: Direct Application of MLLMs to Visual Object Tracking via Reinforcement Learning cites this paper.

R1-Track: Direct Application of MLLMs to Visual Object Tracking via Reinforcement Learning InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T22:18:37.640165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:18:37.640165Z digest=sha256:f3807e25b896b4f3f8fc05ca6d480947297ef911fbd73f82c98c9e73a38f1595

Observation 8f2729da-216a-4dd5-b01d-af2c50343ca5 · inbound

MARBLE: A Hard Benchmark for Multimodal Spatial Reasoning and Planning cites this paper.

MARBLE: A Hard Benchmark for Multimodal Spatial Reasoning and Planning InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T21:58:37.045728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:58:37.045728Z digest=sha256:a1a7da16390ecf0f88cbc06f9cdb8e989952112298604898616e8eadb7265733

Observation 27412aec-328a-47fb-9fbc-581d57699664 · inbound

Ovis-U1 Technical Report cites this paper.

Ovis-U1 Technical Report InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T22:01:21.610269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:01:21.610269Z digest=sha256:ec065b9ce1641980e6300b84fd10b1808a12e245260989ee60e5ad05f4b72e11

Observation 862d55c2-a190-4aa5-983e-591ccce7ca14 · inbound

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering cites this paper.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:18.924233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:18.924233Z digest=sha256:5404bd505dd6d23d9961927afe2ca9ccc5f69ee2605d80a901fb6dadd463ebb1

Observation 30a0a296-4d0d-4cc7-8987-ab2dfd049432 · inbound

ZonUI-3B: A Lightweight Vision-Language Model for Cross-Resolution GUI Grounding cites this paper.

ZonUI-3B: A Lightweight Vision-Language Model for Cross-Resolution GUI Grounding InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T21:44:51.870568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:44:51.870568Z digest=sha256:537b6deddaba7e06d8eb51b8ddbd65001c3933efc527690b46ef9505e227beac

Observation 8878ef5f-2876-42fe-ae8d-403b1ca08e06 · inbound

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World cites this paper.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 101

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:40.777286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:40.777286Z digest=sha256:7a671c130006ad3b99985f12c559b36704e7bd2e0b52efe17f567bc1a06a4756

Observation 69729ed6-3059-4b20-af91-b4a8f81f2fc5 · inbound

Just Noticeable Difference for Large Multimodal Models cites this paper.

Just Noticeable Difference for Large Multimodal Models InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T21:20:49.227807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:20:49.227807Z digest=sha256:b6f93f69218ffc6df4487aeccf10a901f101484107f9bcf5f2f59cac55cc4798

Observation 77218f32-162a-4987-874c-f41374e1bc4f · inbound

Improving the Reasoning of Multi-Image Grounding in MLLMs via Reinforcement Learning cites this paper.

Improving the Reasoning of Multi-Image Grounding in MLLMs via Reinforcement Learning InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-05-19T06:52:08.133120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T06:50:02.607136Z digest=sha256:f25866b804170612ee1f7cd0e051f4a4ba38f22783fa358017a474fc01df9eb6

Observation 22a1f710-e07d-4141-9ec1-bdb9ba20c755 · inbound

GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning cites this paper.

GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 72

Resolution
verified exact
local_arxiv, observed 2026-05-11T04:48:27.253853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-11T04:48:26.355351Z digest=sha256:e8c9b89f43de00313b8abd7fdd917219929b57e312f0e4dd7ea6643d98e94bc9

Observation a7b7c1a9-136d-472a-b324-7cd735351659 · inbound

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement cites this paper.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 105

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:10.771705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:10.771705Z digest=sha256:fb1e4a4fa057a3526b251ec3195ece0c85b955ebe54813f74c8361595457f74a

Observation 0999d1ab-1d5a-4155-9634-3b779e585c12 · inbound

Kwai Keye-VL Technical Report cites this paper.

Kwai Keye-VL Technical Report InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T20:45:10.445998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:45:10.445998Z digest=sha256:4e694523c5386f833812321bb58efb42969605e59b5e15c34104b45f2a73fabd

Observation 152ebfed-9f71-45d7-862a-04f50438a944 · inbound

RoboBrain 2.0 Technical Report cites this paper.

RoboBrain 2.0 Technical Report InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:31.131012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:31.131012Z digest=sha256:826c835803128229b85152498f6857476219e335c1bf830fc63e13b34f764843

Observation db243b1a-9ebf-45a9-931b-edfadf59a114 · inbound

Coling-UniA at SciVQA 2025: Few-Shot Example Retrieval and Confidence-Informed Ensembling for Multimodal Large Language Models cites this paper.

Coling-UniA at SciVQA 2025: Few-Shot Example Retrieval and Confidence-Informed Ensembling for Multimodal Large Language Models InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T20:35:58.842488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:35:58.842488Z digest=sha256:d2018a61494b9189d9ce6714bea8bbb967f836543ef5c2d3f4afebefd0b1f33a

Observation 9419ae14-56e3-48af-a03b-5afd9c9958e8 · inbound

ConceptMix++: Leveling the Playing Field in Text-to-Image Benchmarking via Iterative Prompt Optimization cites this paper.

ConceptMix++: Leveling the Playing Field in Text-to-Image Benchmarking via Iterative Prompt Optimization InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:05.475777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:05.475777Z digest=sha256:488efee24c362f29df3325e3c95bf2c5793677703494309f7f53adf43f7c7a4c

Observation 4e837ba1-77d1-4473-a9aa-2d5a7ee160d0 · inbound

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset cites this paper.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.553874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.553874Z digest=sha256:96ae72b511efe93e3758aa827261af1289a4d592f39679356593f6525cdcf012

Observation fd02ebad-55c7-4314-b2ce-37e4ab195079 · inbound

Foundation versus Domain-specific Models: Performance Comparison, Fusion, and Explainability in Face Recognition cites this paper.

Foundation versus Domain-specific Models: Performance Comparison, Fusion, and Explainability in Face Recognition InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T20:12:20.956555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:12:20.956555Z digest=sha256:c6cfbde68882b49557de6399079b68b4641e4fb0e3a6519878dba9fbeea4a873

Observation 9a05b92b-45ab-4985-8480-3b916fd85525 · inbound

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs cites this paper.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 103

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:12.507168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:12.507168Z digest=sha256:aed3db719e8e1a0599834fef38a7480fd208ff5449b4fd8b0c21bd27b3c1b1dd

Observation 4eaaa038-ee8b-4475-9fcb-d7a3e6d9708e · inbound

BlueLM-2.5-3B Technical Report cites this paper.

BlueLM-2.5-3B Technical Report InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-06T19:20:57.763553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:20:57.763553Z digest=sha256:f6dbb1b413b3bf13eff1a9c05d9cc2cb5f9667c6d66aa268f4daed56d4c48c14

Observation 12c77fae-44e4-4e2f-8811-b35043c4d6db · inbound

Skywork-R1V3 Technical Report cites this paper.

Skywork-R1V3 Technical Report InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T19:14:07.140185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:14:07.140185Z digest=sha256:619659665121b80081d1e6da22a5ac29f57c5b01b3c779582afd24de10b2f97a