Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-10T13:41:07.991012Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 100 of 154 outbound references and 100 inbound Pith citation observations for arXiv:2504.10479.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-10T13:41:07.991012Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-10T20:17:26.566560Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
100 of 154 outbound references displayed
External citation measurements
6
pith, observed 2026-08-05T02:28:24.338817Z
Observation 56017071-d3c3-44f2-9f57-4525753d64e0 · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 2e17a19a-faef-4fb3-89d2-93dccb5fc7e7 · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models CG-bench: Clue-grounded question answering benchmark for long video understanding
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 4cd3784e-e2b9-4c1d-973c-a9cb153bf5dc · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models The claude 3 model family: Opus, sonnet, haiku
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 63863f0e-fa66-4640-bfa0-d32424c8477e · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Program Synthesis with Large Language Models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation dacd29b5-663f-401d-ad1b-1e58422fd67b · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0fa9c2e0-f66a-476b-84f0-a2fde5a65dbe · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Qwen2.5-VL Technical Report
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 56e9debf-ff99-48e2-8b42-93b9f328610b · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Smollm-corpus
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 16bddcd9-47ee-4ec1-8499-4d73fc809b7e · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Scene text visual question answering
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3e75801c-f429-4383-8e31-44b2df9bf678 · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models An augmented benchmark dataset for geometric question answering through dual parallel text encoding
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 13f3b8a6-0ebf-4a73-8942-754f92418f5f · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models MapQA: A Dataset for Question Answering on Choropleth Maps
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 429dc09c-1502-4571-9bb8-5aeaac25600d · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation af2f885b-c983-414d-aaa1-0758eb7f8dc2 · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Are We on the Right Way for Evaluating Large Vision-Language Models?
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 8ad0c483-aaf4-418f-ab69-77dcef1b97ab · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Evaluating Large Language Models Trained on Code
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c777bb48-c43a-44be-9197-3c0d73e0aa55 · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models InternEvo: Efficient Long-sequence Large Language Model Training via Hybrid Parallelism and Redundant Sharding
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3b253375-e068-4098-b7f3-71a665c88a01 · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 56c4f3d4-7eeb-46d2-9c5e-0224e90f59ad · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Theoremqa: A theorem-driven question answering dataset
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation dca35e67-146b-42ea-926a-304bbc25abee · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e18683e8-78b7-4cc3-bf15-7280ace80e77 · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 64555317-3d21-4729-84f6-87b8917bc516 · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 961d9827-e406-43d2-8d20-f8de8d333c19 · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 227cab50-7277-48a6-be3a-21050c0ba8bc · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d93645c5-e167-4955-b7ae-aeebb07c2b9d · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Simple and effective multi-paragraph reading comprehension
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0269d5fd-ab27-485f-92ec-8843919b3837 · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Training Verifiers to Solve Math Word Problems
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 63360a63-bc28-49cd-a554-19ad2935a695 · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Opencompass: A universal evaluation platform for foundation models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation aed4d604-7f67-4fb5-acd5-19d54a362ef9 · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Grok-1.5 vision preview: Connecting the digital and physical worlds with our first multimodal model
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f6e4e7ae-6dc9-4717-8785-6343b01fb55a · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models NVLM: Open Frontier-Class Multimodal LLMs
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e404920f-6e33-477a-ad91-4c1664591cb8 · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Gemini 2.0 is now available to everyone
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 2ee03a1c-19f5-41c9-9997-a3b38fd9f75b · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Introducing gemini 2.0: our new ai model for the agentic era
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e00a2882-d65a-4442-a090-1dfebde268d7 · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 35a7f493-0b75-4099-bfd8-01d8a6304fee · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 47cad76e-e666-46ea-ba99-6a4137ee179a · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Vlmevalkit: An open-source toolkit for evaluating large multi-modality models
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 7b947b91-e3f4-487e-9520-5283c96336b4 · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models The Llama 3 Herd of Models
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a6b6f22f-c849-40be-aee7-4c23dfdc372a · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 1f651706-54cc-4dd6-907f-48668b377f05 · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Learning generative visual models from few training examples: An incremental bayesian approach tested on 101 object categories
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e1a23374-18ff-42bc-945b-9af51eaeeb19 · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 39e2b465-27e0-4d0a-8d4a-4112ab8b9bf3 · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 2a242639-a304-471e-8c5f-5b5ad78f44e1 · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models BLINK: Multimodal Large Language Models Can See but Not Perceive
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 99c40049-2da8-48ca-908d-407fabca8222 · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models G-LLaVA: Solving Geometric Problem with Multi-Modal Large Language Model
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 74a09442-949e-43fa-b0e9-8dcd2b66185e · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 7567a43a-81e0-4548-bf55-f8e0092efe2b · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models V2PE: Improving Multimodal Long-Context Capability of Vision-Language Models with Variable Visual Position Encoding
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6e854932-c0a3-48ab-bd1f-687d4ebbd5ea · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 62830ee8-01c9-4fcd-a88b-38ec23fbcc33 · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Infinity-MM: Scaling Multimodal Performance with Large-Scale and High-Quality Instruction Data
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6f218f3e-0620-4daf-b312-46e0ebea4e4b · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 279fd556-6349-42b0-959c-6fb4965345ff · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Measuring massive multitask language understanding
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 34e69bd0-2ebd-4d68-9c92-ce33ac0843c3 · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Measuring mathematical problem solving with the MATH dataset
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6c1e602a-50df-4a1e-954e-f2cfd5bdf537 · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models C-eval: A multi-level multi-discipline chinese evaluation suite for foundation models
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ad463b37-61ec-4d5f-82f7-14b4a3592bb1 · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Icdar2019 competition on scanned receipt ocr and information extraction
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 97185d93-673d-4468-ad97-ab643a72e62d · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Gqa: A new dataset for real-world visual reasoning and compositional question answering
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c5e544a3-4cd7-4dc9-a46a-cdf8e342dd3a · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models MANTIS: Interleaved Multi-Image Instruction Tuning
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ba68e2b8-5aaa-4396-8d27-03fa089ed5d5 · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 197dbeae-0714-4cb4-b632-28d1edc8729f · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Binary Classifier Optimization for Large Language Model Alignment
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ca1c5b0f-7d92-4288-b8ad-0a6c1657a248 · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Dvqa: Understanding data visualizations via question answering
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 9c4e5890-e323-4baa-91b5-86b83b4d3d67 · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models GeomVerse: A Systematic Evaluation of Large Models for Geometric Reasoning
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation af44c27e-5484-4633-b6b8-e00174edd9bc · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Referitgame: Referring to objects in photographs of natural scenes
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a0f678d5-20ab-4f89-95ec-4107f9bdf393 · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models A diagram is worth a dozen images
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation cbe6585c-29e2-4f76-98d2-f600653cae00 · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Natural questions: a benchmark for question answering research
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation efac3ba4-3f1d-4ee6-8e09-c65a3426cbcf · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models RACE: Large-scale ReAding Comprehension Dataset From Examinations
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 68f46997-cb03-46d6-9992-693bf6a94c74 · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models LLaVA-OneVision: Easy Visual Task Transfer
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6ffa1d7f-aaee-47f1-9efa-8d2f643abf52 · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation cde53698-b607-4154-bbb3-3f3c3d1803a7 · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models R-Bench: Are your Large Multimodal Model Robust to Real-world Corruptions?
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation bc7aac1e-221a-4981-9a6c-2daedea49371 · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models CMMLU: Measuring massive multitask language understanding in Chinese
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 96dde0f3-38d0-406b-b44d-5595cf56c437 · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models VideoChat: Chat-Centric Video Understanding
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ea763f14-65cd-43f4-8f5d-e3f4e8522369 · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Mvbench: A comprehensive multi-modal video understanding benchmark
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e9259743-9772-4012-a169-a6436c8abd83 · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Mvitv2: Improved multiscale vision transformers for classification and detection
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a4326a14-ba66-4440-b8f4-2269e7f0dee1 · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Evaluating object hallucination in large vision-language models
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3984b339-064a-43db-9912-503af55cc6e9 · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5c37d1f1-2132-4333-af88-222f777f6598 · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 7e3c5189-a7ae-4b70-aecc-3c8789e7dcf6 · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Let’s verify step by step
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 112be398-657f-4b46-a2cc-051d2e6efc37 · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Vila: On pre-training for visual language models
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 027b7536-7492-4f42-9d0e-d2ceb2612068 · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models CLEVR-Math: A Dataset for Compositional Language, Visual and Mathematical Reasoning
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0e44c233-ef2e-47ed-8809-536391a18559 · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Visual instruction tuning
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation cacf4376-3d67-4178-b69f-3dd2c259453a · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Grounding dino: Marrying dino with grounded pre-training for open-set object detection
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 4c8570eb-9042-47be-b7d3-e4241449b233 · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models MMBench: Is Your Multi-modal Model an All-around Player?
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation cbc8a068-f9e8-4fda-8e1c-d49fe18adf14 · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 16e18c3b-2b72-4467-b02b-d442d328f963 · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Acemath: Advancing frontier math reasoning with post-training and reward modeling
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 78546f28-51b5-41e9-b238-f48203b0cc7f · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c6f1bbd5-7f6b-4ba7-ac2d-974bb218370d · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Scp-116k: A high-quality problem-solution dataset and a generalized pipeline for automated extraction in the higher education science domain
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 4aa5ed8b-6816-413f-b710-1e86b5f59be5 · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation edf6f1d4-dce2-49f9-a15c-43dbd9a766a7 · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Inter-GPS: Interpretable Geometry Problem Solving with Formal Language and Symbolic Reasoning
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 9a5922a2-f8e7-4a2f-973d-420edea72035 · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Learn to explain: Multimodal reasoning via thought chains for science question answering
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 16e0f5ce-ae4c-47e4-a1fb-bc94f4fbe4aa · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models IconQA: A New Benchmark for Abstract Diagram Understanding and Visual Language Reasoning
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f4566433-df2b-4547-be0c-63e97c43124e · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 50d1bb8a-1bda-4b21-8f31-ac3177e5c7f9 · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models BlueLM-V-3B: Algorithm and System Co-Design for Multimodal Large Language Models on Mobile Devices
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 2b254b4a-fbad-4a33-8c1f-dc89cd8c24f3 · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models WildVision: Evaluating Vision-Language Models in the Wild with Human Preferences
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3296f90c-8885-472d-a340-987a6d7f00a5 · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Improve Mathematical Reasoning in Language Models by Automated Process Supervision
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6f9fd756-dd20-4ea3-bf3f-de1e663be830 · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Generation and comprehension of unambiguous object descriptions
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 7c025a5f-c007-4400-a17c-1a43417b5062 · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models SmolVLM: Redefining small and efficient multimodal models
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e2b99258-9630-4cc9-98c8-43525eddabcb · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Ok-vqa: A visual question answering benchmark requiring external knowledge
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 7a9a8293-78e7-4bf0-ab5e-7b079d1fadb9 · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Chartqa: A benchmark for question answering about charts with visual and logical reasoning
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 1a202c3a-cd35-4fe1-aaa8-f723909d9df5 · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Infographicvqa
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation fc41b8a2-3aa1-4bda-8acc-14efe6d738f8 · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Docvqa: A dataset for vqa on document images
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ee956037-fb1a-4883-9db4-5f2edfa415a8 · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models LLM Critics Help Catch LLM Bugs
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 9ff29946-03b1-4140-bb99-fcbb4ebea566 · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models MMIU: Multimodal Multi-image Understanding for Evaluating Large Vision-Language Models
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 53ffb316-dd3d-4bcd-9465-b5f05e34d038 · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Ocr-vqa: Visual question answering by reading text in images
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation aa8fb495-d1be-44c2-9a5b-44455e210b93 · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Gpt-4v(ision) system card
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 2965058a-3ca8-4656-ad6a-724504d790ef · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Unresolved cited work
Reference 98
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 23cea178-7685-4b9a-b84e-6bc3ed66b0a6 · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Gpt-4o system card
Reference 99
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 91ef0e03-403d-4559-b0fb-82f687bd6a8f · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation eafb055f-075c-42e3-9559-2f6650fb124f · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models UI-TARS: Pioneering Automated GUI Interaction with Native Agents
Reference 101
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 26ecea0e-3d97-473e-863f-5c16f2e232c2 · outbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Direct preference optimization: Your language model is secretly a reward model.Advances in Neural Information Processing Systems, 36
Reference 102
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 05f338c6-f496-4ed5-a922-455affbcbc31 · inbound
Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 124
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 03d07df6-49f5-4699-b8b3-d34a45cb49ec · inbound
Compositional Generative Model of Unbounded 4D Cities InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 123
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b17fde1e-7c7c-4d82-bd6b-44a9f5314801 · inbound
Perception Encoder: The best visual embeddings are not at the output of the network InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 168
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f5318b83-0017-4777-817a-fd88113da0cd · inbound
Benchmarking Large Vision-Language Models on Fine-Grained Image Tasks: A Comprehensive Evaluation InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 1a098fb1-db94-4d62-9439-b3f194309c96 · inbound
DeepEyes: Incentivizing "Thinking with Images" via Reinforcement Learning InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation cda1f626-b060-455b-89a8-a02f3d16b44e · inbound
KORGym: A Dynamic Game Platform for LLM Reasoning Evaluation InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2f50624-c532-4090-9107-576c2a735bad · inbound
VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e74cc5c-5dd6-4746-92a7-02b407ac0979 · inbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9921659-b815-4ea3-a041-8092867677ee · inbound
LENS: Multi-level Evaluation of Multimodal Reasoning with Large Language Models InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f67027c9-4458-48db-9a08-8c0e0be4206f · inbound
STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0994b7bf-7d19-4bf0-8e6a-0ce65e4607dc · inbound
GRIT: Teaching MLLMs to Think with Images InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3ab6f159-b65f-4550-b531-c09f3494b5ba · inbound
PhyX: Does Your Model Have the "Wits" for Physical Reasoning? InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3b24cbf-3b98-4fd6-9d57-9e382a28dd35 · inbound
QuickVideo: Real-Time Long Video Understanding with System Algorithm Co-Design InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75822fd1-1674-4244-a18f-b958e6bb4628 · inbound
Bridging the Dynamic Perception Gap: Training-Free Draft Chain-of-Thought for Dynamic Multimodal Spatial Reasoning InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb45cb7a-5c9c-48e7-83b4-dc1a04123dcb · inbound
RBench-V: A Primary Assessment for Visual Reasoning Models with Multi-modal Outputs InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04e31361-0eda-4b64-b7f2-0025f048c250 · inbound
LaViDa: A Large Diffusion Language Model for Multimodal Understanding InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9886c18f-904d-4ce6-b7fa-422291e00bf2 · inbound
DetailMaster: Can Your Text-to-Image Model Handle Long Prompts? InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8563c017-fdd0-409c-9824-a7e4b58df3d9 · inbound
SpatialScore: Towards Comprehensive Evaluation for Spatial Intelligence InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 114
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a11578a7-f8a2-47b0-86d1-823c5a417955 · inbound
OCR-Reasoning Benchmark: Unveiling the True Capabilities of MLLMs in Complex Text-Rich Image Reasoning InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 239bc833-9097-451a-a98d-d3e16827e969 · inbound
FullFront: Benchmarking MLLMs Across the Full Front-End Engineering Workflow InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 316e5b31-197b-41f0-a4a1-8706a1e0a564 · inbound
VIBE: Annotation-Free Video-to-Text Information Bottleneck Evaluation for TL;DR InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 904ae12a-c11c-4182-9fbe-504249297311 · inbound
EVADE-Bench: Multimodal Benchmark for Evaluating and Enhancing Evasive Content Detection InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8f34cbe-3bb1-4204-aef1-f334d06d5fc7 · inbound
Generative RLHF-V: Learning Principles from Multi-modal Human Preference InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f977bd7-d0b0-4f54-af21-f70a07258c6f · inbound
MLLMs are Deeply Affected by Modality Bias InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68370baa-f6fd-4626-9c4e-34af8b7f9cf4 · inbound
Benchmarking Multimodal Knowledge Conflict for Large Multimodal Models InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d815d66-67ce-4078-a33f-b49ae13ce16b · inbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 118
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f059adeb-1584-42cc-8b4d-52f63e67b90f · inbound
Vad-R1: Towards Video Anomaly Reasoning via Perception-to-Cognition Chain-of-Thought InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bbe0ee83-2b7a-4ec4-b01a-5d9eb00677f7 · inbound
Seeing is Believing, but How Much? A Comprehensive Analysis of Verbalized Calibration in Vision-Language Models InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9652d4f-5803-404f-bfbe-04cbc5d6017a · inbound
Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5188b38-781a-443d-9b52-42126a12fce0 · inbound
MME-Reasoning: A Comprehensive Benchmark for Logical Reasoning in MLLMs InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 779aaedc-eb90-4265-a5f1-bf46bf209fca · inbound
Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning? InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 281c2a70-5c47-4125-9504-d3bffeb7b64b · inbound
ID-Align: RoPE-Conscious Position Remapping for Dynamic High-Resolution Adaptation in Vision-Language Models InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0c5c59e-0ee5-4ba8-94f0-7349e1a74c8a · inbound
UI-Genie: A Self-Improving Approach for Iteratively Boosting MLLM-based Mobile GUI Agents InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e800659-f0a3-4137-b139-2382ac133c49 · inbound
Training Free Stylized Abstraction InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8046bc62-217f-40d9-83c9-4fe1946366d6 · inbound
VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c83732d-0799-4372-a423-e1939376ee40 · inbound
ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0c1c0d0-f2ce-4a7e-8433-21199a2041fd · inbound
mRAG: Elucidating the Design Space of Multi-modal Retrieval-Augmented Generation InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea093f35-2b5b-40ce-9134-c4dd340489b8 · inbound
Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01b07d1d-4d82-4309-8863-116f46653402 · inbound
Seeing is Not Reasoning: MVPBench for Graph-based Evaluation of Multi-path Visual Physical CoT InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 393bbb5a-41ba-444d-872b-ce64e739b817 · inbound
Benchmarking Foundation Models for Zero-Shot Biometric Tasks InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3ee151d-61fb-4225-b0fd-cf459e23e674 · inbound
Draw ALL Your Imagine: A Holistic Benchmark and Agent Framework for Complex Instruction-based Image Generation InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe697fdf-c8d0-45cd-88c1-95da879b32b7 · inbound
Agent-X: Evaluating Deep Multimodal Reasoning in Vision-Centric Agentic Tasks InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e645ceb-1ac1-4c26-88ae-c40d0630f781 · inbound
MedBookVQA: A Systematic and Comprehensive Medical Benchmark Derived from Open-Access Book InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dabf967d-32cf-406a-95d9-51cdf832db31 · inbound
Affordance Benchmark for MLLMs InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab1a015a-c659-4ae2-979b-d101bc05bf37 · inbound
GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6242eb91-4f6f-497f-9b41-262e710bdd93 · inbound
ReFoCUS: Reinforcement-guided Frame Optimization for Contextual Understanding InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24a627c5-91d6-4e12-86c9-66262efea867 · inbound
Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56948934-2ff3-42cf-864d-e24574e7f022 · inbound
Improve Multi-Modal Embedding Learning via Explicit Hard Negative Gradient Amplifying InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c76ed4c-6170-46ba-b043-abac5c72b6c3 · inbound
VS-Bench: Evaluating VLMs for Strategic Abilities in Multi-Agent Environments InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ee514e3a-760c-49f7-8dfb-3f8be28bc8ec · inbound
SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c394f00e-7816-4bb2-b57b-b74c3e7e0d02 · inbound
SVGenius: Benchmarking LLMs in SVG Understanding, Editing and Generation InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4727461-a149-4bd3-a45b-6ea9595f4598 · inbound
VLMs Can Aggregate Scattered Training Patches InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b59dcb72-876f-4160-82fa-5a30bde78ca9 · inbound
Rex-Thinker: Grounded Object Referring via Chain-of-Thought Reasoning InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a93883e8-83a9-4a11-ac5c-a00642dd05e3 · inbound
MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 242fbcc3-9161-4688-add2-97f2000fa578 · inbound
Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6be6bd9a-0430-44a7-9f86-3859b73daa87 · inbound
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 493fb17e-3258-499d-b49a-9c284528201d · inbound
VideoMolmo: Spatio-Temporal Grounding Meets Pointing InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15d0e53f-db6a-4af0-b7c5-10e219a037f9 · inbound
VideoMathQA: Benchmarking Mathematical Reasoning via Multimodal Understanding in Videos InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 797883c6-4fa8-489b-90fa-97838cf7c6db · inbound
SIV-Bench: A Video Benchmark for Social Interaction Understanding and Reasoning InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 54b1c1b9-7d54-497d-8e8f-d898cebd1c34 · inbound
Degradation-Aware Image Enhancement via Vision-Language Classification InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71324b7e-4868-4669-9169-e59b8f1dc13f · inbound
PuzzleWorld: A Benchmark for Multimodal, Open-Ended Reasoning in Puzzlehunts InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ef4cd2f1-7031-42bb-99f4-518c23f15f8a · inbound
Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 482ffa6f-a7bc-4d60-aff8-32ef23cb0176 · inbound
CyberV: Cybernetics for Test-time Scaling in Video Understanding InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49cb0657-4e38-4676-b54e-3acf8100c50d · inbound
GUI-Reflection: Empowering Multimodal GUI Models with Self-Reflection Behavior InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a00704df-5a5f-47aa-b894-baf2e328cd43 · inbound
ReCogDrive: A Reinforced Cognitive Framework for End-to-End Autonomous Driving InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 44f0f79c-8ad7-4331-afd8-bbf13190d480 · inbound
AD^2-Bench: A Hierarchical CoT Benchmark for MLLM in Autonomous Driving under Adverse Conditions InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34043248-d0ab-47ce-a149-911caeab7419 · inbound
EditInspector: A Benchmark for Evaluation of Text-Guided Image Edits InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d1d6a56-dde6-489d-8637-f9f54565b541 · inbound
Breaking Bad Molecules: Are MLLMs Ready for Structure-Level Molecular Detoxification? InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 155
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8344eeb-4409-476e-8514-35f8472891fe · inbound
Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3cbee560-f79c-4c3e-850a-a752a3295f19 · inbound
VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2193e504-0396-4d27-8390-f69c30c0412f · inbound
Dynamic Mixture of Curriculum LoRA Experts for Continual Multimodal Instruction Tuning InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee80fc58-473b-4dee-ad2c-8da8c5ef9b52 · inbound
VGR: Visual Grounded Reasoning InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6054d47b-14a4-44d6-a8ec-76c98143f6de · inbound
DinoCompanion: An Attachment-Theory Informed Multimodal Robot for Emotionally Responsive Child-AI Interaction InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fce9c156-8fc7-47df-8f60-e76731845313 · inbound
Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6cbfbc00-bc4e-4090-88cc-c476612276a2 · inbound
HKD4VLM: A Progressive Hybrid Knowledge Distillation Framework for Robust Multimodal Hallucination and Factuality Detection in VLMs InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39ae18ee-d0a8-474b-9222-8cddd87ad256 · inbound
AceReason-Nemotron 1.1: Advancing Math and Code Reasoning through SFT and RL Synergy InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12a90f74-4402-4da3-8371-388257d89423 · inbound
Dense360: Dense Understanding from Omnidirectional Panoramas InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cba8af1b-81fd-4ced-9f49-ae5ffed51561 · inbound
CLGRPO: Reasoning Ability Enhancement for Small VLMs InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 349009e5-9f3b-4f1c-9ff6-80ab18e6d2ae · inbound
Task-Aware KV Compression For Cost-Effective Long Video Understanding InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c23b3d1-a3c8-4ad6-8de2-7dfd2f22f349 · inbound
LLaVA-Pose: Enhancing Human Pose and Action Understanding via Keypoint-Integrated Instruction Tuning InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9dcd1a2-1e1c-42ab-934b-5e83343b41f0 · inbound
Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 853d58cd-65fb-4489-8823-a2d88a1469ae · inbound
R1-Track: Direct Application of MLLMs to Visual Object Tracking via Reinforcement Learning InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f2729da-216a-4dd5-b01d-af2c50343ca5 · inbound
MARBLE: A Hard Benchmark for Multimodal Spatial Reasoning and Planning InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27412aec-328a-47fb-9fbc-581d57699664 · inbound
Ovis-U1 Technical Report InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 862d55c2-a190-4aa5-983e-591ccce7ca14 · inbound
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30a0a296-4d0d-4cc7-8987-ab2dfd049432 · inbound
ZonUI-3B: A Lightweight Vision-Language Model for Cross-Resolution GUI Grounding InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8878ef5f-2876-42fe-ae8d-403b1ca08e06 · inbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 101
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69729ed6-3059-4b20-af91-b4a8f81f2fc5 · inbound
Just Noticeable Difference for Large Multimodal Models InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77218f32-162a-4987-874c-f41374e1bc4f · inbound
Improving the Reasoning of Multi-Image Grounding in MLLMs via Reinforcement Learning InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 22a1f710-e07d-4141-9ec1-bdb9ba20c755 · inbound
GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a7b7c1a9-136d-472a-b324-7cd735351659 · inbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 105
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0999d1ab-1d5a-4155-9634-3b779e585c12 · inbound
Kwai Keye-VL Technical Report InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 152ebfed-9f71-45d7-862a-04f50438a944 · inbound
RoboBrain 2.0 Technical Report InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db243b1a-9ebf-45a9-931b-edfadf59a114 · inbound
Coling-UniA at SciVQA 2025: Few-Shot Example Retrieval and Confidence-Informed Ensembling for Multimodal Large Language Models InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9419ae14-56e3-48af-a03b-5afd9c9958e8 · inbound
ConceptMix++: Leveling the Playing Field in Text-to-Image Benchmarking via Iterative Prompt Optimization InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e837ba1-77d1-4473-a9aa-2d5a7ee160d0 · inbound
BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd02ebad-55c7-4314-b2ce-37e4ab195079 · inbound
Foundation versus Domain-specific Models: Performance Comparison, Fusion, and Explainability in Face Recognition InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a05b92b-45ab-4985-8480-3b916fd85525 · inbound
VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 103
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4eaaa038-ee8b-4475-9fcb-d7a3e6d9708e · inbound
BlueLM-2.5-3B Technical Report InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12c77fae-44e4-4e2f-8811-b35043c4d6db · inbound
Skywork-R1V3 Technical Report InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.