Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-10T13:23:57.588851Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 100 of 300 outbound references and 100 inbound Pith citation observations for arXiv:2412.05271.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-10T13:23:57.588851Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-10T23:07:14.402030Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
100 of 300 outbound references displayed
External citation measurements
14
pith, observed 2026-08-05T02:28:24.338817Z
Observation f6c45375-722c-4cfa-ac61-f9e4c0d07f36 · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 354b7464-ff37-4753-ad03-5c415cbf0441 · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Tallyqa: Answering complex counting questions
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 93d80056-95b1-492c-94ad-be8acfa3505f · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling GPT-4 Technical Report
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9fba97ac-8164-4ec5-ab38-540fb7d35e16 · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Wave-ui.https://huggingface.co/datasets/agentsea/wave-ui
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 323706d6-fa24-480b-ac17-bda0a7906393 · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Flamingo: a visual language model for few-shot learning
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 44c6580b-d82f-468d-8fd1-3c31b82425bb · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling MathQA: Towards Interpretable Math Word Problem Solving with Operation-Based Formalisms
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e0e95554-2034-4cac-9471-b5bcb336e11c · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling CG-bench: Clue-grounded question answering benchmark for long video understanding
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 46d8af48-b448-47f6-9c96-adc88e7ff32c · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling The claude 3 model family: Opus, sonnet, haiku
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d8f74ace-34ed-467e-aa61-059efcbdbd07 · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Program Synthesis with Large Language Models
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 624abb5d-089f-4e24-be61-b1e657ff8799 · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Scanqa: 3d question answering for spatial scene understanding
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation fa35076e-1a37-480f-a103-61cddbd93720 · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Layer Normalization
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 4714b379-b938-4275-8412-a754b91415ce · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling UIBert: Learning Generic Multimodal Representations for UI Understanding
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 09e339bd-a6fd-402b-8182-6e3b6b40d1ff · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 8e6f0162-5275-409f-b560-26a0c7f33847 · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Beijing Anjie Zhihe Technology Co
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7f6d035e-ec60-4786-9d2c-f490c62ea316 · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Vqa-med: Overview of the medical visual question answering task at imageclef 2019
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ea95d6a2-7514-4ea7-a452-c8045913625b · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Are we done with ImageNet?
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0bb7bda7-fd34-4aba-becf-74f2a1b0f875 · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Scene text visual question answering
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7121e350-55b6-465f-8688-93b65c912694 · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Coco-stuff: Thing and stuff classes in context
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e101ffbc-5196-4ad1-923a-711163632721 · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling InternLM2 Technical Report
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a4363e8f-6050-4b32-8e09-e2d4012c3416 · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling An augmented benchmark dataset for geometric question answering through dual parallel text encoding
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation dd3ab7b6-9103-480c-94e6-fa775efba9b0 · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling openai summarize tldr dataset
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6f6b7ac7-e5ca-4054-b417-548da7b31f1f · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling AMEX: Android Multi-annotation Expo Dataset for Mobile GUI Agents
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation fb483837-30a6-4fa4-9425-392ee7ade3e7 · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling MapQA: A Dataset for Question Answering on Choropleth Maps
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c53053c5-01e5-4ccf-8047-9d75d2bca10e · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling GUI-World: A Video Benchmark and Dataset for Multimodal GUI-oriented Understanding
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 60562b2e-a27c-4d56-b8c2-6b73887b1e0f · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 659889a3-271d-4631-9648-860b79fec7d1 · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling UniGeo: Unifying Geometry Logical Reasoning via Reformulating Mathematical Expression
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 127d7c43-50e8-49af-bbac-9ac9e201ae31 · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1f6485f8-3e12-4f5e-97e5-41ac988bff8e · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Are We on the Right Way for Evaluating Large Vision-Language Models?
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f3a0cecc-bea8-4593-a9bb-c07b62f3935f · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling ShareGPT4V: Improving Large Multi-Modal Models with Better Captions
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 235ff406-5120-49e3-9bd1-921fedb6817a · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7993ecb8-ec6e-47c0-b96e-09826a1b3e5c · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Evaluating Large Language Models Trained on Code
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 22fcdd29-2720-4847-99da-eb49e1a98a7b · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling A simple framework for contrastive learning of visual representations
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ee10c80f-3f0a-4ca0-b821-35c3987f5ae8 · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Theoremqa: A theorem-driven question answering dataset
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 8e3386e5-9654-44b0-99f9-8f203cd8f7e5 · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling LongLoRA: Efficient Fine-tuning of Long-Context Large Language Models
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3d589368-0d9b-4329-bab3-856bd93828e0 · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 32d5ffd0-f0f2-439d-ba2e-b0117d08fd83 · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b8a88a99-2a9f-45e9-ab5a-1e92c7b096b4 · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 40bb2c53-e9d1-49b9-855f-a2b08239daea · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 77e3eb98-8e06-4249-a49c-28d457da9b53 · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Complicated Table Structure Recognition
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2072faca-4cc8-4def-8c17-493f25100f25 · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Icdar2019 robust reading challenge on arbitrary-shaped text-rrc-art
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 8442830c-d1a9-448f-9194-58572393f992 · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Scaling instruction-finetuned language models.Journal of Machine Learning Research, 25(70):1–53
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5d768e4e-9c07-435f-8abe-2c3758eae32e · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Simple and effective multi-paragraph reading comprehension
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 44bf3c12-1781-47b2-af9c-6c485ee89c51 · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Training Verifiers to Solve Math Word Problems
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation cf974630-4a7d-480c-bae4-e9fdbb75f9d2 · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Free dolly: Introducing the world’s first truly open instruction-tuned llm.Company Blog of Databricks
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 95ba00ff-1a82-44d4-8d0d-2b02e1fddd1a · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Mmsegmentation: Openmmlab semantic segmentation toolbox and benchmark
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 47acdf29-cdb4-4dd7-a7e4-4fa9d700e13e · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Opencompass: A universal evaluation platform for foundation models
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 55afcede-c500-4d5a-831c-724fe53fda83 · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Grok-1.5 vision preview: Connecting the digital and physical worlds with our first multimodal model
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 10522019-f25c-41d5-9fb1-e405a54f695c · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling UltraFeedback: Boosting Language Models with Scaled AI Feedback
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f3cd8071-a568-4027-827a-498bba347186 · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling 15M Multimodal Facial Image-Text Dataset
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e43d73ea-e6e2-4505-95b4-506290f133e8 · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling NVLM: Open Frontier-Class Multimodal LLMs
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d9215db8-6b1a-4f96-94c6-4b1a1308751d · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Visual dialog
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1f38ce5d-6752-4d25-a716-95b6b5740df3 · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Deep visual template-free form parsing
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5e330b7a-74b4-4e11-bf4b-817f6e0fb2ab · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Scaling vision transformers to 22 billion parameters
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7286b563-8a32-4138-aa2c-422486401088 · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e83a9c2e-b407-44d3-8d3c-e83bb8c36cc8 · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Rico: A mobile app dataset for building data-driven design applications
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d7f5bb5b-eab9-4755-9c7b-5da408ac3dd0 · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Imagenet: A large-scale hierarchical image database
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation bed1a26e-b107-425b-8654-6f8d9aaae828 · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Mind2web: Towards a generalist agent for the web.Advances in Neural Information Processing Systems, 36
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1fdc4fd8-47a4-4c8e-8316-1f1e03ad236e · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Enhancing Chat Language Models by Scaling High-quality Instructional Conversations
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6310e289-b803-4cb2-b21a-858594f9b6ad · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Vintern-1B: An Efficient Multimodal Large Language Model for Vietnamese
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation caf1b279-98f3-42ab-99e1-0add0841ad00 · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9a81c0a4-de17-496c-b52b-09e5ae113c23 · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling An image is worth 16x16 words: Transformers for image recognition at scale
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a2a7b757-325b-4b77-bb96-7764b6a85952 · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Glm: General language model pretraining with autoregressive blank infilling
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6dbff460-5d56-49d6-b8d8-fab692dc5b19 · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Vlmevalkit: An open-source toolkit for evaluating large multi-modality models
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3170f94b-7a23-46b3-b90b-751fc514e2d5 · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling The Llama 3 Herd of Models
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0bbc87ab-4cc4-4e82-851a-e02b4159ce21 · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 66de3b3f-90be-48f2-8765-9a3ed77d7fe2 · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Eva: Exploring the limits of masked visual representation learning at scale
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 62238b76-6a20-4a0d-8d56-4e96b07e5dc9 · outbound
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e638a4f1-5d1e-4f00-9214-025ed6f888cc · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 4ba4a556-a5f4-46d4-b39d-05c6a1cae324 · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b97e81f7-0bde-406c-b183-6b7e09787a8c · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling BLINK: Multimodal Large Language Models Can See but Not Perceive
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7fedb350-47bb-4f9a-9359-3913c8e6be31 · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5339a34a-c7b7-4878-b74e-45b760824bde · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Overview of the imageclef 2015 medical classification task
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation fc442fde-599a-4e73-85dc-09b7afc08aa7 · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Glaive code assistant v3 dataset
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a38aa9de-01fc-437f-a34c-796538a9b56b · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7e44df2f-7b5f-4361-abe4-6e2430872b61 · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Wukong: A 100 million large-scale chinese cross-modal pre-training benchmark.Advances in Neural Information Processing Systems, 35:26418–26431
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5fae675e-72b8-49f3-8c13-71fc57915129 · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Infinity-MM: Scaling Multimodal Performance with Large-Scale and High-Quality Instruction Data
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a3d66bc5-602d-4f6e-833d-809142c9cb8b · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 85960aff-c758-4087-be01-70127bcd6fee · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Eaten: Entity-aware attention for single shot visual text extraction
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 67a82fb9-b2d9-4c93-a750-623edba794ac · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Synthetic data for text localisation in natural images
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 699a2106-755f-48b2-9e04-9a632f4f974e · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e6bfa9cf-62c2-40fc-91be-52d0081facb4 · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling WanJuan: A Comprehensive Multimodal Dataset for Advancing English and Chinese Large Models
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d0e8539c-9254-4489-90ac-7d2fd41c2461 · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Icpr2018 contest on robust reading for multi-type web images
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 4bf4431e-000a-445a-8ce5-748ec76f7c19 · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling PathVQA: 30000+ Questions for Medical Visual Question Answering
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a5520a15-8b69-4963-bf98-1098b9fc63d0 · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling The many faces of robustness: A critical analysis of out-of-distribution generalization
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b92d0459-6b89-4b34-8e2f-4f09117617e7 · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Measuring massive multitask language understanding
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation df2856e6-2c71-4da3-992e-ebb5d925fba4 · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Measuring mathematical problem solving with the MATH dataset
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation cfd386e3-f8ed-4be4-aef3-019c4f0c951b · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Natural adversarial examples
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1866826c-40cb-464d-baf8-ad6367a71bad · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling understanding
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d695a355-98d6-48cb-b35b-c526d39faaae · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Parsynth-ocr-200k
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9c88df7b-8560-497e-a806-0816a2ed83c7 · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Unnatural Instructions: Tuning Language Models with (Almost) No Human Labor
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7c2ade04-1f1a-4b33-9f6c-9ddf43dfba83 · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Koniq-10k: An ecologically valid database for deep learning of blind image quality assessment.IEEE Transactions on Image Processing, 29:4041–4056
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f9cda3eb-b097-4f41-814d-e2911563940b · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling ScreenQA: Large-Scale Question-Answer Pairs over Mobile App Screenshots
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 36eb7878-7b97-440b-86f7-8d44812fc120 · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 15dd39bf-a818-44c5-916a-6b13c7ab6f3c · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Medical-diff-vqa: a large-scale medical dataset for difference visual question answering on chest x-ray images.PhysioNet
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6f964294-f194-4733-88d5-be647956b2a8 · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Movienet: A holistic dataset for movie understanding
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2f3cc5e8-6445-4223-bfea-2f528fcd8dd9 · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling C-eval: A multi-level multi-discipline chinese evaluation suite for foundation models
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 21468330-e7e7-4cca-9bb7-6c2f79fa3cef · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Icdar2019 competition on scanned receipt ocr and information extraction
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation cab861df-f609-4b41-9152-7f9c9c61b937 · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Gqa: A new dataset for real-world visual reasoning and compositional question answering
Reference 98
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6ef5c445-9ec3-44d0-bcea-6883d0311021 · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Egotaskqa: Understanding human tasks in egocentric videos.Advances in Neural Information Processing Systems, 35:3343–3360
Reference 99
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d3cb90da-ce14-4dc8-937c-e425142bc3da · outbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling MANTIS: Interleaved Multi-Image Instruction Tuning
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation dbf3dc5f-04ef-452f-b47e-d9e69ed2731a · inbound
Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 4a4005a2-c553-4284-93bf-4d199e027499 · inbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe5de7f2-d8da-4945-9db7-4404a8086643 · inbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 4b69a1b7-2758-4655-b24f-0ec49c7eaa31 · inbound
CaReBench: A Fine-Grained Benchmark for Video Captioning and Retrieval Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db4489fd-e6fb-47cb-98ff-2ea6d77c59d0 · inbound
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d5cdff0d-953d-440d-997d-61b082553ac0 · inbound
HLV-1K: A Large-scale Hour-Long Video Benchmark for Time-Specific Long Video Understanding Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fdfc9d99-3ca6-4c53-8d07-d375493f4258 · inbound
Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 087cd7be-d6e1-4c05-b617-a14ff3b46f4a · inbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c85b112d-5f3d-4c60-b896-0c86dd59248e · inbound
LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e14bc7da-57d3-437f-930e-d95c1e01be9e · inbound
Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07d3c61f-97c8-43a9-b894-6fb073d2041b · inbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f56fa86e-db31-4618-a2a3-65e0feffd7b9 · inbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03ae5b0e-f7eb-4be8-9daf-d4fa9ce6d87b · inbound
Omni-Emotion: Extending Video MLLM with Detailed Face and Audio Modeling for Multimodal Emotion Analysis Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fdacd6b9-e11a-4829-a704-0402494f65ce · inbound
AdaFV: Rethinking of Visual-Language alignment for VLM acceleration Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d12f5c2c-075e-46f1-a857-0347cf2ed681 · inbound
A Simple Aerial Detection Baseline of Multimodal Language Models Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8514a535-74bc-4bba-ac64-6b00fd3a5803 · inbound
Is your LLM trapped in a Mental Set? Investigative study on how mental sets affect the reasoning capabilities of LLMs Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fdbc36e7-7bd3-4698-8f73-99efc9412160 · inbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9bc5ee28-ccce-48cb-8d24-fbf9ab89ec04 · inbound
VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6d0295fd-ec2e-40ef-b063-7e45bdf771de · inbound
OCSU: Optical Chemical Structure Understanding for Molecule-centric Scientific Discovery Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bba54c0c-287d-494f-aa6e-d45ed753f2dc · inbound
Ocean-OCR: Towards General OCR Application via a Vision-Language Model Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a099e45-6051-4743-a87f-585d228c25a1 · inbound
AIN: The Arabic INclusive Large Multimodal Model Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43321dff-4be8-48d7-8055-8195a6e15781 · inbound
MRAMG-Bench: A Comprehensive Benchmark for Advancing Multimodal Retrieval-Augmented Multimodal Generation Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2193c28a-8434-4b28-9561-5bd00eefe9ad · inbound
WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d3d61bf3-5cef-4116-b793-1bcec186b665 · inbound
Ola: Pushing the Frontiers of Omni-Modal Language Model Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e497a2f-7f7b-4020-a666-7df7f524aa32 · inbound
SARChat-Bench-2M: A Multi-Task Vision-Language Benchmark for SAR Image Interpretation Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0a61174-98ae-4d45-a0c6-b7deacda45be · inbound
EmoAssist: Emotional Assistant for Visual Impairment Community Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be40dd94-a9cd-4805-a809-abc0d8a832aa · inbound
Qwen2.5-VL Technical Report Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3a059973-2a2b-4636-9b9f-9ab0c3887ebe · inbound
Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 4cdad193-8e50-4b67-9043-fa2598f9f8e8 · inbound
AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 27815057-470f-436b-8557-88f2dbea35fa · inbound
AlphaDrive: Unleashing the Power of VLMs in Autonomous Driving via Reinforcement Learning and Reasoning Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 4cf32f77-ca1b-417f-aa46-6735afb14994 · inbound
R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 454d47c7-8dd6-4678-8e23-d5621c470468 · inbound
Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 272
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 715acb0e-a21b-4696-a749-b7d8769192c9 · inbound
Cosmos-Reason1: From Physical Common Sense To Embodied Reasoning Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f62fed40-d4c2-4ab0-b2b4-19a60f7d09ad · inbound
MathFlow: Enhancing the Perceptual Flow of MLLMs for Visual Mathematical Problems Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d00c3be7-a55c-4580-b867-4b22d5693751 · inbound
OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 413e1df8-f063-46eb-9398-5b712199bbec · inbound
SpaceR: Reinforcing MLLMs in Video Spatial Reasoning Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f97b9fff-840d-4c74-a7a3-fb38cd2134f6 · inbound
VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation dca35e67-146b-42ea-926a-304bbc25abee · inbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2ab87385-289b-4fe2-a7e0-3b51be7dd4cc · inbound
Consensus Entropy: Harnessing Multi-VLM Agreement for Self-Verifying and Self-Improving OCR Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 57bd24ad-6817-4225-ace8-5877e4367350 · inbound
Perception Encoder: The best visual embeddings are not at the output of the network Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ec1683fa-ebdf-461b-bfa0-47116c0ab52a · inbound
We'll Fix it in Post: Improving Text-to-Video Generation with Neuro-Symbolic Feedback Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 48eafba0-689d-44b6-a352-d9810ad2d040 · inbound
CAFES: A Collaborative Multi-Agent Framework for Multi-Granular Multimodal Essay Scoring Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fcbff939-49af-4d50-8fd5-c9dad6aed463 · inbound
ShieldVLM: Safeguarding the Multimodal Implicit Toxicity via Deliberative Reasoning with LVLMs Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1bb37353-e8cb-4ac9-86bc-d6b8d4ebf0d1 · inbound
Visual Agentic Reinforcement Fine-Tuning Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bce2c6ed-8e9d-463c-b9df-bb33676dc6c4 · inbound
DeepEyes: Incentivizing "Thinking with Images" via Reinforcement Learning Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b9eae6d5-b575-4723-9821-2e20512be1e8 · inbound
Towards a Foundation Model for Communication Systems Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d0872ad-c2e1-40e5-a047-d42e50b38f36 · inbound
VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79f87b93-d215-452b-a229-07b36248cee7 · inbound
Emerging Properties in Unified Multimodal Pretraining Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation aeb95aec-7c5b-48a8-9143-5c9c25861d5b · inbound
Clapper: Compact Learning and Video Representation in VLMs Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1d57b38-53f5-49c4-a16d-20c220232ee3 · inbound
STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 294efcf3-6fe5-4ed8-bcf1-da716dec89ff · inbound
Training-Free Reasoning and Reflection in MLLMs Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cfd80a0d-2535-4e0e-8341-55ac8772c364 · inbound
QuickVideo: Real-Time Long Video Understanding with System Algorithm Co-Design Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f10b55dc-d9f9-484d-b772-14fa6ab0c3ed · inbound
Understanding Generative AI Capabilities in Everyday Image Editing Tasks Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a075c7a6-6f15-4551-b397-40d6859966ba · inbound
NTIRE 2025 challenge on Text to Image Generation Model Quality Assessment Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 050d2c3b-ca16-4e0d-857b-5540d678d2de · inbound
Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 724ccaf1-5152-4020-b225-fbaa6e536ca7 · inbound
Bridging the Dynamic Perception Gap: Training-Free Draft Chain-of-Thought for Dynamic Multimodal Spatial Reasoning Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13a8bef6-d4c6-4bb7-9e4f-450488504528 · inbound
DetailMaster: Can Your Text-to-Image Model Handle Long Prompts? Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df65dd4a-7c41-4556-91b1-3026e5e6bdad · inbound
Backdoor Cleaning without External Guidance in MLLM Fine-tuning Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc7986d5-c5af-4279-bd62-9dae15108a73 · inbound
SpatialScore: Towards Comprehensive Evaluation for Spatial Intelligence Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 4aa1239c-a36c-4ac4-8feb-489ff15f1a8f · inbound
ARB: A Comprehensive Arabic Multimodal Reasoning Benchmark Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 917cbe05-127a-47dd-8b36-ba927934ca76 · inbound
FullFront: Benchmarking MLLMs Across the Full Front-End Engineering Workflow Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 822023ef-72e4-4818-a6fb-0a8c6d76f737 · inbound
VIBE: Annotation-Free Video-to-Text Information Bottleneck Evaluation for TL;DR Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83cd6933-ec06-4df0-b1eb-f65ae61f62c2 · inbound
$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84836bc0-651d-46b3-8fbf-c737d4b9694f · inbound
MLLMs are Deeply Affected by Modality Bias Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33c63cf9-3e2e-4215-ada8-85d55d388162 · inbound
VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad2e589f-6bcc-4b8b-8565-678b05f362ab · inbound
CardioCoT: Hierarchical Reasoning for Multimodal Survival Analysis Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4f7ae01-793e-4ea3-90f5-516f7ba1e176 · inbound
Improving Medical Reasoning with Curriculum-Aware Reinforcement Learning Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f295fbe9-90b6-48b8-ad51-5d90ac34222a · inbound
Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c97a7b7-3b6a-40a8-959d-3a065bd0e19d · inbound
Benchmarking Multimodal Knowledge Conflict for Large Multimodal Models Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2699f0e0-d716-4443-96fa-cfcfafcf8b6a · inbound
Align and Surpass Human Camouflaged Perception: Visual Refocus Reinforcement Fine-Tuning Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 037a3a34-ad60-4583-98aa-61a6086261d6 · inbound
VisCRA: A Visual Chain Reasoning Attack for Jailbreaking Multimodal Large Language Models Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2bf1845f-f9b1-4086-930c-d38c6ee018da · inbound
MT$^{3}$: Scaling MLLM-based Text Image Machine Translation via Multi-Task Reinforcement Learning Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc8321d9-4ebe-4ec6-9b16-5f6864a728b2 · inbound
AdaTP: Attention-Debiased Token Pruning for Video Large Language Models Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 373e72bc-1f8a-40de-bfaa-f65fa07922f0 · inbound
Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bbb13462-b46a-4cf4-9581-2759a67ac555 · inbound
Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning? Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0c8d5d8f-a9db-45a7-b901-2184cca4158c · inbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 794782d4-e9a5-40f2-8394-4d03513b5c12 · inbound
ID-Align: RoPE-Conscious Position Remapping for Dynamic High-Resolution Adaptation in Vision-Language Models Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9790918-9649-4cc5-ab71-331c48752194 · inbound
VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab1fe195-2f7d-4b5c-bb0e-a44cd8a1818e · inbound
Reinforced Reasoning for Embodied Planning Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e203eb4-52f9-4748-a37d-59c946346027 · inbound
SridBench: Benchmark of Scientific Research Illustration Drawing of Image Generation Model Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 616cbabc-ee87-4b4a-81fc-493b07965263 · inbound
Zooming from Context to Cue: Hierarchical Preference Optimization for Multi-Image MLLMs Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 233a67b4-7300-4094-bd33-e442cae16063 · inbound
Scaling-up Perceptual Video Quality Assessment Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c51cfa79-1d7a-46af-bebb-b92b42132eb1 · inbound
Universal Visuo-Tactile Video Understanding for Embodied Interaction Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50f5e6e4-7b4e-42e3-81e1-9c05c4c47877 · inbound
Synthetic Document Question Answering in Hungarian Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1633aeeb-f418-4f81-ad87-d005d7b4126e · inbound
Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 064d151c-f6d6-48f8-abe0-6684ee452d39 · inbound
Infi-MMR: Curriculum-based Unlocking Multimodal Reasoning via Phased Reinforcement Learning in Multimodal Small Language Models Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25d24c42-a878-4a4f-a6eb-056edf6a179e · inbound
ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82dd815f-cc4c-4aac-a85e-a3dd84af6992 · inbound
Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24f747fc-509d-416c-80e3-8e1d9704fa43 · inbound
Seeing is Not Reasoning: MVPBench for Graph-based Evaluation of Multi-path Visual Physical CoT Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 538b5115-2b18-4501-8090-17fa6c5d4c27 · inbound
DisTime: Distribution-based Time Representation for Video Large Language Models Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87aac395-fd51-415d-992b-53fc65b37f08 · inbound
Unifying Language Agent Algorithms with Graph-based Orchestration Engine for Reproducible Agent Research Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c113cf7c-5adc-4af2-af5c-8140a0ea177c · inbound
MMAFFBen: A Multilingual and Multimodal Affective Analysis Benchmark for Evaluating LLMs and VLMs Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 082977d1-44db-463c-a5ba-4a4e62f6ed8b · inbound
SORCE: Small Object Retrieval in Complex Environments Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation faf511ce-2e13-437b-8378-3041049502fe · inbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45b1c30d-60b4-4e2d-9249-f678a1b45626 · inbound
EffiVLM-BENCH: A Comprehensive Benchmark for Evaluating Training-Free Acceleration in Large Vision-Language Models Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b01b931d-9111-4054-a48e-48d08084632b · inbound
MedBookVQA: A Systematic and Comprehensive Medical Benchmark Derived from Open-Access Book Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 765a8170-44c3-4b5f-8f69-cb30a2eebca8 · inbound
FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1723077-1fe6-4428-b60f-2d9391758425 · inbound
NavBench: Probing Multimodal Large Language Models for Embodied Navigation Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e6fe2f7-205c-4d23-bd7a-9b9e7cd2f583 · inbound
Abstractive Visual Understanding of Multi-modal Structured Knowledge: A New Perspective for MLLM Evaluation Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b65b9093-4a76-49dc-8d33-1b46a11cc01e · inbound
ReAgent-V: A Reward-Driven Multi-Agent Framework for Video Understanding Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.