{"work":{"id":"e22c3789-9e71-4242-b6ea-3e60e06e2b66","openalex_id":"https://openalex.org/W4413144973","doi":"10.1109/cvpr52734.2025.01245","arxiv_id":"2310.02255","raw_key":null,"title":"MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts","authors":null,"authors_text":"Pan Lu, Hritik Bansal, Tony Xia, Jiacheng Liu, Chunyuan Li, Hannaneh Hajishirzi","year":2023,"venue":"cs.CV","abstract":"Large Language Models (LLMs) and Large Multimodal Models (LMMs) exhibit impressive problem-solving skills in many tasks and domains, but their ability in mathematical reasoning in visual contexts has not been systematically studied. To bridge this gap, we present MathVista, a benchmark designed to combine challenges from diverse mathematical and visual tasks. It consists of 6,141 examples, derived from 28 existing multimodal datasets involving mathematics and 3 newly created datasets (i.e., IQTest, FunctionQA, and PaperQA). Completing these tasks requires fine-grained, deep visual understanding and compositional reasoning, which all state-of-the-art foundation models find challenging. With MathVista, we have conducted a comprehensive, quantitative evaluation of 12 prominent foundation models. The best-performing GPT-4V model achieves an overall accuracy of 49.9%, substantially outperforming Bard, the second-best performer, by 15.1%. Our in-depth analysis reveals that the superiority of GPT-4V is mainly attributed to its enhanced visual perception and mathematical reasoning. However, GPT-4V still falls short of human performance by 10.4%, as it often struggles to understand complex figures and perform rigorous reasoning. This significant gap underscores the critical role that MathVista will play in the development of general-purpose AI agents capable of tackling mathematically intensive and visually rich real-world tasks. We further explore the new ability of self-verification, the application of self-consistency, and the interactive chatbot capabilities of GPT-4V, highlighting its promising potential for future research. The project is available at https://mathvista.github.io/.","external_url":"https://arxiv.org/abs/2310.02255","cited_by_count":6,"metadata_source":"pith","metadata_fetched_at":"2026-08-05T02:28:24.338817+00:00","pith_arxiv_id":"2310.02255","created_at":"2026-05-09T06:55:43.914826+00:00","updated_at":"2026-08-05T02:28:24.338817+00:00","title_quality_ok":true,"display_title":"MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts","render_title":"MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts"},"hub":{"state":{"work_id":"e22c3789-9e71-4242-b6ea-3e60e06e2b66","tier":"super_hub","tier_reason":"100+ Pith inbound or 10,000+ external citations","pith_inbound_count":182,"external_cited_by_count":6,"distinct_field_count":11,"first_pith_cited_at":"2023-05-05T17:59:46+00:00","last_pith_cited_at":"2026-07-09T17:58:29+00:00","author_build_status":"needed","summary_status":"needed","contexts_status":"needed","graph_status":"needed","ask_index_status":"needed","reader_status":"not_needed","recognition_status":"not_needed","updated_at":"2026-08-21T20:39:26.183141+00:00","tier_text":"super_hub"},"tier":"super_hub","role_counts":[{"context_role":"dataset","n":26},{"context_role":"background","n":24},{"context_role":"baseline","n":4}],"polarity_counts":[{"context_polarity":"use_dataset","n":26},{"context_polarity":"background","n":23},{"context_polarity":"baseline","n":4},{"context_polarity":"unclear","n":1}],"runs":{"ask_index":{"job_type":"ask_index","status":"succeeded","result":{"title":"MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts","claims":[{"claim_text":"Large Language Models (LLMs) and Large Multimodal Models (LMMs) exhibit impressive problem-solving skills in many tasks and domains, but their ability in mathematical reasoning in visual contexts has not been systematically studied. To bridge this gap, we present MathVista, a benchmark designed to combine challenges from diverse mathematical and visual tasks. It consists of 6,141 examples, derived from 28 existing multimodal datasets involving mathematics and 3 newly created datasets (i.e., IQTest, FunctionQA, and PaperQA). Completing these tasks requires fine-grained, deep visual understandin","claim_type":"abstract","evidence_strength":"source_metadata"},{"claim_text":"(iii)Correctness-conditionedregularization restricts invariance enforcement to successful trajectories, so the model is not pushed toward becoming consistently but systematically incorrect. We validate ROMA by fine-tuning Qwen3-VL 4B and 8B Instruct models [1] and evaluating visual robustness across seven multimodal reasoning benchmarks: MathVista [18], WeMath [25], ChartQA [21], LogicVista [37], MMStar [4], VisualPuzzles [31], and RealWorldQA [36]. While standard GRPO reaches strong clean-input","claim_type":"dataset","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"imately 25,000 H20 GPU hours. Note on Reproducibility: This work is primarily aimed at MLLM pre-training teams. The computational cost is highly dependent on the scale of proprietary training data and the size of the model. 4.2. Evaluation & Benchmarks Fine-grained Image BenchmarksWe follow Opencom- pass [14] image leaderboard, evaluating on MMBench [50], MathVista [57], HallusionBench [25], OCRBench [52], AI2D [35], MMVet [108], MMStar [9], MMMU [109]. Video BenchmarksWe evaluate on a wide rang","claim_type":"dataset","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"OCR-related benchmarks include: DocVQA test [82], ChartQA test [81], InfographicVQA test [83], TextVQA val [100], and OCRBench [67]. General multimodal benchmarks encompass: MME [26], RealWorldQA [125], AI2D test [39], MMMU val [135], MMBench-EN/CN test [66], CCBench dev [66], MMVet [133], SEED Image [46], and HallusionBench [30]. Additionally, the math dataset includes MathVista testmini [75]. * denotes that Rosetta OCR tokens are used in the testing of TextVQA. The MME results we report are th","claim_type":"dataset","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"most reasoning-intensive subset, as each sample includes an explicit chain-of-thought trace. 15 Benchmark SenseNova-U1 8B-Think Qwen3VL 8B-Think Qwen3.5 9B SenseNova-U1 30BA3B-Think Qwen3VL 30BA3B-Think Qwen3.5 35BA3B Gemma4 26BA4B LongCat-Next 68BA3B STEM & Reasoning MMMU [159] 74.78 74.10 78.40 80.55 76.00 81.40 76.56 70.60 MMMU-Pro [160] 67.69 60.40 70.10 72.83 63.00 75.10 73.80 60.30 MathVistamini [87] 84.20 81.40 85.70 85.30 81.90 86.20 72.70 83.10 MathVision [132] 75.82 62.70 78.90 79.63 6","claim_type":"dataset","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"3 56.0 48.8 49.9 54.1 48.8 BLINKtest - 54.8 48.2 47.0∗ 43.1∗ 56.7 Knowledge/General QA RealWorldQA 70.7 70.1 66.3 68.6 70.1 72.7 AI2D 93.2 84.5 81.4 92.3 83.0 84.7 GQA - - 62.3 - 62.4∗ 64.9 MME - 2344 1998 2219 2327 2102 Mathematical Reasoning. VideoLLaMA3's mathematical reasoning capabilities are evaluated through the MathVista [115] and MathVision [116] benchmarks. These benchmarks focus on evaluating the model's ability to reason about and solve mathematical problems presented in visual forma","claim_type":"dataset","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"by rules or external executors, which will later be incorporated into the RLVR training. 4.3.1 Visual STEM STEM (science, technology, engineering, and mathematics) questions usually have unique and verifiable answers, which are suitable for RLVR. We collect over one million problems with images in STEM fields, mostly on mathematics, from both open-sourced resources [85] and internal K-12 education collections. To prepare the training data, multiple-choice questions were initially transformed int","claim_type":"dataset","confidence":0.95,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts because it crossed a citation-hub threshold. Current citing contexts most often use it as dataset evidence (21 contexts).","role_counts":[{"n":21,"context_role":"dataset"},{"n":10,"context_role":"background"},{"n":1,"context_role":"baseline"}]},"error":null,"updated_at":"2026-05-17T00:49:23.859519+00:00"},"author_expand":{"job_type":"author_expand","status":"succeeded","result":{"authors_linked":[{"id":"c15e46ea-b407-41e3-86bd-d2d90fe983f5","orcid":null,"display_name":"Pan Lu"},{"id":"811ef97e-e496-4c6f-878a-1ba4cf7c8a05","orcid":null,"display_name":"Hritik Bansal"},{"id":"1864c6c2-ae48-4398-8f17-bbdcf484dfe5","orcid":null,"display_name":"Tony Xia"},{"id":"a219b7f3-28b3-43a9-ae1b-81ca07174106","orcid":null,"display_name":"Jiacheng Liu"},{"id":"4c582d1e-c6c5-47f7-b526-e877960cf613","orcid":null,"display_name":"Chunyuan Li"},{"id":"9628ae2e-a10d-4388-809d-a00630df0f88","orcid":null,"display_name":"Hannaneh Hajishirzi"}]},"error":null,"updated_at":"2026-05-17T00:49:23.849341+00:00"},"context_extract":{"job_type":"context_extract","status":"succeeded","result":{"enqueued_papers":25},"error":null,"updated_at":"2026-05-14T09:38:51.491933+00:00"},"graph_features":{"job_type":"graph_features","status":"succeeded","result":{"co_cited":[{"title":"Qwen2.5-VL Technical Report","work_id":"69dffacb-bfe8-442d-be86-48624c60426f","shared_citers":26},{"title":"Qwen3-VL Technical Report","work_id":"1fe243aa-e3c0-4da6-b391-4cbcfc88d5c0","shared_citers":21},{"title":"LLaVA-OneVision: Easy Visual Task Transfer","work_id":"f5f2452b-f2a9-49ac-b38d-c76e18cdfe49","shared_citers":20},{"title":"Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution","work_id":"8abcfe4f-e0fb-44b7-9123-448fac95f90a","shared_citers":20},{"title":"DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models","work_id":"c5006563-f3ec-438a-9e35-b7b484f34828","shared_citers":19},{"title":"Gemini: A Family of Highly Capable Multimodal Models","work_id":"83f7c85b-3f11-450f-ac0c-64d9745220b2","shared_citers":19},{"title":"DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning","work_id":"e6b75ad5-2877-4168-97c8-710407094d20","shared_citers":18},{"title":"Training Verifiers to Solve Math Word Problems","work_id":"acab1aa8-b4d6-40e0-a3ee-25341701dca2","shared_citers":15},{"title":"GPT-4 Technical Report","work_id":"b928e041-6991-4c08-8c81-0359e4097c7b","shared_citers":14},{"title":"MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities","work_id":"7f3bac41-a0a5-4a7a-bfd2-526b616db745","shared_citers":14},{"title":"Proximal Policy Optimization Algorithms","work_id":"240c67fe-d14d-4520-91c1-38a4e272ca19","shared_citers":14},{"title":"Are We on the Right Way for Evaluating Large Vision-Language Models?","work_id":"0d0b977c-a42e-49b1-869e-b7360dca5282","shared_citers":12},{"title":"InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models","work_id":"fe8637aa-12bc-4434-8d36-9f57b5eebcbe","shared_citers":12},{"title":"Measuring multimodal mathematical reasoning with MATH-Vision dataset","work_id":"c59c0707-68e6-4ab4-9b9f-293004398dc7","shared_citers":12},{"title":"OpenAI o1 System Card","work_id":"68d3c334-0fc9-49e3-b7b0-a69afae933e2","shared_citers":12},{"title":"Qwen3 Technical Report","work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e","shared_citers":12},{"title":"Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context","work_id":"80e3e977-f1bb-4c83-8d0c-1ab0a0c5c3f1","shared_citers":11},{"title":"MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models","work_id":"a7e3a737-e007-42bc-be89-c4d34c5ee071","shared_citers":11},{"title":"MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models","work_id":"806d2e73-71b3-4d56-87e0-39d571cc15d6","shared_citers":11},{"title":"The Llama 3 Herd of Models","work_id":"1549a635-88af-4ac1-acfe-51ae7bb53345","shared_citers":11},{"title":"Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models","work_id":"38998646-34ee-4605-b661-ab356f16d6e5","shared_citers":11},{"title":"DAPO: An Open-Source LLM Reinforcement Learning System at Scale","work_id":"64019d00-0b11-4bbd-b173-b46c8fad0157","shared_citers":10},{"title":"Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling","work_id":"ee70bdc8-4656-4849-ada7-ce42a2278d70","shared_citers":10},{"title":"InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency","work_id":"b8f5e260-fff5-444e-bcf5-2c42cfefd83d","shared_citers":10}],"time_series":[{"n":7,"year":2024},{"n":10,"year":2025},{"n":50,"year":2026}],"dependency_candidates":[]},"error":null,"updated_at":"2026-05-14T09:38:35.526440+00:00"},"identity_refresh":{"job_type":"identity_refresh","status":"succeeded","result":{"items":[{"title":"Qwen3 Technical Report","outcome":"unchanged","work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e","resolver":"local_arxiv","confidence":0.98,"old_work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e"}],"counts":{"fixed":0,"merged":0,"unchanged":1,"quarantined":0,"needs_external_resolution":0},"errors":[],"attempted":1},"error":null,"updated_at":"2026-05-14T09:38:55.526014+00:00"},"role_polarity":{"job_type":"role_polarity","status":"succeeded","result":{"title":"MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts","claims":[{"claim_text":"Large Language Models (LLMs) and Large Multimodal Models (LMMs) exhibit impressive problem-solving skills in many tasks and domains, but their ability in mathematical reasoning in visual contexts has not been systematically studied. To bridge this gap, we present MathVista, a benchmark designed to combine challenges from diverse mathematical and visual tasks. It consists of 6,141 examples, derived from 28 existing multimodal datasets involving mathematics and 3 newly created datasets (i.e., IQTest, FunctionQA, and PaperQA). Completing these tasks requires fine-grained, deep visual understandin","claim_type":"abstract","evidence_strength":"source_metadata"},{"claim_text":"(iii)Correctness-conditionedregularization restricts invariance enforcement to successful trajectories, so the model is not pushed toward becoming consistently but systematically incorrect. We validate ROMA by fine-tuning Qwen3-VL 4B and 8B Instruct models [1] and evaluating visual robustness across seven multimodal reasoning benchmarks: MathVista [18], WeMath [25], ChartQA [21], LogicVista [37], MMStar [4], VisualPuzzles [31], and RealWorldQA [36]. While standard GRPO reaches strong clean-input","claim_type":"dataset","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"imately 25,000 H20 GPU hours. Note on Reproducibility: This work is primarily aimed at MLLM pre-training teams. The computational cost is highly dependent on the scale of proprietary training data and the size of the model. 4.2. Evaluation & Benchmarks Fine-grained Image BenchmarksWe follow Opencom- pass [14] image leaderboard, evaluating on MMBench [50], MathVista [57], HallusionBench [25], OCRBench [52], AI2D [35], MMVet [108], MMStar [9], MMMU [109]. Video BenchmarksWe evaluate on a wide rang","claim_type":"dataset","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"OCR-related benchmarks include: DocVQA test [82], ChartQA test [81], InfographicVQA test [83], TextVQA val [100], and OCRBench [67]. General multimodal benchmarks encompass: MME [26], RealWorldQA [125], AI2D test [39], MMMU val [135], MMBench-EN/CN test [66], CCBench dev [66], MMVet [133], SEED Image [46], and HallusionBench [30]. Additionally, the math dataset includes MathVista testmini [75]. * denotes that Rosetta OCR tokens are used in the testing of TextVQA. The MME results we report are th","claim_type":"dataset","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"most reasoning-intensive subset, as each sample includes an explicit chain-of-thought trace. 15 Benchmark SenseNova-U1 8B-Think Qwen3VL 8B-Think Qwen3.5 9B SenseNova-U1 30BA3B-Think Qwen3VL 30BA3B-Think Qwen3.5 35BA3B Gemma4 26BA4B LongCat-Next 68BA3B STEM & Reasoning MMMU [159] 74.78 74.10 78.40 80.55 76.00 81.40 76.56 70.60 MMMU-Pro [160] 67.69 60.40 70.10 72.83 63.00 75.10 73.80 60.30 MathVistamini [87] 84.20 81.40 85.70 85.30 81.90 86.20 72.70 83.10 MathVision [132] 75.82 62.70 78.90 79.63 6","claim_type":"dataset","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"3 56.0 48.8 49.9 54.1 48.8 BLINKtest - 54.8 48.2 47.0∗ 43.1∗ 56.7 Knowledge/General QA RealWorldQA 70.7 70.1 66.3 68.6 70.1 72.7 AI2D 93.2 84.5 81.4 92.3 83.0 84.7 GQA - - 62.3 - 62.4∗ 64.9 MME - 2344 1998 2219 2327 2102 Mathematical Reasoning. VideoLLaMA3's mathematical reasoning capabilities are evaluated through the MathVista [115] and MathVision [116] benchmarks. These benchmarks focus on evaluating the model's ability to reason about and solve mathematical problems presented in visual forma","claim_type":"dataset","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"by rules or external executors, which will later be incorporated into the RLVR training. 4.3.1 Visual STEM STEM (science, technology, engineering, and mathematics) questions usually have unique and verifiable answers, which are suitable for RLVR. We collect over one million problems with images in STEM fields, mostly on mathematics, from both open-sourced resources [85] and internal K-12 education collections. To prepare the training data, multiple-choice questions were initially transformed int","claim_type":"dataset","confidence":0.95,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts because it crossed a citation-hub threshold. Current citing contexts most often use it as dataset evidence (21 contexts).","role_counts":[{"n":21,"context_role":"dataset"},{"n":10,"context_role":"background"},{"n":1,"context_role":"baseline"}]},"error":null,"updated_at":"2026-05-17T00:49:23.855292+00:00"},"summary_claims":{"job_type":"summary_claims","status":"succeeded","result":{"title":"MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts","claims":[{"claim_text":"Large Language Models (LLMs) and Large Multimodal Models (LMMs) exhibit impressive problem-solving skills in many tasks and domains, but their ability in mathematical reasoning in visual contexts has not been systematically studied. To bridge this gap, we present MathVista, a benchmark designed to combine challenges from diverse mathematical and visual tasks. It consists of 6,141 examples, derived from 28 existing multimodal datasets involving mathematics and 3 newly created datasets (i.e., IQTest, FunctionQA, and PaperQA). Completing these tasks requires fine-grained, deep visual understandin","claim_type":"abstract","evidence_strength":"source_metadata"}],"why_cited":"Pith tracks MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts because it crossed a citation-hub threshold.","role_counts":[]},"error":null,"updated_at":"2026-05-14T09:38:55.531830+00:00"}},"summary":{"title":"MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts","claims":[{"claim_text":"Large Language Models (LLMs) and Large Multimodal Models (LMMs) exhibit impressive problem-solving skills in many tasks and domains, but their ability in mathematical reasoning in visual contexts has not been systematically studied. To bridge this gap, we present MathVista, a benchmark designed to combine challenges from diverse mathematical and visual tasks. It consists of 6,141 examples, derived from 28 existing multimodal datasets involving mathematics and 3 newly created datasets (i.e., IQTest, FunctionQA, and PaperQA). Completing these tasks requires fine-grained, deep visual understandin","claim_type":"abstract","evidence_strength":"source_metadata"}],"why_cited":"Pith tracks MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts because it crossed a citation-hub threshold.","role_counts":[]},"graph":{"co_cited":[{"title":"Qwen2.5-VL Technical Report","work_id":"69dffacb-bfe8-442d-be86-48624c60426f","shared_citers":26},{"title":"Qwen3-VL Technical Report","work_id":"1fe243aa-e3c0-4da6-b391-4cbcfc88d5c0","shared_citers":21},{"title":"LLaVA-OneVision: Easy Visual Task Transfer","work_id":"f5f2452b-f2a9-49ac-b38d-c76e18cdfe49","shared_citers":20},{"title":"Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution","work_id":"8abcfe4f-e0fb-44b7-9123-448fac95f90a","shared_citers":20},{"title":"DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models","work_id":"c5006563-f3ec-438a-9e35-b7b484f34828","shared_citers":19},{"title":"Gemini: A Family of Highly Capable Multimodal Models","work_id":"83f7c85b-3f11-450f-ac0c-64d9745220b2","shared_citers":19},{"title":"DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning","work_id":"e6b75ad5-2877-4168-97c8-710407094d20","shared_citers":18},{"title":"Training Verifiers to Solve Math Word Problems","work_id":"acab1aa8-b4d6-40e0-a3ee-25341701dca2","shared_citers":15},{"title":"GPT-4 Technical Report","work_id":"b928e041-6991-4c08-8c81-0359e4097c7b","shared_citers":14},{"title":"MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities","work_id":"7f3bac41-a0a5-4a7a-bfd2-526b616db745","shared_citers":14},{"title":"Proximal Policy Optimization Algorithms","work_id":"240c67fe-d14d-4520-91c1-38a4e272ca19","shared_citers":14},{"title":"Are We on the Right Way for Evaluating Large Vision-Language Models?","work_id":"0d0b977c-a42e-49b1-869e-b7360dca5282","shared_citers":12},{"title":"InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models","work_id":"fe8637aa-12bc-4434-8d36-9f57b5eebcbe","shared_citers":12},{"title":"Measuring multimodal mathematical reasoning with MATH-Vision dataset","work_id":"c59c0707-68e6-4ab4-9b9f-293004398dc7","shared_citers":12},{"title":"OpenAI o1 System Card","work_id":"68d3c334-0fc9-49e3-b7b0-a69afae933e2","shared_citers":12},{"title":"Qwen3 Technical Report","work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e","shared_citers":12},{"title":"Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context","work_id":"80e3e977-f1bb-4c83-8d0c-1ab0a0c5c3f1","shared_citers":11},{"title":"MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models","work_id":"a7e3a737-e007-42bc-be89-c4d34c5ee071","shared_citers":11},{"title":"MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models","work_id":"806d2e73-71b3-4d56-87e0-39d571cc15d6","shared_citers":11},{"title":"The Llama 3 Herd of Models","work_id":"1549a635-88af-4ac1-acfe-51ae7bb53345","shared_citers":11},{"title":"Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models","work_id":"38998646-34ee-4605-b661-ab356f16d6e5","shared_citers":11},{"title":"DAPO: An Open-Source LLM Reinforcement Learning System at Scale","work_id":"64019d00-0b11-4bbd-b173-b46c8fad0157","shared_citers":10},{"title":"Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling","work_id":"ee70bdc8-4656-4849-ada7-ce42a2278d70","shared_citers":10},{"title":"InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency","work_id":"b8f5e260-fff5-444e-bcf5-2c42cfefd83d","shared_citers":10}],"time_series":[{"n":7,"year":2024},{"n":10,"year":2025},{"n":50,"year":2026}],"dependency_candidates":[]},"authors":[{"id":"4c582d1e-c6c5-47f7-b526-e877960cf613","orcid":null,"display_name":"Chunyuan Li","source":"manual","import_confidence":0.72},{"id":"9628ae2e-a10d-4388-809d-a00630df0f88","orcid":null,"display_name":"Hannaneh Hajishirzi","source":"manual","import_confidence":0.72},{"id":"811ef97e-e496-4c6f-878a-1ba4cf7c8a05","orcid":null,"display_name":"Hritik Bansal","source":"manual","import_confidence":0.72},{"id":"a219b7f3-28b3-43a9-ae1b-81ca07174106","orcid":null,"display_name":"Jiacheng Liu","source":"manual","import_confidence":0.72},{"id":"c15e46ea-b407-41e3-86bd-d2d90fe983f5","orcid":null,"display_name":"Pan Lu","source":"manual","import_confidence":0.72},{"id":"1864c6c2-ae48-4398-8f17-bbdcf484dfe5","orcid":null,"display_name":"Tony Xia","source":"manual","import_confidence":0.72}]}}