{"work":{"id":"5f6cf57b-2407-4127-b39c-d8a61494e474","openalex_id":"https://openalex.org/W4417298859","doi":"10.48550/arxiv.2505.14362","arxiv_id":"2505.14362","raw_key":null,"title":"DeepEyes: Incentivizing \"Thinking with Images\" via Reinforcement Learning","authors":null,"authors_text":"Ziwei Zheng, Michael Yang, Jack Hong, Chenxiao Zhao, Guohai Xu, Le Yang","year":2025,"venue":"cs.CV","abstract":"Large Vision-Language Models excel at multimodal understanding but struggle to deeply integrate visual information into their predominantly text-based reasoning processes, a key challenge in mirroring human cognition. To address this, we introduce DeepEyes, a model that learns to \"think with images\", trained end-to-end with reinforcement learning without requiring pre-collected reasoning data for cold-start supervised fine-tuning (SFT). Notably, this ability emerges natively, leveraging the model's own grounding capability as an intrinsic function rather than relying on external specialized models or APIs. We enable this capability through active perception, where the model learns to strategically ground its reasoning in visual information, guided by a tailored data selection and reward strategy. DeepEyes achieves significant performance gains on general perception and reasoning benchmarks and also demonstrates improvement in grounding, hallucination, and mathematical reasoning tasks. Interestingly, we observe the distinct evolution of active perception from initial exploration to efficient and accurate exploitation, and diverse thinking patterns that closely mirror human visual reasoning processes. Code is available at https://github.com/Visual-Agent/DeepEyes.","external_url":"https://arxiv.org/abs/2505.14362","cited_by_count":0,"metadata_source":"pith","metadata_fetched_at":"2026-08-05T02:28:24.338817+00:00","pith_arxiv_id":"2505.14362","created_at":"2026-05-08T17:08:34.333954+00:00","updated_at":"2026-08-05T02:28:24.338817+00:00","title_quality_ok":true,"display_title":"DeepEyes: Incentivizing \"Thinking with Images\" via Reinforcement Learning","render_title":"DeepEyes: Incentivizing \"Thinking with Images\" via Reinforcement Learning"},"hub":{"state":{"work_id":"5f6cf57b-2407-4127-b39c-d8a61494e474","tier":"super_hub","tier_reason":"100+ Pith inbound or 10,000+ external citations","pith_inbound_count":118,"external_cited_by_count":0,"distinct_field_count":6,"first_pith_cited_at":"2025-02-06T18:59:40+00:00","last_pith_cited_at":"2026-07-09T17:58:29+00:00","author_build_status":"needed","summary_status":"needed","contexts_status":"needed","graph_status":"needed","ask_index_status":"needed","reader_status":"not_needed","recognition_status":"not_needed","updated_at":"2026-08-23T19:59:26.754530+00:00","tier_text":"super_hub"},"tier":"super_hub","role_counts":[{"context_role":"background","n":21},{"context_role":"baseline","n":6},{"context_role":"dataset","n":2}],"polarity_counts":[{"context_polarity":"background","n":21},{"context_polarity":"baseline","n":6},{"context_polarity":"use_dataset","n":2}],"runs":{"ask_index":{"job_type":"ask_index","status":"succeeded","result":{"title":"DeepEyes: Incentivizing \"Thinking with Images\" via Reinforcement Learning","claims":[{"claim_text":"Large Vision-Language Models excel at multimodal understanding but struggle to deeply integrate visual information into their predominantly text-based reasoning processes, a key challenge in mirroring human cognition. To address this, we introduce DeepEyes, a model that learns to \"think with images\", trained end-to-end with reinforcement learning without requiring pre-collected reasoning data for cold-start supervised fine-tuning (SFT). Notably, this ability emerges natively, leveraging the model's own grounding capability as an intrinsic function rather than relying on external specialized mo","claim_type":"abstract","evidence_strength":"source_metadata"},{"claim_text":"Geox-bench: Benchmarking cross-view geo-localization and pose estimation capabilities of large multimodal models.arXiv preprint arXiv:2511.13259, 2025. [104] Ziwei Zheng, Michael Yang, Jack Hong, Chenxiao Zhao, Guohai Xu, Le Yang, Chao Shen, and Xing Yu. Deepeyes: Incentivizing\" thinking with images\" via reinforcement learning. arXiv preprint arXiv:2505.14362, 2025. [105] Xiran Zhou, Yi Wen, Zhenfeng Shao, Wenwen Li, Kaiyuan Li, Honghao Li, Xiao Xie, and Zhigang Yan. Cartomark: a benchmark datas","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"ReSearch [108] External Qwen2.5-7B/32B-Instruct /githubGitHub StepSearch [267] External Qwen2.5-3B/7B-Base/Instruct /githubGitHub DeepResearcher [268] External Qwen2.5-7B-Instruct /githubGitHub WebDancer [106] External Qwen2.5-7B/32B, QWQ-32B /githubGitHub WebThinker [269] External QwQ-32B, DeepSeek-R1-Distilled-Qwen, Qwen2.5-32B/githubGitHub WebSailor [105] External Qwen2.5-3B/7B/32B/72B /githubGitHub WebWatcher [270] External Qwen2.5-VL-7B/32B /githubGitHub WebShaper [271] External Qwen-2.5-32","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"3 57.043.9 12.4 29.649.1 14.135.7InfiGUI-G1-3B [21]64.9 20.0 -51.5 16.8 -50.8 25.0 -68.8 32.7 -70.6 32.1 -49.5 15.7 -- -45.2SE-GUI-3B [44]55.8 7.6 35.147.0 4.9 29.038.1 12.5 31.861.8 16.4 43.359.9 24.5 50.940.2 12.4 25.550.4 11.835.9Jedi-3B [36] 61.0 13.8 38.153.5 8.4 34.627.4 9.4 23.054.2 18.2 38.664.4 32.1 57.038.3 9.0 25.049.8 13.736.1GUI-G1-3B [48]50.7 10.3 31.136.6 11.9 26.639.6 9.4 32.261.8 30.0 48.067.2 32.1 59.123.5 10.6 16.149.5 16.837.1RegionFocus-7B [22]53.2 3.4 29.142.9 4.9 27.028.4 ","claim_type":"baseline","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"Prior work on multimodal ICL largely focuses on visual perception tasks (e.g., VQA [2], OCR [47]), where correct answers can often be inferred from surface-level visual similarity without discovering complex patterns across cases. The genuine ICL capabilities of VLMs on reasoning-intensive tasks remain largely unexplored. Therefore, we select three categories of benchmark : VQA [2], MMIQ [7], and MDK12 [ 55], to examine visual perception, logical reasoning, and STEM knowledge, respectively. We p","claim_type":"dataset","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"For any centered vision embedding, PCA gives coefficients and a rank-𝑘reconstruction: c=P ⊤ 𝑘 (v−𝝁), ˆv=P 𝑘c+𝝁.(6) We report the relative reconstruction error as the fraction of centered vision-embedding variance not captured by the rank-𝑘subspace: RelMSE(𝑘)= 1 𝑁 Í𝑁 𝑖=1 v𝑖 − P𝑘P⊤ 𝑘 (v𝑖 −𝝁) +𝝁 \u0001 2 2 1 𝑁 Í𝑁 𝑖=1 ∥v𝑖 −𝝁 ∥2 2 =1− Í𝑘 𝑗=1 𝜆 𝑗 Í𝑑 𝑗=1 𝜆 𝑗 ,(7) where 𝜆 𝑗 are the eigenvalues of the empirical covariance. Lower relative reconstruction error means that the PCA subspace retains more of the aux","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"dynamic modality chain. For analytical tasks requiring transitive logic, we introduce a data-efficient SFT strategy on Music-AVQA. Extensive experiments demonstrate that our approach consistently matches or outperforms specialized models on fine-grained music perception benchmarks (Music-AVQA [15]), open-domain scenarios (AV-Odyssey [7], DailyOmni [39], OmniBench [18], WorldSense [10], AV-Counting [23]), and cross-modal hallucinations (AVH- Bench [24]). Our contributions are summarized as follow","claim_type":"dataset","confidence":0.9,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks DeepEyes: Incentivizing \"Thinking with Images\" via Reinforcement Learning because it crossed a citation-hub threshold. Current citing contexts most often use it as background evidence (20 contexts).","role_counts":[{"n":20,"context_role":"background"},{"n":6,"context_role":"baseline"},{"n":2,"context_role":"dataset"}]},"error":null,"updated_at":"2026-07-02T20:02:56.299934+00:00"},"author_expand":{"job_type":"author_expand","status":"succeeded","result":{"authors_linked":[{"id":"635e58dc-ad4a-4c68-8579-5f2299255298","orcid":null,"display_name":"Ziwei Zheng"},{"id":"237df7e5-88bb-4c81-8426-fda281b6df52","orcid":null,"display_name":"Michael Yang"},{"id":"98011042-bb39-4251-a56a-a8e7faf8780e","orcid":null,"display_name":"Jack Hong"},{"id":"4c35a727-7b1c-4999-8fbc-93130b176c85","orcid":null,"display_name":"Chenxiao Zhao"},{"id":"a22e70a1-7b28-43ce-a2af-3da9391add10","orcid":null,"display_name":"Guohai Xu"},{"id":"4efa2858-9c19-41bf-a92d-123c0950dd89","orcid":null,"display_name":"Le Yang"}]},"error":null,"updated_at":"2026-07-02T20:02:56.296330+00:00"},"context_extract":{"job_type":"context_extract","status":"succeeded","result":{"enqueued_papers":25},"error":null,"updated_at":"2026-05-14T18:06:23.654596+00:00"},"graph_features":{"job_type":"graph_features","status":"succeeded","result":{"co_cited":[{"title":"Qwen2.5-VL Technical Report","work_id":"69dffacb-bfe8-442d-be86-48624c60426f","shared_citers":23},{"title":"Qwen3-VL Technical Report","work_id":"1fe243aa-e3c0-4da6-b391-4cbcfc88d5c0","shared_citers":20},{"title":"DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models","work_id":"c5006563-f3ec-438a-9e35-b7b484f34828","shared_citers":15},{"title":"Pixel Reasoner: Incentivizing Pixel-Space Reasoning with Curiosity-Driven Reinforcement Learning","work_id":"878c3e90-ce55-4ba3-a588-2abe369013e6","shared_citers":15},{"title":"Thinking with Images for Multimodal Reasoning: Foundations, Methods, and Future Frontiers","work_id":"760ebd7d-977d-4280-afae-adb421d49ed4","shared_citers":13},{"title":"InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models","work_id":"fe8637aa-12bc-4434-8d36-9f57b5eebcbe","shared_citers":12},{"title":"Deepeyesv2: Toward agentic multimodal model","work_id":"0bc8f779-a287-4c6f-95b2-daffd15ad044","shared_citers":11},{"title":"DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning","work_id":"e6b75ad5-2877-4168-97c8-710407094d20","shared_citers":11},{"title":"GPT-4o System Card","work_id":"f37bf1c7-4964-4e56-9762-d20da8d9009f","shared_citers":11},{"title":"Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution","work_id":"8abcfe4f-e0fb-44b7-9123-448fac95f90a","shared_citers":11},{"title":"Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities","work_id":"008df105-2fdd-45d8-857a-8e35868aecb6","shared_citers":10},{"title":"Thyme: Think beyond images","work_id":"f91f31cb-6ce5-43a8-b71e-9fc90a2b4160","shared_citers":10},{"title":"InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency","work_id":"b8f5e260-fff5-444e-bcf5-2c42cfefd83d","shared_citers":9},{"title":"Latent visual reasoning","work_id":"b6468cfa-4f13-4e02-b0ec-24ff5cd6785a","shared_citers":9},{"title":"Mini-o3: Scaling up reasoning patterns and interaction turns for visual search","work_id":"8d8763c6-6201-4243-9dab-975ba02a78db","shared_citers":9},{"title":"Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models","work_id":"38998646-34ee-4605-b661-ab356f16d6e5","shared_citers":9},{"title":"LLaVA-OneVision: Easy Visual Task Transfer","work_id":"f5f2452b-f2a9-49ac-b38d-c76e18cdfe49","shared_citers":8},{"title":"Openthinkimg: Learning to think with images via visual tool reinforcement learning","work_id":"3939edf9-d5d5-4c79-bd9c-02e7961fda21","shared_citers":8},{"title":"Chain-of-focus: Adaptive visual search and zooming for multimodal reasoning via rl","work_id":"c157db40-f1a3-4bf3-b004-55183727b771","shared_citers":6},{"title":"Mme-realworld: Could your multimodal llm challenge high-resolution real-world scenarios that are difficult for humans?","work_id":"140d79fb-a3d2-4af2-a436-9d997c171f61","shared_citers":6},{"title":"Monet: Reasoning in latent visual space beyond images and language","work_id":"ce4ac531-7907-4d8a-afe2-9a0dae21ffa5","shared_citers":6},{"title":"DAPO: An Open-Source LLM Reinforcement Learning System at Scale","work_id":"64019d00-0b11-4bbd-b173-b46c8fad0157","shared_citers":5},{"title":"Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling","work_id":"ee70bdc8-4656-4849-ada7-ce42a2278d70","shared_citers":5},{"title":"GPT-4 Technical Report","work_id":"b928e041-6991-4c08-8c81-0359e4097c7b","shared_citers":5}],"time_series":[{"n":37,"year":2026}],"dependency_candidates":[]},"error":null,"updated_at":"2026-05-14T18:06:23.692550+00:00"},"identity_refresh":{"job_type":"identity_refresh","status":"succeeded","result":{"items":[{"title":"Qwen3 Technical Report","outcome":"unchanged","work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e","resolver":"local_arxiv","confidence":0.98,"old_work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e"}],"counts":{"fixed":0,"merged":0,"unchanged":1,"quarantined":0,"needs_external_resolution":0},"errors":[],"attempted":1},"error":null,"updated_at":"2026-05-14T18:06:44.531111+00:00"},"role_polarity":{"job_type":"role_polarity","status":"succeeded","result":{"title":"DeepEyes: Incentivizing \"Thinking with Images\" via Reinforcement Learning","claims":[{"claim_text":"Large Vision-Language Models excel at multimodal understanding but struggle to deeply integrate visual information into their predominantly text-based reasoning processes, a key challenge in mirroring human cognition. To address this, we introduce DeepEyes, a model that learns to \"think with images\", trained end-to-end with reinforcement learning without requiring pre-collected reasoning data for cold-start supervised fine-tuning (SFT). Notably, this ability emerges natively, leveraging the model's own grounding capability as an intrinsic function rather than relying on external specialized mo","claim_type":"abstract","evidence_strength":"source_metadata"},{"claim_text":"Geox-bench: Benchmarking cross-view geo-localization and pose estimation capabilities of large multimodal models.arXiv preprint arXiv:2511.13259, 2025. [104] Ziwei Zheng, Michael Yang, Jack Hong, Chenxiao Zhao, Guohai Xu, Le Yang, Chao Shen, and Xing Yu. Deepeyes: Incentivizing\" thinking with images\" via reinforcement learning. arXiv preprint arXiv:2505.14362, 2025. [105] Xiran Zhou, Yi Wen, Zhenfeng Shao, Wenwen Li, Kaiyuan Li, Honghao Li, Xiao Xie, and Zhigang Yan. Cartomark: a benchmark datas","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"ReSearch [108] External Qwen2.5-7B/32B-Instruct /githubGitHub StepSearch [267] External Qwen2.5-3B/7B-Base/Instruct /githubGitHub DeepResearcher [268] External Qwen2.5-7B-Instruct /githubGitHub WebDancer [106] External Qwen2.5-7B/32B, QWQ-32B /githubGitHub WebThinker [269] External QwQ-32B, DeepSeek-R1-Distilled-Qwen, Qwen2.5-32B/githubGitHub WebSailor [105] External Qwen2.5-3B/7B/32B/72B /githubGitHub WebWatcher [270] External Qwen2.5-VL-7B/32B /githubGitHub WebShaper [271] External Qwen-2.5-32","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"3 57.043.9 12.4 29.649.1 14.135.7InfiGUI-G1-3B [21]64.9 20.0 -51.5 16.8 -50.8 25.0 -68.8 32.7 -70.6 32.1 -49.5 15.7 -- -45.2SE-GUI-3B [44]55.8 7.6 35.147.0 4.9 29.038.1 12.5 31.861.8 16.4 43.359.9 24.5 50.940.2 12.4 25.550.4 11.835.9Jedi-3B [36] 61.0 13.8 38.153.5 8.4 34.627.4 9.4 23.054.2 18.2 38.664.4 32.1 57.038.3 9.0 25.049.8 13.736.1GUI-G1-3B [48]50.7 10.3 31.136.6 11.9 26.639.6 9.4 32.261.8 30.0 48.067.2 32.1 59.123.5 10.6 16.149.5 16.837.1RegionFocus-7B [22]53.2 3.4 29.142.9 4.9 27.028.4 ","claim_type":"baseline","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"Prior work on multimodal ICL largely focuses on visual perception tasks (e.g., VQA [2], OCR [47]), where correct answers can often be inferred from surface-level visual similarity without discovering complex patterns across cases. The genuine ICL capabilities of VLMs on reasoning-intensive tasks remain largely unexplored. Therefore, we select three categories of benchmark : VQA [2], MMIQ [7], and MDK12 [ 55], to examine visual perception, logical reasoning, and STEM knowledge, respectively. We p","claim_type":"dataset","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"For any centered vision embedding, PCA gives coefficients and a rank-𝑘reconstruction: c=P ⊤ 𝑘 (v−𝝁), ˆv=P 𝑘c+𝝁.(6) We report the relative reconstruction error as the fraction of centered vision-embedding variance not captured by the rank-𝑘subspace: RelMSE(𝑘)= 1 𝑁 Í𝑁 𝑖=1 v𝑖 − P𝑘P⊤ 𝑘 (v𝑖 −𝝁) +𝝁 \u0001 2 2 1 𝑁 Í𝑁 𝑖=1 ∥v𝑖 −𝝁 ∥2 2 =1− Í𝑘 𝑗=1 𝜆 𝑗 Í𝑑 𝑗=1 𝜆 𝑗 ,(7) where 𝜆 𝑗 are the eigenvalues of the empirical covariance. Lower relative reconstruction error means that the PCA subspace retains more of the aux","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"dynamic modality chain. For analytical tasks requiring transitive logic, we introduce a data-efficient SFT strategy on Music-AVQA. Extensive experiments demonstrate that our approach consistently matches or outperforms specialized models on fine-grained music perception benchmarks (Music-AVQA [15]), open-domain scenarios (AV-Odyssey [7], DailyOmni [39], OmniBench [18], WorldSense [10], AV-Counting [23]), and cross-modal hallucinations (AVH- Bench [24]). Our contributions are summarized as follow","claim_type":"dataset","confidence":0.9,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks DeepEyes: Incentivizing \"Thinking with Images\" via Reinforcement Learning because it crossed a citation-hub threshold. Current citing contexts most often use it as background evidence (20 contexts).","role_counts":[{"n":20,"context_role":"background"},{"n":6,"context_role":"baseline"},{"n":2,"context_role":"dataset"}]},"error":null,"updated_at":"2026-07-02T20:02:56.302409+00:00"},"summary_claims":{"job_type":"summary_claims","status":"succeeded","result":{"title":"DeepEyes: Incentivizing \"Thinking with Images\" via Reinforcement Learning","claims":[{"claim_text":"Large Vision-Language Models excel at multimodal understanding but struggle to deeply integrate visual information into their predominantly text-based reasoning processes, a key challenge in mirroring human cognition. To address this, we introduce DeepEyes, a model that learns to \"think with images\", trained end-to-end with reinforcement learning without requiring pre-collected reasoning data for cold-start supervised fine-tuning (SFT). Notably, this ability emerges natively, leveraging the model's own grounding capability as an intrinsic function rather than relying on external specialized mo","claim_type":"abstract","evidence_strength":"source_metadata"}],"why_cited":"Pith tracks DeepEyes: Incentivizing \"Thinking with Images\" via Reinforcement Learning because it crossed a citation-hub threshold.","role_counts":[]},"error":null,"updated_at":"2026-05-14T18:06:40.172223+00:00"}},"summary":{"title":"DeepEyes: Incentivizing \"Thinking with Images\" via Reinforcement Learning","claims":[{"claim_text":"Large Vision-Language Models excel at multimodal understanding but struggle to deeply integrate visual information into their predominantly text-based reasoning processes, a key challenge in mirroring human cognition. To address this, we introduce DeepEyes, a model that learns to \"think with images\", trained end-to-end with reinforcement learning without requiring pre-collected reasoning data for cold-start supervised fine-tuning (SFT). Notably, this ability emerges natively, leveraging the model's own grounding capability as an intrinsic function rather than relying on external specialized mo","claim_type":"abstract","evidence_strength":"source_metadata"}],"why_cited":"Pith tracks DeepEyes: Incentivizing \"Thinking with Images\" via Reinforcement Learning because it crossed a citation-hub threshold.","role_counts":[]},"graph":{"co_cited":[{"title":"Qwen2.5-VL Technical Report","work_id":"69dffacb-bfe8-442d-be86-48624c60426f","shared_citers":23},{"title":"Qwen3-VL Technical Report","work_id":"1fe243aa-e3c0-4da6-b391-4cbcfc88d5c0","shared_citers":20},{"title":"DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models","work_id":"c5006563-f3ec-438a-9e35-b7b484f34828","shared_citers":15},{"title":"Pixel Reasoner: Incentivizing Pixel-Space Reasoning with Curiosity-Driven Reinforcement Learning","work_id":"878c3e90-ce55-4ba3-a588-2abe369013e6","shared_citers":15},{"title":"Thinking with Images for Multimodal Reasoning: Foundations, Methods, and Future Frontiers","work_id":"760ebd7d-977d-4280-afae-adb421d49ed4","shared_citers":13},{"title":"InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models","work_id":"fe8637aa-12bc-4434-8d36-9f57b5eebcbe","shared_citers":12},{"title":"Deepeyesv2: Toward agentic multimodal model","work_id":"0bc8f779-a287-4c6f-95b2-daffd15ad044","shared_citers":11},{"title":"DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning","work_id":"e6b75ad5-2877-4168-97c8-710407094d20","shared_citers":11},{"title":"GPT-4o System Card","work_id":"f37bf1c7-4964-4e56-9762-d20da8d9009f","shared_citers":11},{"title":"Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution","work_id":"8abcfe4f-e0fb-44b7-9123-448fac95f90a","shared_citers":11},{"title":"Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities","work_id":"008df105-2fdd-45d8-857a-8e35868aecb6","shared_citers":10},{"title":"Thyme: Think beyond images","work_id":"f91f31cb-6ce5-43a8-b71e-9fc90a2b4160","shared_citers":10},{"title":"InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency","work_id":"b8f5e260-fff5-444e-bcf5-2c42cfefd83d","shared_citers":9},{"title":"Latent visual reasoning","work_id":"b6468cfa-4f13-4e02-b0ec-24ff5cd6785a","shared_citers":9},{"title":"Mini-o3: Scaling up reasoning patterns and interaction turns for visual search","work_id":"8d8763c6-6201-4243-9dab-975ba02a78db","shared_citers":9},{"title":"Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models","work_id":"38998646-34ee-4605-b661-ab356f16d6e5","shared_citers":9},{"title":"LLaVA-OneVision: Easy Visual Task Transfer","work_id":"f5f2452b-f2a9-49ac-b38d-c76e18cdfe49","shared_citers":8},{"title":"Openthinkimg: Learning to think with images via visual tool reinforcement learning","work_id":"3939edf9-d5d5-4c79-bd9c-02e7961fda21","shared_citers":8},{"title":"Chain-of-focus: Adaptive visual search and zooming for multimodal reasoning via rl","work_id":"c157db40-f1a3-4bf3-b004-55183727b771","shared_citers":6},{"title":"Mme-realworld: Could your multimodal llm challenge high-resolution real-world scenarios that are difficult for humans?","work_id":"140d79fb-a3d2-4af2-a436-9d997c171f61","shared_citers":6},{"title":"Monet: Reasoning in latent visual space beyond images and language","work_id":"ce4ac531-7907-4d8a-afe2-9a0dae21ffa5","shared_citers":6},{"title":"DAPO: An Open-Source LLM Reinforcement Learning System at Scale","work_id":"64019d00-0b11-4bbd-b173-b46c8fad0157","shared_citers":5},{"title":"Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling","work_id":"ee70bdc8-4656-4849-ada7-ce42a2278d70","shared_citers":5},{"title":"GPT-4 Technical Report","work_id":"b928e041-6991-4c08-8c81-0359e4097c7b","shared_citers":5}],"time_series":[{"n":37,"year":2026}],"dependency_candidates":[]},"authors":[{"id":"4c35a727-7b1c-4999-8fbc-93130b176c85","orcid":null,"display_name":"Chenxiao Zhao","source":"manual","import_confidence":0.72},{"id":"a22e70a1-7b28-43ce-a2af-3da9391add10","orcid":null,"display_name":"Guohai Xu","source":"manual","import_confidence":0.72},{"id":"98011042-bb39-4251-a56a-a8e7faf8780e","orcid":null,"display_name":"Jack Hong","source":"manual","import_confidence":0.72},{"id":"4efa2858-9c19-41bf-a92d-123c0950dd89","orcid":null,"display_name":"Le Yang","source":"manual","import_confidence":0.72},{"id":"237df7e5-88bb-4c81-8426-fda281b6df52","orcid":null,"display_name":"Michael Yang","source":"manual","import_confidence":0.72},{"id":"635e58dc-ad4a-4c68-8579-5f2299255298","orcid":null,"display_name":"Ziwei Zheng","source":"manual","import_confidence":0.72}]}}