{"work":{"id":"f0dc2969-bd90-4a5d-81df-d557352690b6","openalex_id":"https://openalex.org/W7124446719","doi":"10.48550/arxiv.2601.10611","arxiv_id":"2601.10611","raw_key":null,"title":"Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding","authors":null,"authors_text":"Christopher Clark, Jieyu Zhang, Zixian Ma, Jae Sung Park, Mohammadreza Salehi, Rohun Tripathi","year":2026,"venue":"cs.CV","abstract":"Today's strongest video-language models (VLMs) remain proprietary. The strongest open-weight models either rely on synthetic data from proprietary VLMs, effectively distilling from them, or do not disclose their training data or recipe. As a result, the open-source community lacks the foundations needed to improve on the state-of-the-art video (and image) language models. Crucially, many downstream applications require more than just high-level video understanding; they require grounding -- either by pointing or by tracking in pixels. Even proprietary models lack this capability. We present Molmo2, a new family of VLMs that are state-of-the-art among open-source models and demonstrate exceptional new capabilities in point-driven grounding in single image, multi-image, and video tasks. Our key contribution is a collection of 7 new video datasets and 2 multi-image datasets, including a dataset of highly detailed video captions for pre-training, a free-form video Q&A dataset for fine-tuning, a new object tracking dataset with complex queries, and an innovative new video pointing dataset, all collected without the use of closed VLMs. We also present a training recipe for this data utilizing an efficient packing and message-tree encoding scheme, and show bi-directional attention on vision tokens and a novel token-weight strategy improves performance. Our best-in-class 8B model outperforms others in the class of open weight and data models on short videos, counting, and captioning, and is competitive on long-videos. On video-grounding Molmo2 significantly outperforms existing open-weight models like Qwen3-VL (35.5 vs 29.6 accuracy on video counting) and surpasses proprietary models like Gemini 3 Pro on some tasks (38.4 vs 20.0 F1 on video pointing and 56.2 vs 41.1 J&F on video tracking).","external_url":"https://arxiv.org/abs/2601.10611","cited_by_count":1,"metadata_source":"pith","metadata_fetched_at":"2026-08-05T02:28:24.338817+00:00","pith_arxiv_id":"2601.10611","created_at":"2026-05-09T22:34:07.286810+00:00","updated_at":"2026-08-05T02:28:24.338817+00:00","title_quality_ok":true,"display_title":"Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding","render_title":"Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding"},"hub":{"state":{"work_id":"f0dc2969-bd90-4a5d-81df-d557352690b6","tier":"hub","tier_reason":"10+ Pith inbound or 1,000+ external citations","pith_inbound_count":50,"external_cited_by_count":1,"distinct_field_count":6,"first_pith_cited_at":"2026-02-26T09:15:34+00:00","last_pith_cited_at":"2026-07-01T16:04:24+00:00","author_build_status":"not_needed","summary_status":"needed","contexts_status":"needed","graph_status":"needed","ask_index_status":"not_needed","reader_status":"not_needed","recognition_status":"not_needed","updated_at":"2026-08-23T12:29:32.857711+00:00","tier_text":"hub"},"tier":"hub","role_counts":[{"context_role":"background","n":6},{"context_role":"baseline","n":3},{"context_role":"other","n":1}],"polarity_counts":[{"context_polarity":"background","n":6},{"context_polarity":"baseline","n":3},{"context_polarity":"unclear","n":1}],"runs":{},"summary":{},"graph":{},"authors":[]}}