Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-15T01:55:12.501409Z
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 100 of 137 outbound references and 79 inbound Pith citation observations for arXiv:2409.17146.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-15T01:55:12.501409Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-05T12:28:29.027544Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
100 of 137 outbound references displayed
External citation measurements
8
pith, observed 2026-08-05T02:28:24.338817Z
Observation 8d916b44-044f-43ce-999e-1e801b25f4e3 · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4f641411-604c-47e7-b3b8-01b6f4a3ac00 · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models TallyQA: Answering complex counting questions
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 66afdab2-c2de-4c04-a32c-70269fe46219 · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Pixtral 12B
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 175d3db8-aadb-45a9-a5b7-84b6d9cb67bd · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Yi: Open Foundation Models by 01.AI
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8d3e417b-c926-43b2-a6db-3b7376cf41a0 · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models The Llama 3 Herd of Models
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c6af7ef7-eff2-4d66-8b95-a607f242baab · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Flamingo: a visual language model for few-shot learning
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9f6b7d5b-83e7-4505-93f4-18e264455955 · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models The claude 3 model family: Opus, sonnet, haiku
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e187df9f-bee6-468b-940d-dec9ac9dc948 · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Layer normalization
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 27e7d071-5587-4757-8e1c-08f729938d42 · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Fuyu-8b: A multimodal architecture for ai agents
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation aa522005-f79e-433f-adfd-34aef21ec019 · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models PaliGemma: A versatile 3B VLM for transfer
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8c16d6ae-a6dd-4ee4-a60f-c6ac72f5042f · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Scene text visual question answering
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2b1945b4-3e19-4b00-866e-9e9ffca1aefe · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Honeybee: Locality-enhanced projector for multimodal llm
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1b9acf5c-c89a-4316-912f-b40cbaa05b3a · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0c4271e4-c97b-4810-9b10-48639d805bf4 · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models EVLM: An Efficient Vision-Language Model for Visual Understanding
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3d13dbdd-8cfd-4bf8-9514-8fd46658da54 · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models ShareGPT4V: Improving Large Multi-Modal Models with Better Captions
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 10afaa3d-1884-4f6b-a3e5-33fb556baf50 · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Evaluating Large Language Models Trained on Code
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e27abcfe-0a82-4689-8963-6ac137d7c451 · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Microsoft COCO Captions: Data Collection and Evaluation Server
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c9085aad-b652-4417-917f-ead2dd30e829 · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models PaLI-3 Vision Language Models: Smaller, Faster, Stronger
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 76eee52b-6244-42d1-9162-682f15aa8c30 · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b2d6ab35-3a86-4b40-b859-530942728310 · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Reproducible scaling laws for contrastive language-image learning
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 989c6a9f-11af-4e98-976c-94a7c3d554ac · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Chatbot arena: An open platform for evaluating LLMs by human preference
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3a7116c2-28ef-48f8-b48b-af9dfecf4bad · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 13fcd1d4-5482-468c-a429-22af94e90abe · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f65f18fc-2f99-4113-bd09-49470ff73d89 · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Training Verifiers to Solve Math Word Problems
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5100c3e0-82f6-4c46-8d65-fe867ba98c06 · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models On implementing 2d rectangular assignment algo- rithms
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 51d4a3e9-1270-49c6-b3de-5a379c419fed · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Unresolved cited work
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 16ae0048-5fc7-49ca-96b3-56406959f01c · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models NVLM: Open Frontier-Class Multimodal LLMs
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 92d86baa-5452-4c60-85bb-c5ede3751c52 · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models FlashAttention-2: Faster attention with better parallelism and work partitioning
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d855a28e-3556-498c-928b-b503bd2928f3 · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Fu, Stefano Ermon, Atri Rudra, and Christopher R´e
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0f776ec2-dd71-4348-8b5c-0f340008bf15 · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation cbc0f833-6031-4516-88c4-a398773fc0e1 · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models An image is worth 16x16 words: Transformers for image recognition at scale
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0c9c9db9-68f1-43e9-a6a6-5311c3768a0b · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models VILA$^2$: VILA Augmented VILA
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 320a91aa-1659-43bd-a7df-d3d95a7f6577 · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Devise: A deep visual-semantic embedding model
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5c6a522d-e36f-4c50-bf24-b9a962dfe8a9 · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models VITA: Towards Open-Source Interactive Omni Multimodal LLM
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6681f19c-668e-4764-9aab-1cf21dff7002 · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Scaling Synthetic Data Creation with 1,000,000,000 Personas
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0a1676c9-5182-4139-90bd-abeaf258e485 · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Making the V in VQA matter: Elevating the role of image understanding in visual question answering
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 363513ea-6dbc-40b0-b13d-08ef640f4f16 · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Unresolved cited work
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b85563b5-15ee-4540-9494-26b27446cbd5 · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Measuring massive multitask language understanding
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b418d3cf-72ca-4d4f-8b7b-5937e73adeaa · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Mea- suring mathematical problem solving with the math dataset
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 86082a46-60d2-4806-867e-d9c659cf60b6 · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Accu- mulated gradient normalization
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3ed25625-ffcc-4338-b771-77e891f93c84 · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models CogVLM2: Visual Language Models for Image and Video Understanding
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation bde59e58-be69-461e-a6a5-8f734c3a563e · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models mplug-docowl 1.5: Unified structure learning for ocr-free document understanding
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 235d57ff-4e77-4de5-898e-1407a60d7755 · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Scaling up visual and vision-language representation learning with noisy text supervision
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 95518f9a-3b1a-4faa-8955-5d1b15b44eb1 · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models MANTIS: Interleaved Multi-Image Instruction Tuning
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 18e20941-9771-4b83-aea7-d2db56a4fc80 · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models A shortest augmenting path algo- rithm for dense and sparse linear assignment problems.Computing
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 69ae4a02-b940-450c-943e-560be1802068 · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models DVQA: Understanding data visualizations via question answering
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 809132a1-3500-48e1-be6e-5059b88d2933 · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models FigureQA: An Annotated Figure Dataset for Visual Reasoning
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 295792de-3296-4a4f-bbc7-0ba24a18c730 · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Prismatic VLMs: Investigating the Design Space of Visually-Conditioned Language Models
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 58941f00-4e33-426e-ba9c-28ea8bb3938e · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models A diagram is worth a dozen images
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f93cecd4-64e1-474e-a157-5ad3b6a8f9ab · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Adam: A method for stochastic optimization
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e35b5c95-b6f8-4a6c-a3c7-fc9c18609abd · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Berg, Wan-Yen Lo, Piotr Doll´ar, and Ross Girshick
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 22c41068-de0b-4278-883e-0220f1872a23 · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Shamma, Michael S
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1b0193d8-2822-4c1f-adfc-bc2c5f70e459 · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models The open images dataset v4: Unified image classification, object detection, and visual relation- ship detection at scale
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 33e2b272-aef7-43ed-a3d4-b2fd70c979bb · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Building and better understanding vision-language models: insights and future directions
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 88a0a4bb-f973-4747-8366-f7b26dc263b2 · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Unlocking the conversion of Web Screenshots into HTML Code with the WebSight Dataset
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 73c0fa3e-d4f1-46eb-bce9-1b34fabd9625 · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models OtterHD: A High-Resolution Multi-modality Model
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b975245f-b26e-4cf2-ae75-f3254cc1188d · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models MIMIC-IT: Multi-Modal In-Context Instruction Tuning
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 733a87e2-4dcf-4f13-88ca-c5e9c83c261b · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Otter: A Multi-Modal Model with In-Context Instruction Tuning
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0a1fa6ab-bb1e-4d10-8053-2d0f6c1ab2e4 · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models LLaVA-OneVision: Easy Visual Task Transfer
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 175a6c54-d799-48dd-b88f-a9074bf358cd · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Blip: Boot- strapping language-image pre-training for unified vision-language understanding and generation
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation efaa6182-b356-4a6f-a8b5-8cab6baa616d · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Covlm: Composing visual entities and relationships in large language models via communicative de- coding
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1f7d1fa6-f636-4c3c-8fce-e625e4fc7b27 · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models On the Effects of Data Scale on UI Control Agents
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation fd9b7035-1361-4cf5-ab62-a68f0c2c147b · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Monkey: Image resolution and text label are important things for large multi-modal models
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7c4f82bd-232a-4ce2-9930-e2a40a6011cb · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models MoE-LLaVA: Mixture of Experts for Large Vision-Language Models
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ae981ab4-9564-4a22-b692-c3f0b5fc71df · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Mi- crosoft coco: Common objects in context
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4cd61f70-2dbe-4925-a00c-1afc98be521d · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models GRES: Generalized referring expression segmentation
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b9c48908-59ac-48cb-9350-c02d431efc71 · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models SPHINX-x: Scaling data and parameters for a family of multi- modal large language models
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 00907bf3-cd35-4657-b388-5e313e8e44d5 · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models MMC: Advancing Multimodal Chart Understanding with Large-scale Instruction Tuning
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a154ac9f-d09e-4fd9-b233-65d1d867ba7c · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Visual instruction tuning
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 259a9fae-5cc0-40ef-86d4-71f275b4b70d · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Improved baselines with visual instruction tuning
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 47de1e46-8e0d-4f37-8efe-49916253543d · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Llava-next: Improved reasoning, ocr, and world knowledge
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 33474c40-71df-42f0-8198-39abc54cb99f · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models SGDR: Stochastic Gradient Descent with Warm Restarts
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7d8bb6a7-159b-4a39-bb7e-48ec9315cbf4 · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Decoupled weight decay regu- larization
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 81c71248-cc10-452a-9eae-4158779acee1 · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models DeepSeek-VL: Towards Real-World Vision-Language Understanding
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2e3ce511-9b32-4b58-b1c2-d783d8aa6441 · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Unified-io 2: Scaling autoregressive multimodal models with vision language audio and action
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d7bc6368-5151-44a8-aa86-f1930c8bc96c · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Learn to explain: Multimodal reasoning via thought chains for sci- ence question answering
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a8b3a234-af38-4831-92ef-2a68e20ae143 · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Dy- namic prompt learning via policy gradient for semi-structured math- ematical reasoning
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5d625152-6f44-4ac7-b117-d7fc2b6a6040 · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models MathVista: Evaluating mathematical reasoning of foundation models in visual contexts
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 55c7b206-57f1-43f5-a9d1-8e5377195d9d · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Cheap and quick: Efficient vision-language in- struction tuning for large language models
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9245dc91-9630-459b-b637-553008f44f89 · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models ExpertQA: Expert-Curated Questions and Attributed Answers
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b5284158-b423-4a2f-b560-1280095a16aa · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models OK-VQA: A visual question answering benchmark re- quiring external knowledge
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 957c4882-cb5c-40eb-8ecb-80e9b9472453 · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models ChartQA: A benchmark for question answering about charts with visual and logical reasoning
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation bf158c1b-e9cc-4f3d-8506-32749ddfa62f · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models DocVQA: A dataset for VQA on document images
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3dd108be-bf0c-4aa3-b9a1-4729fa220f9d · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models InfographicVQA
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation db67f8a3-8083-40ea-b8f4-c41282c2bd7a · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation fe9fa82c-70a6-4603-998d-e2ad55fb2439 · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models PlotQA: Reasoning over scientific plots
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7e25d38f-c1c0-45d2-8339-650a94a7e581 · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models OLMoE: Open Mixture-of-Experts Language Models
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ab04819f-78a2-43de-af8c-ee4b016baa7a · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models GPT-4 Technical Report
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 46f43994-70be-4b43-becc-ad381b784019 · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models GPT-4o mini system card
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7d7e216d-8211-46a2-bbbb-d199f3f7dd8a · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models GPT-4o System Card
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e13f6784-27fc-4f31-a098-5fce1a425133 · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models DINOv2: Learning Robust Visual Features without Supervision
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0b673372-18ef-4d2c-ac6e-5f632337b2b9 · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Kosmos-2: Grounding Multimodal Large Language Models to the World
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f7b0621a-0b12-4e6a-945b-0833c445c289 · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Connecting vision and language with lo- calized narratives
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation eac89216-e82d-4cc1-bfbe-f2966e4a05f8 · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Jack of all tasks, master of many: Designing general- purpose coarse-to-fine vision-language model
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f24fbc4b-1c53-4e5e-9523-7b9dcdbd4951 · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Improving language understanding by generative pre- training
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1472ef67-4680-49b7-a62a-ec13abb25d32 · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models 28 Learning transferable visual models from natural language supervi- sion
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8b170e06-f38b-4061-a0ed-ea4481ae3150 · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models GLaMM: Pixel grounding large multimodal model
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e717be9b-542a-4e2a-a2b6-0b48076ffc46 · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models LAION-5B: An open large-scale dataset for training next generation image-text models
Reference 98
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1d538a46-71da-429e-bd0c-eee4f3fb3589 · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models A-OKVQA: A benchmark for vi- sual question answering using world knowledge
Reference 99
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b8dae70c-7937-498b-a12e-bfb158ab969e · outbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Towards VQA models that can read
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8af8519b-85eb-4d2b-a7eb-050fdd0a6cad · inbound
Pixtral 12B Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 20e8761d-11f8-4f0f-8008-9f08393c93ac · inbound
Self-Correcting Text-to-Video Generation with Misalignment Detection and Localized Refinement Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ba0311bd-3799-4d48-89c7-1eed439d2f76 · inbound
PaliGemma 2: A Family of Versatile VLMs for Transfer Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 30682552-5cf8-4c71-98dd-c76d6aa030b7 · inbound
Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b688c074-fbc0-45fe-a1fa-7b85d8d9a822 · inbound
NVILA: Efficient Frontier Visual Language Models Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7286b563-8a32-4138-aa2c-422486401088 · inbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c81f48ad-3d61-48e7-944c-fa3a6ae5ab3e · inbound
DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e70fe4f2-ee1e-4a49-ba6e-5e76eff24780 · inbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b1eea010-638d-4c12-84de-b93cf2eb9da3 · inbound
VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 951b5426-f4ed-4f0a-983e-cc0efc7b2ecd · inbound
Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2f8a1c20-5fed-4c20-bdc4-7856133ce447 · inbound
Seeing is Understanding: Unlocking Causal Attention into Modality-Mutual Attention for Multimodal LLMs Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 516d51f7-c38a-44a3-bc4d-355dbeb60f78 · inbound
Capacity-Aware Inference: Mitigating the Straggler Effect in Mixture of Experts Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0a8b3b20-0fc7-480b-8f5d-bd4d388550d3 · inbound
GuideDog: A Real-World Egocentric Multimodal Dataset for Blind and Low-Vision Accessibility-Aware Guidance Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5e0975f2-ce4c-4ade-b961-4636735b67c7 · inbound
SmolVLM: Redefining small and efficient multimodal models Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 83a0b490-906e-40ed-b2b5-23d309d7b10e · inbound
FLARE: Fully Integration of Vision-Language Representations for Deep Cross-Modal Understanding Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e00a2882-d65a-4442-a090-1dfebde268d7 · inbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c7776a7b-b344-498d-ab3f-2aeff41dcc31 · inbound
Perception Encoder: The best visual embeddings are not at the output of the network Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4903c79d-2e7a-4715-ba10-a683ddc771bf · inbound
$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e4e88b71-6c17-4e55-9f22-d64811bd6c3d · inbound
Seed1.5-VL Technical Report Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 684ca5b8-d3d0-49be-9698-c82e50151a9e · inbound
FieldWorkArena: Agentic AI Benchmark for Real Field Work Tasks Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 228a6080-83c3-463e-8d1a-0a851b2d3eb3 · inbound
Grounded Reinforcement Learning for Visual Reasoning Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 261929bf-d31d-44c9-9e91-484d0399b0e0 · inbound
Common Inpainted Objects In-N-Out of Context Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 003da85c-77e9-43e6-b3b6-1f9d9fa16fb6 · inbound
A Survey on Vision-Language-Action Models: An Action Tokenization Perspective Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 182
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7f2e98f7-69b6-4ef1-a6ce-cad52c7a2f7e · inbound
High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 15df2ffb-9ffc-4e60-b1fb-34703a5eb7c9 · inbound
When Seeing Overrides Knowing: Disentangling Knowledge Conflicts in Vision-Language Models Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 36458622-c441-4ed3-9332-8a5e2e1554c7 · inbound
Embodied-R1: Reinforced Embodied Reasoning for General Robotic Manipulation Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0cce3dfb-280c-4057-949d-81b8e2770e95 · inbound
InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f696242a-dd65-4674-8dfc-64c84568f856 · inbound
Kwai Keye-VL 1.5 Technical Report Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90da2d0b-414b-43cc-878b-39ed1d204b76 · inbound
Improving Large Vision and Language Models by Learning from a Panel of Peers Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4026b94f-1a6f-412c-81d3-6e399801e536 · inbound
Reinforced Visual Perception with Tools Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd4df457-7e3a-4575-93f0-db03883c2852 · inbound
LExI: Layer-Adaptive Active Experts for Efficient MoE Model Inference Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44068fd0-5849-4e91-bf7d-b33f4cf17e88 · inbound
Weakly-Supervised Learning of Dense Functional Correspondences Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d539798-7a22-43b1-a53d-fab53a9356a0 · inbound
Promptception: How Sensitive Are Large Multimodal Models to Prompts? Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 130cf91b-85ff-4fd1-989f-523aa9f82f7d · inbound
Towards Better Dental AI: A Multimodal Benchmark and Instruction Dataset for Panoramic X-ray Analysis Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4964d9db-8401-4b26-bb24-e2b2d0c1c663 · inbound
Improving Fungi Prototype Representations for Few-Shot Classification Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7460dde3-c7a3-4c46-9163-9122caeb58c0 · inbound
InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9e3b4d55-2dda-411b-94b2-20e4b964ecfd · inbound
VisCoder2: Building Multi-Language Visualization Coding Agents Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 68d4ffd2-bfd0-4552-8927-d520ca3ab30f · inbound
BOP-ASK: Object-Interaction Reasoning for Vision-Language Models Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8250275f-ec65-4132-b714-23b7b4b0cd99 · inbound
Steerable Vision-Language-Action Policies for Embodied Reasoning and Hierarchical Control Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ec8ba832-38b1-4df0-ae51-2522ed3c6018 · inbound
Visual Para-Thinker: Divide-and-Conquer Reasoning for Visual Comprehension Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4710ed8f-9d25-446f-9330-8a9faee53e69 · inbound
TrajTok: Learning Trajectory Tokens enables better Video Understanding Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8cdd0965-97dd-46f2-9f63-c90efe162aaf · inbound
TrajTok: Learning Trajectory Tokens enables better Video Understanding Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da3b6ecc-b7da-44d6-93c5-a7b33b8e36b0 · inbound
TiPToP: A Modular Open-Vocabulary Robot Manipulation System That Plans Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a115856-c175-4260-8ee8-46c105ffe1f3 · inbound
Evaluating Vision Foundation Models for Pixel and Object Classification in Microscopy Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 271ab0f9-1f78-4a54-88dc-b679a7b170dd · inbound
A1: A Fully Transparent Open-Source, Adaptive and Efficient Truncated Vision-Language-Action Model Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 211cac15-8573-4f9a-b6de-edb2618fc9f0 · inbound
ABMAMBA: Multimodal Large Language Model with Aligned Hierarchical Bidirectional Scan for Efficient Video Captioning Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4828e8f5-3a9c-4eac-a049-45900830e3db · inbound
Entropy-Gradient Grounding: Training-Free Evidence Retrieval in Vision-Language Models Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5e4c86e9-9082-4271-9169-156300307c1c · inbound
Anthropogenic Regional Adaptation in Multimodal Vision-Language Model Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation cb95ea42-25b9-47d4-87b4-59507d6e45a9 · inbound
UniMesh: Unifying 3D Mesh Understanding and Generation Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 10fe9bf2-e95c-4301-af01-7f0d9f1eacf2 · inbound
PRTS: A Primitive Reasoning and Tasking System via Contrastive Representations Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 65f28d80-14f7-4cae-a642-f088e0e49f94 · inbound
Visibility-Aware Mobile Grasping in Dynamic Environments Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 298e1376-4c87-480c-92b1-cf60ddfbc751 · inbound
Visibility-Aware Mobile Grasping in Dynamic Environments Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c1a3f6cf-72a5-4c82-bce2-7e8d88395cb3 · inbound
When Relations Break: Analyzing Relation Hallucination in Vision-Language Model Under Rotation and Noise Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ee9e83c5-b2d8-4f7a-8f53-39c69e9ad358 · inbound
When Relations Break: Analyzing Relation Hallucination in Vision-Language Model Under Rotation and Noise Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 605ca1b7-7702-4d2d-af34-4d419a2969c6 · inbound
20/20 Vision Language Models: A Prescription for Better VLMs through Data Curation Alone Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 99da1db6-1974-4923-912d-a95c4bcd97d1 · inbound
20/20 Vision Language Models: A Prescription for Better VLMs through Data Curation Alone Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f2b47308-92f2-48b3-a1bb-140b9b36d1ea · inbound
Beyond Waypoints: Dual-Heatmap Grounding for Cross-Embodiment Semantic Navigation Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3754e66a-e946-447b-a587-ed2e7a417482 · inbound
Binding Visual Features Point by Point Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 09a2392a-73ee-43b9-86d8-e8694f9d8ab5 · inbound
LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0b7fb262-8dd6-44cb-ab57-d65d0ccf9d8e · inbound
VCap: Hypergeometric Rewards for Weak-to-Strong Visual Captioning Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8338fa87-ef7c-4a6a-a574-7e2476c4d907 · inbound
OpenEAI-Platform: An Open-source Embodied Artificial Intelligence Hardware-Software Unified Platform Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 186f8c83-491b-466e-a3c9-99a916e9a67c · inbound
TurtleAI: Benchmarking Multimodal Models for Visual Programming in Turtle Graphics Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2991a7a4-e583-4535-9552-f3ac570c3285 · inbound
FindIt: A Format-Informed Visual Detection Benchmark for Generalist Multimodal LLMs Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 429dbaf1-c476-4e4c-8d6e-99e6ec024e7e · inbound
VoLo: A Physical Orchestrator for Open-Vocabulary Long-Horizon Manipulation Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation dff2304c-45b6-45ef-ba6e-d4833df9e37c · inbound
Mitigating Manifold Departure: Uncertainty-Aware Subspace Rectification for Trustworthy MLLM Decoding Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d41d686e-968f-4083-9174-6917fd592dfe · inbound
Kwai Keye-VL-2.0 Technical Report Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8dc55f05-b5d0-4062-baa6-a9031859f552 · inbound
Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 126
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 171174b7-0028-4b2f-b4ed-77c362ac8346 · inbound
Confidence Calibration for Multimodal LLMs: An Empirical Study through Medical VQA Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 87be315a-80fe-4ba8-a9d9-6b8367d7d947 · inbound
Efficient Visual Pointing for Embodied AI:Agent-Driven Data Synthesis, Cross-Block Attention, and Iterative Correction Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 24835933-4e67-4a74-9600-2c350fc5ca88 · inbound
Program-as-Weights: A Programming Paradigm for Fuzzy Functions Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ca4cabcf-6b18-4f73-b78c-c9855e5d509d · inbound
BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 231
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 505b980e-1fb0-44cb-938a-fcb73327db0e · inbound
MentalThink: Shaping Thoughts in Mental SVG World Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 284
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81e3cb87-82b7-4d68-8baa-9e4f70a9d5ec · inbound
ACE-Brain-0.5: A Unified Embodied Foundational Model for Physical Agentic AI Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 159
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d30f5a49-2cbe-4992-96d7-716a8ed7b9b0 · inbound
LookME: Lookup-Based Multimodal Embeddings for Layer Injection in Vision-Language Models Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef0405c2-9e41-413c-b7e9-71406a9ff105 · inbound
MVEI & EmObserver: Empowering MLLM-Oriented Visual Emotional Intelligence via Emotion Statement Judgement Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ef0d572-a1e8-495c-afcc-558fd2e7a8f6 · inbound
Data Pyramid for Embodied Manipulation Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8dfa9635-ff28-466f-8fe5-ac4da1bda456 · inbound
Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3506be23-9ce8-4aaa-8b1c-b20d03c98b34 · inbound
$\pi\mathbf{R}^2$: Reactive Real-time Flow Policies Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1728ff29-10d0-4f86-96de-62aa7cfc815f · inbound
ParVL: Parallel Scaling and Expandable Compute Allocation for Multimodal LLMs Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.