Pith. sign in

Paper Citation Record · LEDGER

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

As of 5 August 2026, this Paper Citation Record lists 100 of 137 outbound references and 79 inbound Pith citation observations for arXiv:2409.17146.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2409.17146 v2

Coverage vector

measured 100 of 137 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-15T01:55:12.501409Z

measured 179 of 179 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 79 of 79 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T12:28:29.027544Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

100 of 137 outbound references displayed

  • verified exact39
  • verified fuzzy55
  • unresolved2
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch4

External citation measurements

8
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 8d916b44-044f-43ce-999e-1e801b25f4e3 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-15T01:55:12.614106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:38c23b1e65f38a7eec8ee4d292c1962f3cf2ce74936f562d3e7c2e23e88d966d

Observation 4f641411-604c-47e7-b3b8-01b6f4a3ac00 · outbound

This paper cites TallyQA: Answering complex counting questions.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models TallyQA: Answering complex counting questions

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T01:55:12.775441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:6dd1a784d3f9c6377a3ff3d85771b5087b7e552cc648cc13e9a664283d6c5f96

Observation 66afdab2-c2de-4c04-a32c-70269fe46219 · outbound

This paper cites Pixtral 12B.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Pixtral 12B

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-15T01:55:12.555156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:171847e7216aa20868287779f7776b97131b2e74eeeeed7daf0ca038ba47c598

Observation 175d3db8-aadb-45a9-a5b7-84b6d9cb67bd · outbound

This paper cites Yi: Open Foundation Models by 01.AI.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Yi: Open Foundation Models by 01.AI

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T01:55:12.558215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:aec308b6ccd33a48bc6d2106397e62b6b90ab25bc32f902ced7a41891d8cb89a

Observation 8d3e417b-c926-43b2-a6db-3b7376cf41a0 · outbound

This paper cites The Llama 3 Herd of Models.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models The Llama 3 Herd of Models

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-15T01:55:12.561796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:1fcc0cfdb81fbc6b4ad2ca3697e648bafff091cafef74f7b1fb84e18f4f1cb18

Observation c6af7ef7-eff2-4d66-8b95-a607f242baab · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Flamingo: a visual language model for few-shot learning

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T01:55:12.783582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:99bf0f6d7e675ca9befffd8f2b3ecb27ca3d5ad86f23cfc872335497856374fa

Observation 9f6b7d5b-83e7-4505-93f4-18e264455955 · outbound

This paper cites The claude 3 model family: Opus, sonnet, haiku.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models The claude 3 model family: Opus, sonnet, haiku

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T01:55:12.785439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:d59bcd270511ca79b0e377be72c15bd8a07890391f6b9ae280117c635b89a799

Observation e187df9f-bee6-468b-940d-dec9ac9dc948 · outbound

This paper cites Layer normalization.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Layer normalization

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T01:55:12.787144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:6b18300859155076de8c8c0bd6aba0b580453a198f938b9fc6db5c417ad2776e

Observation 27e7d071-5587-4757-8e1c-08f729938d42 · outbound

This paper cites Fuyu-8b: A multimodal architecture for ai agents.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Fuyu-8b: A multimodal architecture for ai agents

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T01:55:12.788966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:152896f9bac3246cdc9f9485de40a430e2ac5b2d691f9b7171171f3babdd2726

Observation aa522005-f79e-433f-adfd-34aef21ec019 · outbound

This paper cites PaliGemma: A versatile 3B VLM for transfer.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models PaliGemma: A versatile 3B VLM for transfer

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-15T01:55:12.564804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:7a8f4b28f0d2eab5e2fff72d889ea8806c10da5471b10452c3eab9ea1e23c727

Observation 8c16d6ae-a6dd-4ee4-a60f-c6ac72f5042f · outbound

This paper cites Scene text visual question answering.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Scene text visual question answering

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T01:55:12.792436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:f07a6cd13f32636a073df620f1535929c8d767a009983100a1343777bd98aa82

Observation 2b1945b4-3e19-4b00-866e-9e9ffca1aefe · outbound

This paper cites Honeybee: Locality-enhanced projector for multimodal llm.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Honeybee: Locality-enhanced projector for multimodal llm

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T01:55:12.794555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:34149815bd7fbc446f259e818ff911d5b5eb0496fffed74912ca0de3f607caba

Observation 1b9acf5c-c89a-4316-912f-b40cbaa05b3a · outbound

This paper cites ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-23T22:20:22.000193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:12eaa0db068d18f858d56ae5e113057c58a648c01d06d785c388f3f1cdb1c046

Observation 0c4271e4-c97b-4810-9b10-48639d805bf4 · outbound

This paper cites EVLM: An Efficient Vision-Language Model for Visual Understanding.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models EVLM: An Efficient Vision-Language Model for Visual Understanding

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-15T01:55:12.572829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:231eb9cc292c0a7868379515b6dc14999b2ec91c5cc4ba832b9636e9949df211

Observation 3d13dbdd-8cfd-4bf8-9514-8fd46658da54 · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-15T01:55:12.576121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:32fc53828b4ebf708b8a0d2e333152e85988267f6b154eaa0d5c3291d8f2fabc

Observation 10afaa3d-1884-4f6b-a3e5-33fb556baf50 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Evaluating Large Language Models Trained on Code

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-15T01:55:12.579503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:f8ccd5eeaf7d0e68ace7ab111cae23078562020445247b453c69034c175e5624

Observation e27abcfe-0a82-4689-8963-6ac137d7c451 · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-15T01:55:12.582208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:8e71f7755f22d18e47072e8a718af6ba59c8f7c89b068ed3a8049f74a780756e

Observation c9085aad-b652-4417-917f-ead2dd30e829 · outbound

This paper cites PaLI-3 Vision Language Models: Smaller, Faster, Stronger.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models PaLI-3 Vision Language Models: Smaller, Faster, Stronger

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-15T01:55:12.585432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:ea7e5feeafe8f2e0a9cedb4ed2dd1df4c2b1fc224b5ad3db62a73b96a351601c

Observation 76eee52b-6244-42d1-9162-682f15aa8c30 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-15T01:55:12.588717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:20334c943088eebc45d75102c741fe648da5dae9671d58a84f08fa59a8d2deb2

Observation b2d6ab35-3a86-4b40-b859-530942728310 · outbound

This paper cites Reproducible scaling laws for contrastive language-image learning.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Reproducible scaling laws for contrastive language-image learning

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T01:55:12.809689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:a2f1ea014ff6ff114999495eddd738cbdb9f8a8f11d0342edf13eb6604a682e9

Observation 989c6a9f-11af-4e98-976c-94a7c3d554ac · outbound

This paper cites Chatbot arena: An open platform for evaluating LLMs by human preference.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Chatbot arena: An open platform for evaluating LLMs by human preference

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T01:55:12.811427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:08c3cae6bcf2491e3b442debd852fd6d76790c7a21878513338904fdb6e38f39

Observation 3a7116c2-28ef-48f8-b48b-af9dfecf4bad · outbound

This paper cites MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-16T16:35:38.347525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:72be9f4f1d06b409b45f5255997cb7d76c407d05a61a8d45ff91551fddcd41a8

Observation 13fcd1d4-5482-468c-a429-22af94e90abe · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-15T01:55:12.595184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:a6a0b448c4035fb2529ba6714f57c4d243c65e962d2b932a0db36bdcd250101e

Observation f65f18fc-2f99-4113-bd09-49470ff73d89 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Training Verifiers to Solve Math Word Problems

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-15T01:55:12.598652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:396ba8d6d54509573d1ba583ff83d29bab140a672cebcb914371def45ab6de5d

Observation 5100c3e0-82f6-4c46-8d65-fe867ba98c06 · outbound

This paper cites On implementing 2d rectangular assignment algo- rithms.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models On implementing 2d rectangular assignment algo- rithms

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T01:55:12.819569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:8e7aa831f766084ce984e34cccf1a1cc4b084defd6920a13a0bf4e60db023526

Observation 51d4a3e9-1270-49c6-b3de-5a379c419fed · outbound

This paper cites an unresolved cited work.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-05-15T01:55:12.821489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:ecea20ad2a574a7f60708428d926709f1957a13a6ff0ffc25a5555fa7d14aa40

Observation 16ae0048-5fc7-49ca-96b3-56406959f01c · outbound

This paper cites NVLM: Open Frontier-Class Multimodal LLMs.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models NVLM: Open Frontier-Class Multimodal LLMs

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-15T01:55:12.601804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:ae17bcaf6880cdce864d01ffbe1dc80827c653a0a741e6951a62ff743cc5055c

Observation 92d86baa-5452-4c60-85bb-c5ede3751c52 · outbound

This paper cites FlashAttention-2: Faster attention with better parallelism and work partitioning.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models FlashAttention-2: Faster attention with better parallelism and work partitioning

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T01:55:12.825179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:ade9fdef12d7da11146c34661d362fd2954fa0eba2cab1ddf902e78051cd6d53

Observation d855a28e-3556-498c-928b-b503bd2928f3 · outbound

This paper cites Fu, Stefano Ermon, Atri Rudra, and Christopher R´e.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Fu, Stefano Ermon, Atri Rudra, and Christopher R´e

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T01:55:12.827016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:cb5dc31a5d121824be182ce8eb4add2f44b28df5881887bfdda5041db1224b9a

Observation 0f776ec2-dd71-4348-8b5c-0f340008bf15 · outbound

This paper cites InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-15T01:55:12.605209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:dc989d2ca56c0b329daa1be083afdf183f8116e77bb0f0392463c683a0f35855

Observation cbc0f833-6031-4516-88c4-a398773fc0e1 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models An image is worth 16x16 words: Transformers for image recognition at scale

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T01:55:12.831454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:7d5d10711d1b7bc099f996d35007bccfe4d8e33c4de92e4db6f2d4ab44ebe69f

Observation 0c9c9db9-68f1-43e9-a6a6-5311c3768a0b · outbound

This paper cites VILA$^2$: VILA Augmented VILA.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models VILA$^2$: VILA Augmented VILA

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-15T01:55:12.608096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:4606adcc1ecd57218f1b6d7071a171b663c7d5dc88939bb5ad0fb288ba3c3937

Observation 320a91aa-1659-43bd-a7df-d3d95a7f6577 · outbound

This paper cites Devise: A deep visual-semantic embedding model.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Devise: A deep visual-semantic embedding model

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T01:55:12.836760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:fa3616411e54459bb4e369934c8ba9cd3161aaace0da4c66f392a0bd5c9e7c71

Observation 5c6a522d-e36f-4c50-bf24-b9a962dfe8a9 · outbound

This paper cites VITA: Towards Open-Source Interactive Omni Multimodal LLM.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models VITA: Towards Open-Source Interactive Omni Multimodal LLM

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-15T01:55:12.611554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:794cd5c4e04e45e67ec0ba9ad760801eaa1e26760be6abcaaeeecd52b66e3404

Observation 6681f19c-668e-4764-9aab-1cf21dff7002 · outbound

This paper cites Scaling Synthetic Data Creation with 1,000,000,000 Personas.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Scaling Synthetic Data Creation with 1,000,000,000 Personas

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-16T00:03:55.809567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:d7bff9a9a23ccc26e62a676e6e48f40b14fe43c4504befe5a28f6beba96b4586

Observation 0a1676c9-5182-4139-90bd-abeaf258e485 · outbound

This paper cites Making the V in VQA matter: Elevating the role of image understanding in visual question answering.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Making the V in VQA matter: Elevating the role of image understanding in visual question answering

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T01:55:12.844532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:67a3d1416bb9feeb24f0dc1089a7898f62792dc540b476196c4a0e946fbd36ea

Observation 363513ea-6dbc-40b0-b13d-08ef640f4f16 · outbound

This paper cites an unresolved cited work.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-05-15T01:55:12.847046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:48c736938fc292ded8c0f71ba3f50cba4afd5565da90ac85b75a8e358088a4bc

Observation b85563b5-15ee-4540-9494-26b27446cbd5 · outbound

This paper cites Measuring massive multitask language understanding.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Measuring massive multitask language understanding

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T01:55:12.849514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:cba0d905b7ed3cc8ac16941c6dd5fab91d9a3f9f4bc2c7467d2b8715944edc63

Observation b418d3cf-72ca-4d4f-8b7b-5937e73adeaa · outbound

This paper cites Mea- suring mathematical problem solving with the math dataset.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Mea- suring mathematical problem solving with the math dataset

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T01:55:12.851870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:0beb625fc98e363745b20e20c9ffb819681c1d71273da0a8d7e358669856105d

Observation 86082a46-60d2-4806-867e-d9c659cf60b6 · outbound

This paper cites Accu- mulated gradient normalization.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Accu- mulated gradient normalization

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T01:55:12.854077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:c20482cf2712907312f9d52eceaa2affc31210c08e40aec9265591862b7a1651

Observation 3ed25625-ffcc-4338-b771-77e891f93c84 · outbound

This paper cites CogVLM2: Visual Language Models for Image and Video Understanding.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models CogVLM2: Visual Language Models for Image and Video Understanding

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-16T20:10:28.051835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:578725febf316029bc66d47e2fea3efa0de8b54692615835263b7d5096c4efe7

Observation bde59e58-be69-461e-a6a5-8f734c3a563e · outbound

This paper cites mplug-docowl 1.5: Unified structure learning for ocr-free document understanding.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models mplug-docowl 1.5: Unified structure learning for ocr-free document understanding

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T01:55:12.859466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:a755d35f659ec1c21e9228ed3e7f721cdc8d0d6c11f110e418eb07c00fe5be29

Observation 235d57ff-4e77-4de5-898e-1407a60d7755 · outbound

This paper cites Scaling up visual and vision-language representation learning with noisy text supervision.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Scaling up visual and vision-language representation learning with noisy text supervision

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T01:55:12.861657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:d1a713bd2d0bc41b0fe4e25eab691e242d6788debd7e5cbe6c4887a2c2788124

Observation 95518f9a-3b1a-4faa-8955-5d1b15b44eb1 · outbound

This paper cites MANTIS: Interleaved Multi-Image Instruction Tuning.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models MANTIS: Interleaved Multi-Image Instruction Tuning

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T01:55:12.620895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:fd90659a8e0ed4bd14c26bfe6dbbf28d32171d66ee2a40b56c27a03dee7d4a37

Observation 18e20941-9771-4b83-aea7-d2db56a4fc80 · outbound

This paper cites A shortest augmenting path algo- rithm for dense and sparse linear assignment problems.Computing.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models A shortest augmenting path algo- rithm for dense and sparse linear assignment problems.Computing

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T01:55:12.865776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:95b1449d0a1995bd572a1ee9b27c5aabe9810ee84bd8f8ce60d11437543c1fad

Observation 69ae4a02-b940-450c-943e-560be1802068 · outbound

This paper cites DVQA: Understanding data visualizations via question answering.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models DVQA: Understanding data visualizations via question answering

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T01:55:12.868532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:67a591fa074c8bbe745e0b7b194f8a601ea80bf959d3de60c3316c724fb27170

Observation 809132a1-3500-48e1-be6e-5059b88d2933 · outbound

This paper cites FigureQA: An Annotated Figure Dataset for Visual Reasoning.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models FigureQA: An Annotated Figure Dataset for Visual Reasoning

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-05-15T01:55:12.623827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:2f386a9f4b3ab6822188e5b6708a6ed221a6c441f88a36bcf8367c7be763190b

Observation 295792de-3296-4a4f-bbc7-0ba24a18c730 · outbound

This paper cites Prismatic VLMs: Investigating the Design Space of Visually-Conditioned Language Models.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Prismatic VLMs: Investigating the Design Space of Visually-Conditioned Language Models

Reference 48

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T01:55:12.626959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:5c69f465910455fa5917f5fc4c98ba2e8b0e435aa0e51775814af58fe089b4a8

Observation 58941f00-4e33-426e-ba9c-28ea8bb3938e · outbound

This paper cites A diagram is worth a dozen images.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models A diagram is worth a dozen images

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T01:55:12.876438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:5d61cdd4b035f4a830ec49ae901e52049f3ed6d6aaadc4cf86731dcd582eee51

Observation f93cecd4-64e1-474e-a157-5ad3b6a8f9ab · outbound

This paper cites Adam: A method for stochastic optimization.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Adam: A method for stochastic optimization

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T01:55:12.878933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:4084a2ab525653cf7a8ed7a099e20e681c0ac4dd1dc4e2ce468a3e5bc287b938

Observation e35b5c95-b6f8-4a6c-a3c7-fc9c18609abd · outbound

This paper cites Berg, Wan-Yen Lo, Piotr Doll´ar, and Ross Girshick.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Berg, Wan-Yen Lo, Piotr Doll´ar, and Ross Girshick

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T01:55:12.881620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:37e2bac0fd177f52a3023b6c167631c60013f36a0fdf3c73958ac011f727379b

Observation 22c41068-de0b-4278-883e-0220f1872a23 · outbound

This paper cites Shamma, Michael S.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Shamma, Michael S

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T01:55:12.883633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:e69b68304b7ed5c8a6489301ae54ca785db5350434030c6c33e00e46d8d0b776

Observation 1b0193d8-2822-4c1f-adfc-bc2c5f70e459 · outbound

This paper cites The open images dataset v4: Unified image classification, object detection, and visual relation- ship detection at scale.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models The open images dataset v4: Unified image classification, object detection, and visual relation- ship detection at scale

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T01:55:12.885894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:515de142d6341269232a1158e04f61f1132f6d314b743ba17600d43fd2f128a0

Observation 33e2b272-aef7-43ed-a3d4-b2fd70c979bb · outbound

This paper cites Building and better understanding vision-language models: insights and future directions.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Building and better understanding vision-language models: insights and future directions

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-15T01:55:12.630563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:0ed83ba913e9a84e78334e7ef76b709ee72b55f0fe0373e5c517bff48c82236b

Observation 88a0a4bb-f973-4747-8366-f7b26dc263b2 · outbound

This paper cites Unlocking the conversion of Web Screenshots into HTML Code with the WebSight Dataset.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Unlocking the conversion of Web Screenshots into HTML Code with the WebSight Dataset

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-15T01:55:12.633620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:aa6a0af2999bdbd6f1064a63305c2267f67d25629755727e9847da24306c119a

Observation 73c0fa3e-d4f1-46eb-bce9-1b34fabd9625 · outbound

This paper cites OtterHD: A High-Resolution Multi-modality Model.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models OtterHD: A High-Resolution Multi-modality Model

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-15T01:55:12.636644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:3b50ca099daffdb7f97a58dde20ad8d86762625c23dd78b6fcf15408f38c866d

Observation b975245f-b26e-4cf2-ae75-f3254cc1188d · outbound

This paper cites MIMIC-IT: Multi-Modal In-Context Instruction Tuning.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models MIMIC-IT: Multi-Modal In-Context Instruction Tuning

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-15T01:55:12.639738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:0c8c143970d2177e553ef59572f6f45b0a322b6c8277338e5efbc4526ea44f57

Observation 733a87e2-4dcf-4f13-88ca-c5e9c83c261b · outbound

This paper cites Otter: A Multi-Modal Model with In-Context Instruction Tuning.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Otter: A Multi-Modal Model with In-Context Instruction Tuning

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:43:48.053680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:913691b261dca8574c256549e7d67fec7fb1fbb34fc41d5f690a3b38cd0f76f9

Observation 0a1fa6ab-bb1e-4d10-8053-2d0f6c1ab2e4 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models LLaVA-OneVision: Easy Visual Task Transfer

Reference 59

Resolution
verified exact
local_arxiv, observed 2026-05-15T01:55:12.646452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:f4454487cdc9f621821890a53ac0857e68481620305aa6c636668369eb97a0cd

Observation 175a6c54-d799-48dd-b88f-a9074bf358cd · outbound

This paper cites Blip: Boot- strapping language-image pre-training for unified vision-language understanding and generation.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Blip: Boot- strapping language-image pre-training for unified vision-language understanding and generation

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T01:55:12.726853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:f91fa40b0c5e30c2eff86d775a5007653f3a2ef35affb7dea112166db7a46fb5

Observation efaa6182-b356-4a6f-a8b5-8cab6baa616d · outbound

This paper cites Covlm: Composing visual entities and relationships in large language models via communicative de- coding.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Covlm: Composing visual entities and relationships in large language models via communicative de- coding

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T01:55:12.728901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:6584a68c52bd1fcf8717b50bc7c81d331a289588835d20e7a365d7c849ec9464

Observation 1f7d1fa6-f636-4c3c-8fce-e625e4fc7b27 · outbound

This paper cites On the Effects of Data Scale on UI Control Agents.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models On the Effects of Data Scale on UI Control Agents

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-15T01:55:12.649806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:cb0d02726ab50e6230c587e60d588db22ce0c3413962133ab7c031bce748ae45

Observation fd9b7035-1361-4cf5-ab62-a68f0c2c147b · outbound

This paper cites Monkey: Image resolution and text label are important things for large multi-modal models.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Monkey: Image resolution and text label are important things for large multi-modal models

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T01:55:12.732507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:10e47d90efa9d8cfe53b1bea388d00ef40129eb5baaf5dc993f471a362557216

Observation 7c4f82bd-232a-4ce2-9930-e2a40a6011cb · outbound

This paper cites MoE-LLaVA: Mixture of Experts for Large Vision-Language Models.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models MoE-LLaVA: Mixture of Experts for Large Vision-Language Models

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-16T02:33:30.459169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:c4c21b7f2ae9e622ec2ce82c5d5f7a7557ac0c1f970cac3204f2b0d948bf1492

Observation ae981ab4-9564-4a22-b692-c3f0b5fc71df · outbound

This paper cites Mi- crosoft coco: Common objects in context.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Mi- crosoft coco: Common objects in context

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T01:55:12.736464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:47392305c0fe80a798eadf272a57f33543cb896b8965899030a5ac8f44b16948

Observation 4cd61f70-2dbe-4925-a00c-1afc98be521d · outbound

This paper cites GRES: Generalized referring expression segmentation.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models GRES: Generalized referring expression segmentation

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T01:55:12.738341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:1ebf5a19526414e26bfa81efa6fa421d96401959593418de4f23054afa144a15

Observation b9c48908-59ac-48cb-9350-c02d431efc71 · outbound

This paper cites SPHINX-x: Scaling data and parameters for a family of multi- modal large language models.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models SPHINX-x: Scaling data and parameters for a family of multi- modal large language models

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T01:55:12.740252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:d201a0666d052472b3f9f5d9f60d1966226a9ce126e3211429114485626e217e

Observation 00907bf3-cd35-4657-b388-5e313e8e44d5 · outbound

This paper cites MMC: Advancing Multimodal Chart Understanding with Large-scale Instruction Tuning.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models MMC: Advancing Multimodal Chart Understanding with Large-scale Instruction Tuning

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-15T01:55:12.656437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:f347c521a837a102f0bebbd517df40de49145521f6f8e48e4195779383336da7

Observation a154ac9f-d09e-4fd9-b233-65d1d867ba7c · outbound

This paper cites Visual instruction tuning.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Visual instruction tuning

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T01:55:12.743742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:4e7d580b3dd285126fe59d812ae6c4df1dc3138b59769b8a17a22ef28ccdef77

Observation 259a9fae-5cc0-40ef-86d4-71f275b4b70d · outbound

This paper cites Improved baselines with visual instruction tuning.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Improved baselines with visual instruction tuning

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T01:55:12.745606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:756b386c48ea574723daad0b22ec52f066edc64a2d8ad55c5463339a0b069a67

Observation 47de1e46-8e0d-4f37-8efe-49916253543d · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Llava-next: Improved reasoning, ocr, and world knowledge

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T01:55:12.747825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:5d577de9e89e808b7db0e3830205a84cad5b47a38018920d40e20d9ca4887c56

Observation 33474c40-71df-42f0-8198-39abc54cb99f · outbound

This paper cites SGDR: Stochastic Gradient Descent with Warm Restarts.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models SGDR: Stochastic Gradient Descent with Warm Restarts

Reference 72

Resolution
verified exact
local_arxiv, observed 2026-05-15T01:55:12.659286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:ec32db65fe0abd43f735c2c714b5d8f54cb95248a25a90055deb314f9f0e2895

Observation 7d8bb6a7-159b-4a39-bb7e-48ec9315cbf4 · outbound

This paper cites Decoupled weight decay regu- larization.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Decoupled weight decay regu- larization

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T01:55:12.751473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:e14e1851106e21a9dd5dc8eb88d4740d4c6ed6446b5eb4b25578f99d009731a7

Observation 81c71248-cc10-452a-9eae-4158779acee1 · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 74

Resolution
verified exact
local_arxiv, observed 2026-05-15T01:55:12.662275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:f152fe7dcfc660b791639e19502526a6e816afaa1d6e57ba138c14b5d4159a29

Observation 2e3ce511-9b32-4b58-b1c2-d783d8aa6441 · outbound

This paper cites Unified-io 2: Scaling autoregressive multimodal models with vision language audio and action.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Unified-io 2: Scaling autoregressive multimodal models with vision language audio and action

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T01:55:12.755069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:430bcc7403210852442ae0c2984e25263607bca7ec5dac148175633dfa50eb11

Observation d7bc6368-5151-44a8-aa86-f1930c8bc96c · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for sci- ence question answering.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Learn to explain: Multimodal reasoning via thought chains for sci- ence question answering

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T01:55:12.756941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:6b628fbcc77bc9bc32a7274c1544382cc7d840a3538921bf57764f7d6361e38c

Observation a8b3a234-af38-4831-92ef-2a68e20ae143 · outbound

This paper cites Dy- namic prompt learning via policy gradient for semi-structured math- ematical reasoning.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Dy- namic prompt learning via policy gradient for semi-structured math- ematical reasoning

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T01:55:12.759195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:b1418bf7c5bdb248cf6972f5fe1800ee1d0630e68f98eb7771795b3dc246577d

Observation 5d625152-6f44-4ac7-b117-d7fc2b6a6040 · outbound

This paper cites MathVista: Evaluating mathematical reasoning of foundation models in visual contexts.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models MathVista: Evaluating mathematical reasoning of foundation models in visual contexts

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T01:55:12.761070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:834598d489403a1584d274913361ed1bcbd4e9a28f4ee86bcbf4c1833c9d956a

Observation 55c7b206-57f1-43f5-a9d1-8e5377195d9d · outbound

This paper cites Cheap and quick: Efficient vision-language in- struction tuning for large language models.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Cheap and quick: Efficient vision-language in- struction tuning for large language models

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T01:55:12.763583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:bd8ceb4c3d971499dd3fb22ed0b547fbdb547006ae6941ff5b34e0cf39ec1923

Observation 9245dc91-9630-459b-b637-553008f44f89 · outbound

This paper cites ExpertQA: Expert-Curated Questions and Attributed Answers.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models ExpertQA: Expert-Curated Questions and Attributed Answers

Reference 80

Resolution
verified exact
arxiv_id, observed 2026-05-15T01:55:12.665306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:97bf5a1f378fd7fa7802165498828d123194faf2702febd1dbd92e81ab84d9fa

Observation b5284158-b423-4a2f-b560-1280095a16aa · outbound

This paper cites OK-VQA: A visual question answering benchmark re- quiring external knowledge.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models OK-VQA: A visual question answering benchmark re- quiring external knowledge

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T01:55:12.767156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:96efcfdc742ab6136bcc2c0cfccaf0d1db7c403349b34275ce48962e3b533a09

Observation 957c4882-cb5c-40eb-8ecb-80e9b9472453 · outbound

This paper cites ChartQA: A benchmark for question answering about charts with visual and logical reasoning.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models ChartQA: A benchmark for question answering about charts with visual and logical reasoning

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T01:55:12.769089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:aa78b8aeb6739542e54e3d158b96f8a6024b9b83edead65c706a83ec3d418684

Observation bf158c1b-e9cc-4f3d-8506-32749ddfa62f · outbound

This paper cites DocVQA: A dataset for VQA on document images.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models DocVQA: A dataset for VQA on document images

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T01:55:12.771707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:8f8745b21b8714143c42a387565fa5f38aa3dac7d9badf2bc59b00a0129cb95a

Observation 3dd108be-bf0c-4aa3-b9a1-4729fa220f9d · outbound

This paper cites InfographicVQA.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models InfographicVQA

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T01:55:12.773612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:1c3a4a3ad46fc649426c0b1e7f6298f0e6463cdb4e171eb1667a5773bfeeb716

Observation db67f8a3-8083-40ea-b8f4-c41282c2bd7a · outbound

This paper cites MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training

Reference 85

Resolution
verified exact
arxiv_id, observed 2026-05-16T04:09:36.761640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:b0f882abfe1103f3115a9a29be54e514ed9106ea2ad5b5dfff6c52dbce0bc8aa

Observation fe9fa82c-70a6-4603-998d-e2ad55fb2439 · outbound

This paper cites PlotQA: Reasoning over scientific plots.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models PlotQA: Reasoning over scientific plots

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T01:55:12.779430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:67bb04cdc3bdf79ab9c398fb7ab59ecd22065351ff75c06a8e284274e2895dd0

Observation 7e25d38f-c1c0-45d2-8339-650a94a7e581 · outbound

This paper cites OLMoE: Open Mixture-of-Experts Language Models.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models OLMoE: Open Mixture-of-Experts Language Models

Reference 87

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T14:43:13.722529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:42dd7a02712b69ea7679c3ef25f0caec51dca132b6cf2ffa3c5a98aea75cc4bb

Observation ab04819f-78a2-43de-af8c-ee4b016baa7a · outbound

This paper cites GPT-4 Technical Report.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models GPT-4 Technical Report

Reference 88

Resolution
verified exact
local_arxiv, observed 2026-05-15T01:55:12.673580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:12995c37262dbabdc2f4b79441f6d107af47dcc855d3e555e4b95de4700d5c0f

Observation 46f43994-70be-4b43-becc-ad381b784019 · outbound

This paper cites GPT-4o mini system card.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models GPT-4o mini system card

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T01:55:12.796338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:cf515d6f9e253ad9829c899492339033d12920765f10acee483335d912af179e

Observation 7d7e216d-8211-46a2-bbbb-d199f3f7dd8a · outbound

This paper cites GPT-4o System Card.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models GPT-4o System Card

Reference 90

Resolution
verified exact
local_arxiv, observed 2026-05-15T01:55:12.676404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:8e8f6713075ef0b605b7951904e4113a4288b6ae2d5d302c16d5b718b4ba99e8

Observation e13f6784-27fc-4f31-a098-5fce1a425133 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models DINOv2: Learning Robust Visual Features without Supervision

Reference 91

Resolution
verified exact
local_arxiv, observed 2026-05-15T01:55:12.679057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:4979ad758e84832e9d586e3f801ed26c26a1a9a5c7fe1b5d770564f7e56b5300

Observation 0b673372-18ef-4d2c-ac6e-5f632337b2b9 · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 92

Resolution
verified exact
local_arxiv, observed 2026-05-15T01:55:12.681581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:62dd1ca3a85263896596ba1559622fffc8adb07330fee014511929bd535f49b3

Observation f7b0621a-0b12-4e6a-945b-0833c445c289 · outbound

This paper cites Connecting vision and language with lo- calized narratives.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Connecting vision and language with lo- calized narratives

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T01:55:12.803801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:b1ba354827f418ba1853bc58b02b7141da7cb8e7e422bd7a81f38a5e399fe6d6

Observation eac89216-e82d-4cc1-bfbe-f2966e4a05f8 · outbound

This paper cites Jack of all tasks, master of many: Designing general- purpose coarse-to-fine vision-language model.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Jack of all tasks, master of many: Designing general- purpose coarse-to-fine vision-language model

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T01:55:12.806020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:c593d254f9fd01d3854c1048da1968a081257a7a13b814dd686e3fe015651922

Observation f24fbc4b-1c53-4e5e-9523-7b9dcdbd4951 · outbound

This paper cites Improving language understanding by generative pre- training.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Improving language understanding by generative pre- training

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T01:55:12.807886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:c86be8c5ec47b30ae3fa77a6d61c41dcdc332dc0b029e75a646b6a5e5def165c

Observation 1472ef67-4680-49b7-a62a-ec13abb25d32 · outbound

This paper cites 28 Learning transferable visual models from natural language supervi- sion.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models 28 Learning transferable visual models from natural language supervi- sion

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T01:55:12.813303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:07d73dd251292775cd1c9b691806cfa5aac04f055a44bf383bafd5877312e161

Observation 8b170e06-f38b-4061-a0ed-ea4481ae3150 · outbound

This paper cites GLaMM: Pixel grounding large multimodal model.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models GLaMM: Pixel grounding large multimodal model

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T01:55:12.815073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:907a235d75835b76aa8e1f59ad933e0c7b5729360512d24e31c3a7ffca0a0d5c

Observation e717be9b-542a-4e2a-a2b6-0b48076ffc46 · outbound

This paper cites LAION-5B: An open large-scale dataset for training next generation image-text models.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models LAION-5B: An open large-scale dataset for training next generation image-text models

Reference 98

Resolution
verified exact
local_arxiv, observed 2026-05-15T01:55:12.684621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:2f7bfc354c099eea7607c2daf1cad8a4b2a7fb5966d710974fabb3614c84b959

Observation 1d538a46-71da-429e-bd0c-eee4f3fb3589 · outbound

This paper cites A-OKVQA: A benchmark for vi- sual question answering using world knowledge.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models A-OKVQA: A benchmark for vi- sual question answering using world knowledge

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T01:55:12.823262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:c5bbadbd2f399a857f199404de2d5a410de9e2f97b632afbcc70384fe76d6b21

Observation b8dae70c-7937-498b-a12e-bfb158ab969e · outbound

This paper cites Towards VQA models that can read.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Towards VQA models that can read

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T01:55:12.829132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:c0c3d5c64ac07e432a8e45ea1a21aead92899ce39668720c05fb5cdd72c8acd2

Pith citing papers

Observation 8af8519b-85eb-4d2b-a7eb-050fdd0a6cad · inbound

Pixtral 12B cites this paper.

Pixtral 12B Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T01:55:12.888993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T23:53:29.862702Z digest=sha256:813ad3982eb5c3c5f0b8636273498341d3561102f159a2869b1f5436c58107f5

Observation 20e8761d-11f8-4f0f-8008-9f08393c93ac · inbound

Self-Correcting Text-to-Video Generation with Misalignment Detection and Localized Refinement cites this paper.

Self-Correcting Text-to-Video Generation with Misalignment Detection and Localized Refinement Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-23T08:25:29.404962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T08:25:01.468957Z digest=sha256:9fe5b79ef834f18bcb1beff9d4c394d54d85ab73a42a322e2d37dc7863caa138

Observation ba0311bd-3799-4d48-89c7-1eed439d2f76 · inbound

PaliGemma 2: A Family of Versatile VLMs for Transfer cites this paper.

PaliGemma 2: A Family of Versatile VLMs for Transfer Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-15T09:15:07.659639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T09:15:07.523565Z digest=sha256:7dd11c8af6b10055653adfbcd7a105fff66f3f95fa2bd1202a7c8494e5075015

Observation 30682552-5cf8-4c71-98dd-c76d6aa030b7 · inbound

Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction cites this paper.

Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 75

Resolution
verified exact
local_arxiv, observed 2026-05-18T04:09:41.566717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T04:09:41.494136Z digest=sha256:cc13e0c7131db8e1c2f27c3271ebc413a463131be5d275f6b935ec831d6231a4

Observation b688c074-fbc0-45fe-a1fa-7b85d8d9a822 · inbound

NVILA: Efficient Frontier Visual Language Models cites this paper.

NVILA: Efficient Frontier Visual Language Models Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-23T07:42:43.122814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:59483094b7c3e99180416cfdf12614f7c23bf606f52bb5774fbee941c12ed3e1

Observation 7286b563-8a32-4138-aa2c-422486401088 · inbound

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling cites this paper.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-15T01:55:12.888993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:962d97dec21b3239d94478b5a4e11191f7e7af1a53fb9ba62b343671b1d30e8b

Observation c81f48ad-3d61-48e7-944c-fa3a6ae5ab3e · inbound

DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding cites this paper.

DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-15T01:55:12.888993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T10:09:21.542356Z digest=sha256:458a563568bc145abd05e84d020e2f7aa6664581a77e578a704e76c53455fc95

Observation e70fe4f2-ee1e-4a49-ba6e-5e76eff24780 · inbound

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning cites this paper.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:33:26.807978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:ca35b4a456e0b0cdd63265082609e6256f6c83dac0e4b70d258ed156f7649613

Observation b1eea010-638d-4c12-84de-b93cf2eb9da3 · inbound

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding cites this paper.

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T01:55:12.888993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:19:59.603343Z digest=sha256:2380ec51d9a4fffdb20197815905fd03da34bc6c52b4027009fccac84b6b61dc

Observation 951b5426-f4ed-4f0a-983e-cc0efc7b2ecd · inbound

Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs cites this paper.

Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-15T01:55:12.888993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T22:22:27.455361Z digest=sha256:bf1b117b7c706480994a6ef1ea47b265695f1a2cc47e1a3ad3b9350eeed3f912

Observation 2f8a1c20-5fed-4c20-bdc4-7856133ce447 · inbound

Seeing is Understanding: Unlocking Causal Attention into Modality-Mutual Attention for Multimodal LLMs cites this paper.

Seeing is Understanding: Unlocking Causal Attention into Modality-Mutual Attention for Multimodal LLMs Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-23T01:25:16.462613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T01:23:01.892612Z digest=sha256:49e7bf2f58d0db263dab189ffc40a85580913bf63b9f0c2982316a482e579340

Observation 516d51f7-c38a-44a3-bc4d-355dbeb60f78 · inbound

Capacity-Aware Inference: Mitigating the Straggler Effect in Mixture of Experts cites this paper.

Capacity-Aware Inference: Mitigating the Straggler Effect in Mixture of Experts Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-23T00:42:19.055339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T00:41:06.928901Z digest=sha256:7fe652678eae30d041b9dfba4394cd403170e3c57d6f8c2e9af55beb5fb8a32a

Observation 0a8b3b20-0fc7-480b-8f5d-bd4d388550d3 · inbound

GuideDog: A Real-World Egocentric Multimodal Dataset for Blind and Low-Vision Accessibility-Aware Guidance cites this paper.

GuideDog: A Real-World Egocentric Multimodal Dataset for Blind and Low-Vision Accessibility-Aware Guidance Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-22T23:52:17.046567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T23:49:53.830731Z digest=sha256:44f7dded3ccdb033fc2ec436aea349a76e56b487055cb1d5374c3dade5b7168a

Observation 5e0975f2-ce4c-4ade-b961-4636735b67c7 · inbound

SmolVLM: Redefining small and efficient multimodal models cites this paper.

SmolVLM: Redefining small and efficient multimodal models Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T01:55:12.888993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T20:23:50.552549Z digest=sha256:d9aac29be4156778f1d5092976acb670f4c964a40754c2bc4c6810c12fc530e9

Observation 83a0b490-906e-40ed-b2b5-23d309d7b10e · inbound

FLARE: Fully Integration of Vision-Language Representations for Deep Cross-Modal Understanding cites this paper.

FLARE: Fully Integration of Vision-Language Representations for Deep Cross-Modal Understanding Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-05-22T19:52:01.951522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T19:49:00.961388Z digest=sha256:53e9324f1fb933aa465538bd2180cb9876c0c35f6853717e6170bd0b28b649a8

Observation e00a2882-d65a-4442-a090-1dfebde268d7 · inbound

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models cites this paper.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-15T01:55:12.888993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:74189455e64b3de37531fbbe2f9562463f0c922bb65947830887d93576f06d9a

Observation c7776a7b-b344-498d-ab3f-2aeff41dcc31 · inbound

Perception Encoder: The best visual embeddings are not at the output of the network cites this paper.

Perception Encoder: The best visual embeddings are not at the output of the network Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T01:55:12.888993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T22:21:15.681336Z digest=sha256:efd4885d61b16de8b489d16d81f42b9906fb7dea9a75b9eefed242fcb5f7e861

Observation 4903c79d-2e7a-4715-ba10-a683ddc771bf · inbound

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization cites this paper.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-22T18:05:00.971979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:abf20f990bad62f0ceb51e543831f5e6bcad91df1e2c0daae054f930d9266e0b

Observation e4e88b71-6c17-4e55-9f22-d64811bd6c3d · inbound

Seed1.5-VL Technical Report cites this paper.

Seed1.5-VL Technical Report Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-15T01:55:12.888993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:e52a1b7e342c07891b0238c65dc3362ae793bbc7560c1efeff86c8a2019a677d

Observation 684ca5b8-d3d0-49be-9698-c82e50151a9e · inbound

FieldWorkArena: Agentic AI Benchmark for Real Field Work Tasks cites this paper.

FieldWorkArena: Agentic AI Benchmark for Real Field Work Tasks Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T14:37:35.827439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T14:36:30.827705Z digest=sha256:7adc864a1a592a8e94fb80a5da5063f604b57d640736634e6a1b22a795c6be10

Observation 228a6080-83c3-463e-8d1a-0a851b2d3eb3 · inbound

Grounded Reinforcement Learning for Visual Reasoning cites this paper.

Grounded Reinforcement Learning for Visual Reasoning Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-22T01:05:52.112924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T01:05:18.801388Z digest=sha256:1d47810acfa63fff945513fbb9c3802c02f9fcdc1f8c6a8b6dbbdf3c8ed3b947

Observation 261929bf-d31d-44c9-9e91-484d0399b0e0 · inbound

Common Inpainted Objects In-N-Out of Context cites this paper.

Common Inpainted Objects In-N-Out of Context Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-19T11:42:15.998907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:38:44.069681Z digest=sha256:daab989bb9145435df074a596bc9fce160795fa66c95e1a8ec5f66e727910172

Observation 003da85c-77e9-43e6-b3b6-1f9d9fa16fb6 · inbound

A Survey on Vision-Language-Action Models: An Action Tokenization Perspective cites this paper.

A Survey on Vision-Language-Action Models: An Action Tokenization Perspective Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 182

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T14:08:35.062910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T14:08:34.893876Z digest=sha256:5451758e3bfb729cce586ecd58b08937f0d136f4f5fb259ebc1c441dd06dc3e3

Observation 7f2e98f7-69b6-4ef1-a6ce-cad52c7a2f7e · inbound

High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning cites this paper.

High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-19T06:12:07.024774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T06:10:57.219445Z digest=sha256:e9d14e2935ceb057fcf68c5f482f637da2f762761ef783d5ba3e1b9adb27810d

Observation 15df2ffb-9ffc-4e60-b1fb-34703a5eb7c9 · inbound

When Seeing Overrides Knowing: Disentangling Knowledge Conflicts in Vision-Language Models cites this paper.

When Seeing Overrides Knowing: Disentangling Knowledge Conflicts in Vision-Language Models Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T03:52:01.496231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T03:51:17.765919Z digest=sha256:723726831cf68a630611aee865438ab205a17a46cafedbd3123284f5f00eba1f

Observation 36458622-c441-4ed3-9332-8a5e2e1554c7 · inbound

Embodied-R1: Reinforced Embodied Reasoning for General Robotic Manipulation cites this paper.

Embodied-R1: Reinforced Embodied Reasoning for General Robotic Manipulation Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-18T22:06:51.660243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T22:04:34.235731Z digest=sha256:c380cb38612fecd7d2d5deb562ea1a0acb46c97e3d5b404485e73e7af0f33a23

Observation 0cce3dfb-280c-4057-949d-81b8e2770e95 · inbound

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency cites this paper.

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-15T01:55:12.888993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T11:58:58.660564Z digest=sha256:7c6615d9e69e4a0b28cf18326bc65d7298a7afaa1556954b44645fa556a46dfa

Observation f696242a-dd65-4674-8dfc-64c84568f856 · inbound

Kwai Keye-VL 1.5 Technical Report cites this paper.

Kwai Keye-VL 1.5 Technical Report Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T12:28:29.027544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:28:29.027544Z digest=sha256:1545730d33418441d68669c1ae4732438f7f432d4a50e323f8645b39b5658f1d

Observation 90da2d0b-414b-43cc-878b-39ed1d204b76 · inbound

Improving Large Vision and Language Models by Learning from a Panel of Peers cites this paper.

Improving Large Vision and Language Models by Learning from a Panel of Peers Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.318382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.318382Z digest=sha256:53c4168b807855da5a6220526e08f45a00625da7c3fd58744490ef4a60746f74

Observation 4026b94f-1a6f-412c-81d3-6e399801e536 · inbound

Reinforced Visual Perception with Tools cites this paper.

Reinforced Visual Perception with Tools Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:04.911262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:27:04.911262Z digest=sha256:42583071c0bbafa8d65f0d374e83d2a9f33b13259cb1e7a0ff4a4656efc84815

Observation dd4df457-7e3a-4575-93f0-db03883c2852 · inbound

LExI: Layer-Adaptive Active Experts for Efficient MoE Model Inference cites this paper.

LExI: Layer-Adaptive Active Experts for Efficient MoE Model Inference Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-05T11:30:12.840402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:30:12.840402Z digest=sha256:128d85924162481de0fd98d39a90914585cd698ee6eb3737da6abc8323823d24

Observation 44068fd0-5849-4e91-bf7d-b33f4cf17e88 · inbound

Weakly-Supervised Learning of Dense Functional Correspondences cites this paper.

Weakly-Supervised Learning of Dense Functional Correspondences Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T10:37:52.209256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:37:52.209256Z digest=sha256:eee26bbb36b23e1c1e75cde13b02d00acd3faf25b13936b848e838c328291ac5

Observation 8d539798-7a22-43b1-a53d-fab53a9356a0 · inbound

Promptception: How Sensitive Are Large Multimodal Models to Prompts? cites this paper.

Promptception: How Sensitive Are Large Multimodal Models to Prompts? Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-05T10:34:18.332552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:34:18.332552Z digest=sha256:71ae5e1657e665b9aeb8223ef35f5da498a6f3aa44d733dd4baf105f9965be0a

Observation 130cf91b-85ff-4fd1-989f-523aa9f82f7d · inbound

Towards Better Dental AI: A Multimodal Benchmark and Instruction Dataset for Panoramic X-ray Analysis cites this paper.

Towards Better Dental AI: A Multimodal Benchmark and Instruction Dataset for Panoramic X-ray Analysis Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T19:27:40.625058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:27:40.625058Z digest=sha256:45d9189f45bd119757532526ec3bb6d8368bed08b0271b2ad24fb29bc23e46e8

Observation 4964d9db-8401-4b26-bb24-e2b2d0c1c663 · inbound

Improving Fungi Prototype Representations for Few-Shot Classification cites this paper.

Improving Fungi Prototype Representations for Few-Shot Classification Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T17:16:45.186493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:16:45.186493Z digest=sha256:f72ec0938b067cab31fc08e025d8422df05763bbe2ec00e93f4c222764d44d9c

Observation 7460dde3-c7a3-4c46-9163-9122caeb58c0 · inbound

InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy cites this paper.

InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-15T01:55:12.888993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T20:09:39.677347Z digest=sha256:608339965cd30d52ae3ef0f217876f8566cccfa350034fa0c306c67b5db17c13

Observation 9e3b4d55-2dda-411b-94b2-20e4b964ecfd · inbound

VisCoder2: Building Multi-Language Visualization Coding Agents cites this paper.

VisCoder2: Building Multi-Language Visualization Coding Agents Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-18T04:15:51.813368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T04:13:59.418835Z digest=sha256:45e74a392c02d6f3dc9d960e78c34e50248e34486654ae89ed6f4c92683451fe

Observation 68d4ffd2-bfd0-4552-8927-d520ca3ab30f · inbound

BOP-ASK: Object-Interaction Reasoning for Vision-Language Models cites this paper.

BOP-ASK: Object-Interaction Reasoning for Vision-Language Models Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:00:10.747887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T19:58:19.309634Z digest=sha256:2042f07f441fdd4fed1b212b7c57a76f102a5fa03db1d1bda9c5da184416fa74

Observation 8250275f-ec65-4132-b714-23b7b4b0cd99 · inbound

Steerable Vision-Language-Action Policies for Embodied Reasoning and Hierarchical Control cites this paper.

Steerable Vision-Language-Action Policies for Embodied Reasoning and Hierarchical Control Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 49

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T22:06:42.948643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T22:05:39.797848Z digest=sha256:ae630ab3d657c7b857d4db4fe65dacb6067989d8684ffbf8fb42d104f0770226

Observation ec8ba832-38b1-4df0-ae51-2522ed3c6018 · inbound

Visual Para-Thinker: Divide-and-Conquer Reasoning for Visual Comprehension cites this paper.

Visual Para-Thinker: Divide-and-Conquer Reasoning for Visual Comprehension Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-16T03:30:33.058511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T03:27:53.694506Z digest=sha256:bd7b54277aa5e5b2ec83d4dc0238eb1b5fa2cc37e0e68ef89a2f9803e4b5d3ac

Observation 4710ed8f-9d25-446f-9330-8a9faee53e69 · inbound

TrajTok: Learning Trajectory Tokens enables better Video Understanding cites this paper.

TrajTok: Learning Trajectory Tokens enables better Video Understanding Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-15T19:16:31.831336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T19:11:52.694778Z digest=sha256:500b54c647910f73dcfe1e857b864829dbfd202c2215a23a7c4d363db46529b0

Observation 8cdd0965-97dd-46f2-9f63-c90efe162aaf · inbound

TrajTok: Learning Trajectory Tokens enables better Video Understanding cites this paper.

TrajTok: Learning Trajectory Tokens enables better Video Understanding Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T20:38:37.346586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:38:37.346586Z digest=sha256:4ab8dab9ffd7f8fe4522db2516f6756b878e7e0cfba4814e99ce556d143fe07b

Observation da3b6ecc-b7da-44d6-93c5-a7b33b8e36b0 · inbound

TiPToP: A Modular Open-Vocabulary Robot Manipulation System That Plans cites this paper.

TiPToP: A Modular Open-Vocabulary Robot Manipulation System That Plans Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T18:31:40.819754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:31:40.819754Z digest=sha256:ba6fd45f219dadff2ccee92b9457bc78cb88ceeba3fe8bfa61ed0960e4f2f5d4

Observation 3a115856-c175-4260-8ee8-46c105ffe1f3 · inbound

Evaluating Vision Foundation Models for Pixel and Object Classification in Microscopy cites this paper.

Evaluating Vision Foundation Models for Pixel and Object Classification in Microscopy Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T17:52:42.744282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T17:52:42.744282Z digest=sha256:e0f0aa528967d3d073beb4565a5238fa820e5597bf1e79ea7adc9389e08f2be9

Observation 271ab0f9-1f78-4a54-88dc-b679a7b170dd · inbound

A1: A Fully Transparent Open-Source, Adaptive and Efficient Truncated Vision-Language-Action Model cites this paper.

A1: A Fully Transparent Open-Source, Adaptive and Efficient Truncated Vision-Language-Action Model Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T01:55:12.888993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:31:23.255452Z digest=sha256:dcffec5d99360078f393f8c60fdbeda99d1c6fac5154c73a92b52118961de429

Observation 211cac15-8573-4f9a-b6de-edb2618fc9f0 · inbound

ABMAMBA: Multimodal Large Language Model with Aligned Hierarchical Bidirectional Scan for Efficient Video Captioning cites this paper.

ABMAMBA: Multimodal Large Language Model with Aligned Hierarchical Bidirectional Scan for Efficient Video Captioning Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T01:55:12.888993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T17:55:12.551127Z digest=sha256:b048ae52be8e1b966cc6099f7516407bc94d9dfedfc0223891c1c7cc3f823706

Observation 4828e8f5-3a9c-4eac-a049-45900830e3db · inbound

Entropy-Gradient Grounding: Training-Free Evidence Retrieval in Vision-Language Models cites this paper.

Entropy-Gradient Grounding: Training-Free Evidence Retrieval in Vision-Language Models Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T01:55:12.888993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T17:14:11.941977Z digest=sha256:e70a44ec4678954c01d106385b784261e2af4eaee73693cf085ef2d732578922

Observation 5e4c86e9-9082-4271-9169-156300307c1c · inbound

Anthropogenic Regional Adaptation in Multimodal Vision-Language Model cites this paper.

Anthropogenic Regional Adaptation in Multimodal Vision-Language Model Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-15T01:55:12.888993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T16:29:12.064221Z digest=sha256:79bc6b0432498a9699d4eb84fabc79166b943d95e264485f69af62d66505e8b3

Observation cb95ea42-25b9-47d4-87b4-59507d6e45a9 · inbound

UniMesh: Unifying 3D Mesh Understanding and Generation cites this paper.

UniMesh: Unifying 3D Mesh Understanding and Generation Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T01:55:12.888993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:55:42.679323Z digest=sha256:c02e7e31486887ebf52dcd03f7c9885804e3e7fe689db8e4f685980df0ba1dc4

Observation 10fe9bf2-e95c-4301-af01-7f0d9f1eacf2 · inbound

PRTS: A Primitive Reasoning and Tasking System via Contrastive Representations cites this paper.

PRTS: A Primitive Reasoning and Tasking System via Contrastive Representations Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T01:55:12.888993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T08:56:32.164424Z digest=sha256:5b8071c38a182fcebb786a57831d1f02f0c482a0087fed01864a5dc8450ea22e

Observation 65f28d80-14f7-4cae-a642-f088e0e49f94 · inbound

Visibility-Aware Mobile Grasping in Dynamic Environments cites this paper.

Visibility-Aware Mobile Grasping in Dynamic Environments Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-15T01:55:12.888993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T18:01:30.332804Z digest=sha256:c13eafbed0831ac0eec1edbdeb36535f62bcb071f0252a49991d0c46cc108ad5

Observation 298e1376-4c87-480c-92b1-cf60ddfbc751 · inbound

Visibility-Aware Mobile Grasping in Dynamic Environments cites this paper.

Visibility-Aware Mobile Grasping in Dynamic Environments Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-15T01:55:12.888993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T02:28:09.146674Z digest=sha256:b8c00ce186e2fbf615186616a067bbe359be5f3415649c0ae253abff9a28076c

Observation c1a3f6cf-72a5-4c82-bce2-7e8d88395cb3 · inbound

When Relations Break: Analyzing Relation Hallucination in Vision-Language Model Under Rotation and Noise cites this paper.

When Relations Break: Analyzing Relation Hallucination in Vision-Language Model Under Rotation and Noise Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T01:55:12.888993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T17:40:14.089081Z digest=sha256:f4a630aa11ea3f234da2268d6f9178996adac69fc76d8c93a12d8e1e3d241a8c

Observation ee9e83c5-b2d8-4f7a-8f53-39c69e9ad358 · inbound

When Relations Break: Analyzing Relation Hallucination in Vision-Language Model Under Rotation and Noise cites this paper.

When Relations Break: Analyzing Relation Hallucination in Vision-Language Model Under Rotation and Noise Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T01:55:12.888993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T03:52:44.531016Z digest=sha256:d1a338b5c546ec7b983c2b2b3cb0db3508fab3f3f24edaf7fa42955559bb24fb

Observation 605ca1b7-7702-4d2d-af34-4d419a2969c6 · inbound

20/20 Vision Language Models: A Prescription for Better VLMs through Data Curation Alone cites this paper.

20/20 Vision Language Models: A Prescription for Better VLMs through Data Curation Alone Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-15T01:55:12.888993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T02:52:43.674969Z digest=sha256:67c64140804af0532c0bf03e3b1194f965c3d1730591c171e77f669d02f21d3c

Observation 99da1db6-1974-4923-912d-a95c4bcd97d1 · inbound

20/20 Vision Language Models: A Prescription for Better VLMs through Data Curation Alone cites this paper.

20/20 Vision Language Models: A Prescription for Better VLMs through Data Curation Alone Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-15T01:55:12.888993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T21:28:37.680681Z digest=sha256:5196aed61978275e0b87d27998d010eab0330a1e9848557beb88324939730801

Observation f2b47308-92f2-48b3-a1bb-140b9b36d1ea · inbound

Beyond Waypoints: Dual-Heatmap Grounding for Cross-Embodiment Semantic Navigation cites this paper.

Beyond Waypoints: Dual-Heatmap Grounding for Cross-Embodiment Semantic Navigation Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-20T05:58:05.388019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T05:53:31.627150Z digest=sha256:a5b4fac53a45f646deae8ef98787a75d9a816cccf91b031abec121243f21a004

Observation 3754e66a-e946-447b-a587-ed2e7a417482 · inbound

Binding Visual Features Point by Point cites this paper.

Binding Visual Features Point by Point Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-06-29T23:44:03.290485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T22:56:44.793896Z digest=sha256:34426606eef99c62d926f58f4b8de8ef986b1c58efd56d6d8b19fdde5cbcd3c4

Observation 09a2392a-73ee-43b9-86d8-e8694f9d8ab5 · inbound

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence cites this paper.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T22:13:59.623529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:1f16f5ef7fa11b2085f6f6daf0c174a0d058da38f6c373d385c6b7f42461dab4

Observation 0b7fb262-8dd6-44cb-ab57-d65d0ccf9d8e · inbound

VCap: Hypergeometric Rewards for Weak-to-Strong Visual Captioning cites this paper.

VCap: Hypergeometric Rewards for Weak-to-Strong Visual Captioning Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 12

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T13:23:28.479602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T13:13:57.599970Z digest=sha256:abed86289ef273b3626945d8d181d989b9be506652ec2a766730ef875facb20e

Observation 8338fa87-ef7c-4a6a-a574-7e2476c4d907 · inbound

OpenEAI-Platform: An Open-source Embodied Artificial Intelligence Hardware-Software Unified Platform cites this paper.

OpenEAI-Platform: An Open-source Embodied Artificial Intelligence Hardware-Software Unified Platform Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 25

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T03:56:35.248704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T09:31:56.289129Z digest=sha256:50a77414eb042a6f914f19556b6a7cc46259347ce3f4732a343246db7c8b1479

Observation 186f8c83-491b-466e-a3c9-99a916e9a67c · inbound

TurtleAI: Benchmarking Multimodal Models for Visual Programming in Turtle Graphics cites this paper.

TurtleAI: Benchmarking Multimodal Models for Visual Programming in Turtle Graphics Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 2

Resolution
malformed identifier
local_arxiv, observed 2026-07-02T02:56:29.434215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T10:28:00.090116Z digest=sha256:c6e0cb49c9375f34d2969f7777d3bfbdc73f9348dadff5a2cd8190b2d1da4ce7

Observation 2991a7a4-e583-4535-9552-f3ac570c3285 · inbound

FindIt: A Format-Informed Visual Detection Benchmark for Generalist Multimodal LLMs cites this paper.

FindIt: A Format-Informed Visual Detection Benchmark for Generalist Multimodal LLMs Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-07-02T03:06:30.404545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T10:16:53.915716Z digest=sha256:4d3471993b08824c05feed7d6452aab74cf5c020daa1d2ec0c9a9c41387343d9

Observation 429dbaf1-c476-4e4c-8d6e-99e6ec024e7e · inbound

VoLo: A Physical Orchestrator for Open-Vocabulary Long-Horizon Manipulation cites this paper.

VoLo: A Physical Orchestrator for Open-Vocabulary Long-Horizon Manipulation Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 12

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T19:07:18.187904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T21:40:00.330510Z digest=sha256:f776cc790d755cd8f5a528aa5b0ba65a403c3ab2b75c5141c384ac03f73ba0d2

Observation dff2304c-45b6-45ef-ba6e-d4833df9e37c · inbound

Mitigating Manifold Departure: Uncertainty-Aware Subspace Rectification for Trustworthy MLLM Decoding cites this paper.

Mitigating Manifold Departure: Uncertainty-Aware Subspace Rectification for Trustworthy MLLM Decoding Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 43

Resolution
metadata mismatch
local_arxiv, observed 2026-06-28T17:52:27.179206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-28T17:51:38.268304Z digest=sha256:dcb2e9b640c1be2fc65bd4dd6fa885dadc07397c33cbe965b8d55942f3b34fec

Observation d41d686e-968f-4083-9174-6917fd592dfe · inbound

Kwai Keye-VL-2.0 Technical Report cites this paper.

Kwai Keye-VL-2.0 Technical Report Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 47

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T04:27:37.148270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-27T13:53:10.352603Z digest=sha256:9a396cc028a2279fb3687208d7fc25d2f44add7b3cd76632756d1f26ba6b82dd

Observation 8dc55f05-b5d0-4062-baa6-a9031859f552 · inbound

Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients cites this paper.

Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 126

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:48:56.137353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T01:08:52.981296Z digest=sha256:13c146ad2059c1f1b7de1954901b6158d1d96b518d286901d514bdc66b335c9f

Observation 171174b7-0028-4b2f-b4ed-77c362ac8346 · inbound

Confidence Calibration for Multimodal LLMs: An Empirical Study through Medical VQA cites this paper.

Confidence Calibration for Multimodal LLMs: An Empirical Study through Medical VQA Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T03:39:30.305306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T17:49:38.151224Z digest=sha256:3231c78baa37a8ae9a30d89b8196eb456025ec85d5c2dd3ac4ddafd8ae0f87ec

Observation 87be315a-80fe-4ba8-a9d9-6b8367d7d947 · inbound

Efficient Visual Pointing for Embodied AI:Agent-Driven Data Synthesis, Cross-Block Attention, and Iterative Correction cites this paper.

Efficient Visual Pointing for Embodied AI:Agent-Driven Data Synthesis, Cross-Block Attention, and Iterative Correction Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-06-30T06:44:18.638769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T06:43:39.125052Z digest=sha256:00e58e558af813b7a53383de4cb92bf3d50f0bf77bbfb17e43ed15387850853a

Observation 24835933-4e67-4a74-9600-2c350fc5ca88 · inbound

Program-as-Weights: A Programming Paradigm for Fuzzy Functions cites this paper.

Program-as-Weights: A Programming Paradigm for Fuzzy Functions Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 43

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T16:18:37.231396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-07-03T16:17:59.522449Z digest=sha256:fd8d1bd9b5ca9179c4eaef22d9f36bf8834f540b45938eaceb840b47d863f391

Observation ca4cabcf-6b18-4f73-b78c-c9855e5d509d · inbound

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception cites this paper.

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 231

Resolution
unresolved
no resolver link, observed 2026-07-12T04:17:40.198357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T04:17:40.198357Z digest=sha256:e12cca8503062d1a95426bf27b727b1cec2dd23afa29ce839b93b532d3f091a3

Observation 505b980e-1fb0-44cb-938a-fcb73327db0e · inbound

MentalThink: Shaping Thoughts in Mental SVG World cites this paper.

MentalThink: Shaping Thoughts in Mental SVG World Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 284

Resolution
unresolved
no resolver link, observed 2026-07-12T01:50:59.184754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T01:50:59.184754Z digest=sha256:f534e62fca5cd59ff7505d3b4dc244b53eacb9b08181abce52fc8946cdffc3c2

Observation 81e3cb87-82b7-4d68-8baa-9e4f70a9d5ec · inbound

ACE-Brain-0.5: A Unified Embodied Foundational Model for Physical Agentic AI cites this paper.

ACE-Brain-0.5: A Unified Embodied Foundational Model for Physical Agentic AI Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 159

Resolution
unresolved
no resolver link, observed 2026-07-11T19:16:57.396710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:16:57.396710Z digest=sha256:899b881df2f7290bcb1fe74e84a0be55681bc1380c1c97953b584ddbbbf864e1

Observation d30f5a49-2cbe-4992-96d7-716a8ed7b9b0 · inbound

LookME: Lookup-Based Multimodal Embeddings for Layer Injection in Vision-Language Models cites this paper.

LookME: Lookup-Based Multimodal Embeddings for Layer Injection in Vision-Language Models Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T06:33:14.240999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:33:14.240999Z digest=sha256:ae13becb0de951ee1efc9eb695c648c4ee3823e47a2cd058b76aae6eb7adfb42

Observation ef0405c2-9e41-413c-b7e9-71406a9ff105 · inbound

MVEI & EmObserver: Empowering MLLM-Oriented Visual Emotional Intelligence via Emotion Statement Judgement cites this paper.

MVEI & EmObserver: Empowering MLLM-Oriented Visual Emotional Intelligence via Emotion Statement Judgement Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-01T08:37:42.483019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:37:42.483019Z digest=sha256:c34ed67c9e85d5cd3412d0a0a7a4d225a5409e5d5c6d85bd312877d465060c83

Observation 5ef0d572-a1e8-495c-afcc-558fd2e7a8f6 · inbound

Data Pyramid for Embodied Manipulation cites this paper.

Data Pyramid for Embodied Manipulation Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 87

Resolution
unresolved
no resolver link, observed 2026-07-31T06:18:55.420407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:18:55.420407Z digest=sha256:61c9e76a73ab5d6261a05ccda8ce5e71cc4aba31390849a415eab60f64a98225

Observation 8dfa9635-ff28-466f-8fe5-ac4da1bda456 · inbound

Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model cites this paper.

Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-07-31T06:20:13.807875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:20:13.807875Z digest=sha256:bca5946b38a5010df9d2e100048ef7fd0f6db3a516e5ec804f2814172788e4ff

Observation 3506be23-9ce8-4aaa-8b1c-b20d03c98b34 · inbound

$\pi\mathbf{R}^2$: Reactive Real-time Flow Policies cites this paper.

$\pi\mathbf{R}^2$: Reactive Real-time Flow Policies Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T00:49:33.132703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:49:33.132703Z digest=sha256:8e53dde9e93cee3a0aca02f42a194edfc754dcef40ab89848d8407c1dd261c51

Observation 1728ff29-10d0-4f86-96de-62aa7cfc815f · inbound

ParVL: Parallel Scaling and Expandable Compute Allocation for Multimodal LLMs cites this paper.

ParVL: Parallel Scaling and Expandable Compute Allocation for Multimodal LLMs Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-05T04:16:08.346639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:16:08.346639Z digest=sha256:72d348c8455b49e3ff9622d10f22f580f2272bab82c3f90a26bc652405d3e0d6