Pith. sign in

Paper Citation Record · LEDGER

DataComp-VLM: Improved Open Datasets for Vision-Language Models

As of 9 August 2026, this Paper Citation Record lists 100 of 299 outbound references and 1 inbound Pith citation observation for arXiv:2606.28551.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.28551 v2

Coverage vector

measured 100 of 299 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-02T21:10:10.548489Z

measured 101 of 101 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T04:09:49.764855Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 299 outbound references displayed

  • verified exact49
  • verified fuzzy24
  • unresolved25
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6e93af41-3080-4578-a7da-f416cd7b35bc · outbound

This paper cites SemDeDup: Data-efficient learning at web-scale through semantic deduplication.

DataComp-VLM: Improved Open Datasets for Vision-Language Models SemDeDup: Data-efficient learning at web-scale through semantic deduplication

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-02T21:17:24.184389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:c574f855ed696edccf36ce15efae8e2b2a3dc8420291c97400b181969cba2b83

Observation 1dbb84d9-7d5f-4521-8f1f-2e2488615cc7 · outbound

This paper cites Effective pruning of web-scale datasets based on complexity of concept clusters.

DataComp-VLM: Improved Open Datasets for Vision-Language Models Effective pruning of web-scale datasets based on complexity of concept clusters

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:17:24.196571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:bf685b71d417fb9fd5c40f2c15613103948b2cbdbdc9657911ea9731be9f25a6

Observation 92b8b338-fd0c-46d3-950c-39ef8e11c2d6 · outbound

This paper cites Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs.

DataComp-VLM: Improved Open Datasets for Vision-Language Models Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-02T21:17:24.209083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:244f910e61800089a8cb6d72f23a98c388f7e4dd528efb716d604250eb5b4ddd

Observation c9721fb6-9e1c-4c25-8c11-06dcdae83067 · outbound

This paper cites Acharya, K.

DataComp-VLM: Improved Open Datasets for Vision-Language Models Acharya, K

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T20:01:24.988261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:569f7bc904dbd95693f97cc844a861d5bc2b2252383371a61bccb57f5aecb0d9

Observation a19966a5-5ef1-4531-aff7-7c132b2c0fc7 · outbound

This paper cites Agnolucci, L.

DataComp-VLM: Improved Open Datasets for Vision-Language Models Agnolucci, L

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T20:01:25.003818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:b4dfdf62203959bcae6c7ddf7db664b3935de9a757b2ca80c8ce22e87820d4f9

Observation 35187fac-7be8-451f-b041-fcd3754ea30a · outbound

This paper cites Ainslie, J.

DataComp-VLM: Improved Open Datasets for Vision-Language Models Ainslie, J

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T20:01:24.858093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:cf7b1c36643fd355117b5caefce0846a576e6db00bbad364772a19d90f2ab62e

Observation 7aa4936f-a048-4046-9ebe-68b1217a0253 · outbound

This paper cites arXiv preprint arXiv:2510.03264 , year=.

DataComp-VLM: Improved Open Datasets for Vision-Language Models arXiv preprint arXiv:2510.03264 , year=

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:17:24.218774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:a44a895bd266ab1bf5c75a6e2de06b98a8392894b7b43f166c67a041ad9a90cf

Observation 88d5015f-3ba6-492c-af90-442eefdc1846 · outbound

This paper cites Alayrac, J.

DataComp-VLM: Improved Open Datasets for Vision-Language Models Alayrac, J

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T20:01:24.927086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:397fbebb315d68c2a21de95ccee60911356d622d73da70b82c88bcf1c6ac0050

Observation d7c0c30d-27c1-439b-a705-0de9b654387c · outbound

This paper cites SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model.

DataComp-VLM: Improved Open Datasets for Vision-Language Models SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-07-02T21:17:24.196896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:ec55121e9cbb85250b9855e579dd5b9ef8226d51054a90dc2e279390e5ad9841

Observation 760d78d4-c1c1-4d09-a3b5-52df8b41f577 · outbound

This paper cites Physics of Language Models: Part 3.1, Knowledge Storage and Extraction.

DataComp-VLM: Improved Open Datasets for Vision-Language Models Physics of Language Models: Part 3.1, Knowledge Storage and Extraction

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:17:24.222155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:548b0c9f81c3df4f0043f4d7b5a0ce63dcd0d45be1450789efa3c85e15c467ab

Observation 57810d9f-af30-46be-8290-53a4d261a0c9 · outbound

This paper cites MathQA: Towards interpretable math word problem solving with operation-based formalisms.

DataComp-VLM: Improved Open Datasets for Vision-Language Models MathQA: Towards interpretable math word problem solving with operation-based formalisms

Reference 11

Resolution
verified exact
doi, observed 2026-07-02T21:17:23.589184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:f2cf0794fcfa68bd2a0901e9754f315098d156c8289ad123984bf803ff389f12

Observation 9b976148-2360-4601-8c58-cad5225c4ca2 · outbound

This paper cites LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training.

DataComp-VLM: Improved Open Datasets for Vision-Language Models LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-07-02T21:17:24.206224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:f9faecca950206743222c95f62f46a0db72ddaf9d818143bee877eec4c60ac1a

Observation dd19a761-50a1-4700-83c2-6c2c5b3407c9 · outbound

This paper cites Perplexed by Perplexity: Perplexity-Based Data Pruning With Small Reference Models.

DataComp-VLM: Improved Open Datasets for Vision-Language Models Perplexed by Perplexity: Perplexity-Based Data Pruning With Small Reference Models

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:17:24.199460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:8de6edc6672610b947f0e26a532d604619d289fe887d55b7ca443c6fa0598f5b

Observation 4a2ed4ba-a2f3-4e81-aa63-e2713f101d7d · outbound

This paper cites Awadalla, L.

DataComp-VLM: Improved Open Datasets for Vision-Language Models Awadalla, L

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T20:01:24.810010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:98d2deb68484eafd8baad08037f92fbfb61a1a22d504a342c245e3feab82f91b

Observation 84b509c1-55a7-482a-baae-eb625d641538 · outbound

This paper cites Louis Bethune, David Grangier, Dan Busbridge, Eleonora Gualdoni, Marco Cuturi, and Pierre Ablin.

DataComp-VLM: Improved Open Datasets for Vision-Language Models Louis Bethune, David Grangier, Dan Busbridge, Eleonora Gualdoni, Marco Cuturi, and Pierre Ablin

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:17:24.212191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:f9a27919883064c2e9798d8941e09fa43fe267b1cffe42126f8caf041ce0c1cb

Observation 60179341-d8db-4ec6-aee5-708dfd95b1c7 · outbound

This paper cites Qwen3-VL Technical Report.

DataComp-VLM: Improved Open Datasets for Vision-Language Models Qwen3-VL Technical Report

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-02T21:17:24.225217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:8f4f2b6e724a2417ac88e7c036e6ac2209f6040cfb28a5a600a26a1e75391ec9

Observation 417e579a-6ac5-4336-bdbb-15e06ac79014 · outbound

This paper cites Qwen2.5-VL Technical Report.

DataComp-VLM: Improved Open Datasets for Vision-Language Models Qwen2.5-VL Technical Report

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-07-02T21:17:24.227419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:8742a6c2a12c21705b0f47924114611c57c19a87a87664abc46358ce907313a9

Observation 661780f2-b57d-4473-b90a-2263d4b3b875 · outbound

This paper cites Berasi, M.

DataComp-VLM: Improved Open Datasets for Vision-Language Models Berasi, M

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:17:24.224911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:31a9d864e41d9f777b630ccc91318c1a7b7d204d35f37ceae1e75747cb3b2870

Observation 786015be-ffec-452d-a7ff-8e5cb19e3c93 · outbound

This paper cites Le Khac, Sanath Narayan, Wamiq Reyaz Para, and Ankit Singh.

DataComp-VLM: Improved Open Datasets for Vision-Language Models Le Khac, Sanath Narayan, Wamiq Reyaz Para, and Ankit Singh

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:17:24.228358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:56ff95d043e33c2f0518cac0990d156560f521c7747374fac1b6a274e3e37981

Observation 34fdb3dd-ed02-4667-b007-9dd9c774bfc9 · outbound

This paper cites an unresolved cited work.

DataComp-VLM: Improved Open Datasets for Vision-Language Models Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-07-05T20:01:24.881034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:dc31253e28bdf654fdf78b70ae627c95d926a575d72ad78a03403d825ac828ed

Observation 919dbdc0-bc6b-416f-b187-6225b4a60a50 · outbound

This paper cites PaliGemma: A versatile 3B VLM for transfer.

DataComp-VLM: Improved Open Datasets for Vision-Language Models PaliGemma: A versatile 3B VLM for transfer

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-07-02T21:17:24.231016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:c60add9c81f27cee4a85fc043153cecab3061195f75589c44fb6150e3daaaf18

Observation f0cb20fc-e2e4-4968-9001-a79d08d1197c · outbound

This paper cites an unresolved cited work.

DataComp-VLM: Improved Open Datasets for Vision-Language Models Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-07-05T20:01:24.799342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:ec404049dd6a9f8fafcf1959c060a2932d1a404e22af0a2ee0f5aaad344e4b68

Observation 8669fb0f-409e-4772-96fc-699582c30a03 · outbound

This paper cites How Much Can We Forget about Data Contamination?.

DataComp-VLM: Improved Open Datasets for Vision-Language Models How Much Can We Forget about Data Contamination?

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:17:24.199652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:68ed1cf7b2064c537991bbeee2a80a8de7403c04fb24d85a51ebd832dc893597

Observation 70f1263d-1545-4d23-bf3c-b86c54ab10b5 · outbound

This paper cites Breuel and WebDataset Contributors.

DataComp-VLM: Improved Open Datasets for Vision-Language Models Breuel and WebDataset Contributors

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T20:01:24.797685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:be64de0721a29d22335ed97770b06e68360dcee0ce61f9fdd2c7fbae685af7c1

Observation e1c4e277-2a1f-49a4-96a5-4a98cd2c6f33 · outbound

This paper cites an unresolved cited work.

DataComp-VLM: Improved Open Datasets for Vision-Language Models Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-07-05T20:01:24.882725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:f2c59474b270b80b1e0f7ec2744769d42a9aeddff88f20e24a51fbbadc8515c5

Observation 27096540-c382-4909-a9a6-c09bfda33a23 · outbound

This paper cites Cahyawijaya, H.

DataComp-VLM: Improved Open Datasets for Vision-Language Models Cahyawijaya, H

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T20:01:24.840995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:012e6142cccbe96edee1f601a21bf2c66081be064fd4d4b017d6cadc86f05eab

Observation b28904d8-093d-42e9-b5b0-18326859e04a · outbound

This paper cites Cao and J.

DataComp-VLM: Improved Open Datasets for Vision-Language Models Cao and J

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T20:01:24.877586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:3d79101351b5ca13c65262f814a929246b4ae252a641dd2c89a38922267edb88

Observation 9656a82b-aa32-4516-b14c-beec3e55b167 · outbound

This paper cites Carlini, D.

DataComp-VLM: Improved Open Datasets for Vision-Language Models Carlini, D

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T20:01:24.847679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:279289c082627eeadb762cdaab7dd2e1ad5a16b862e628c26281952de06d6b39

Observation bf1fb155-8ee7-4788-8d64-4f9e02f876a7 · outbound

This paper cites an unresolved cited work.

DataComp-VLM: Improved Open Datasets for Vision-Language Models Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-07-05T20:01:24.995342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:7b88dcdf4bd4e4c82f2da97bb6d8ce8b1cca972456c07559e318a3deb8c17176

Observation a9a59111-5442-4dce-a9fe-cf8faa980881 · outbound

This paper cites MapQA: A Dataset for Question Answering on Choropleth Maps.

DataComp-VLM: Improved Open Datasets for Vision-Language Models MapQA: A Dataset for Question Answering on Choropleth Maps

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:17:24.222935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:e88051e584c5ce4e4202343d5a63c5f20f3dfe1cfd6bd5ea4fbb2c495b6132de

Observation fb48b3f6-02de-4cb2-bfad-18fec1c71b60 · outbound

This paper cites Changpinyo, P.

DataComp-VLM: Improved Open Datasets for Vision-Language Models Changpinyo, P

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T20:01:24.851123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:d9de730ffc11a6db10a7c31164a8c8a159f2b39ba467aaa0f7496fd4f642d165

Observation 6d30b0ad-a591-435d-953b-7ea19ac1832f · outbound

This paper cites ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models.

DataComp-VLM: Improved Open Datasets for Vision-Language Models ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-07-02T21:17:24.218321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:b05c44d7dc85bdae9d6bfc7a22359fe91630c3b00e61f46fac4dbb1704a43fcf

Observation cb59312b-568b-4db8-a103-08933619a761 · outbound

This paper cites an unresolved cited work.

DataComp-VLM: Improved Open Datasets for Vision-Language Models Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-07-05T20:01:25.029700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:ca12b12608c1151694502651a001a2ccb4a9f86f741bccbe16c52a30e241d968

Observation beef3e49-aa6a-41c9-856e-d18e04d87460 · outbound

This paper cites an unresolved cited work.

DataComp-VLM: Improved Open Datasets for Vision-Language Models Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-07-05T20:01:24.962838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:e3185e6b35cdc99669c6b3ef038312c35ca90ad8e9005c2a5813a775050ef1ea

Observation 5323987d-1522-4440-8f20-ef3705b929f9 · outbound

This paper cites an unresolved cited work.

DataComp-VLM: Improved Open Datasets for Vision-Language Models Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-07-05T20:01:24.973266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:d82e3dd209f02035a43774a9ebc619a2405b9b90ab35343cef24c77ff1b30717

Observation 7899606d-f6bc-4a54-a636-b5bad062d6ab · outbound

This paper cites Evaluating Large Language Models Trained on Code.

DataComp-VLM: Improved Open Datasets for Vision-Language Models Evaluating Large Language Models Trained on Code

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-07-02T21:17:23.903267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:b30e8283aac389e3af5d8f4ede4b80040c9ebb7dfc0380937a522ac206cf70b3

Observation 0daac061-3bac-4b38-b46f-63ecbed1da60 · outbound

This paper cites arXiv preprint arXiv:2602.12237 , year=.

DataComp-VLM: Improved Open Datasets for Vision-Language Models arXiv preprint arXiv:2602.12237 , year=

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:17:24.031180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:d3727ae0a28eec8a6780ac78a211ac67c86fe2caa7762185c502f109a4111e94

Observation 7bf8825d-b14a-4ba6-b3c8-220beb2b887e · outbound

This paper cites an unresolved cited work.

DataComp-VLM: Improved Open Datasets for Vision-Language Models Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-07-05T20:01:24.887864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:e49577c79952c1e52ebeb19045f4781a30318e1210174b3d1150bdcb1c000e29

Observation 8545cc04-a8bd-481b-a9e0-0cb918d0dbcf · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

DataComp-VLM: Improved Open Datasets for Vision-Language Models Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-07-02T21:17:24.158886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:21e48906262cb62efe0c8bdddb3c966ab54b264204b5c97193aecd5989b65cbf

Observation 170a5d6c-2b23-4094-9e14-38e8094149c6 · outbound

This paper cites an unresolved cited work.

DataComp-VLM: Improved Open Datasets for Vision-Language Models Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-07-05T20:01:24.835752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:099e3dab8475a28b40ed8019823a9135d24fbc985a0e537d5918badf1da11c48

Observation ca29f30a-af17-41df-a2dc-3397d8529791 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

DataComp-VLM: Improved Open Datasets for Vision-Language Models Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-07-02T21:17:24.155917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:d426cfba14a39a0c46154b94b2119efcc1c09ba9426637e802d842a909666e44

Observation 98048545-6f73-44d9-9381-16acfa803ead · outbound

This paper cites an unresolved cited work.

DataComp-VLM: Improved Open Datasets for Vision-Language Models Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-07-05T20:01:24.967960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:e11015cd74997b01119db6dd2fc5c86ecef63ed1fe38b7d88d5648522f9a76dd

Observation 840141e3-04e7-4d2b-8350-355c8dc16d54 · outbound

This paper cites an unresolved cited work.

DataComp-VLM: Improved Open Datasets for Vision-Language Models Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-07-05T20:01:24.774560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:d89d55cff9834fc864935f984a96241a0bfd41135fb1d00d653570b0d1abe362

Observation 115430e1-3ed7-4c6d-be2e-554e0056aa28 · outbound

This paper cites PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding.

DataComp-VLM: Improved Open Datasets for Vision-Language Models PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:17:24.161656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:58670cc46277aee51b03da82f56f61c052b8bfec43fe9ff300295a5ab9167df1

Observation 5e1aa486-8110-4749-93f2-469f182bb37f · outbound

This paper cites Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding.

DataComp-VLM: Improved Open Datasets for Vision-Language Models Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-07-02T21:17:24.130067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:7a77002852ac177499385e6012e2253797619160f8a49baf49a949fa7b8ca796

Observation 23750341-1bda-435a-be39-5c9e75df0413 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

DataComp-VLM: Improved Open Datasets for Vision-Language Models Training Verifiers to Solve Math Word Problems

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-07-02T21:17:24.121375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:7feedb38dc125863dbb9037613520efa0e9c9ec166aa1cf596d025865ed73f9b

Observation e015c339-d7d9-48c9-968f-2334fe51fc04 · outbound

This paper cites Conover, M.

DataComp-VLM: Improved Open Datasets for Vision-Language Models Conover, M

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T20:01:24.776347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:bfee2a397a89389f2299a69ba22eb93cf93409cb5f7537e484ea4a40c9f79121

Observation 026175f3-87f4-493c-82a1-e5ddb31e0600 · outbound

This paper cites Contributors.

DataComp-VLM: Improved Open Datasets for Vision-Language Models Contributors

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T20:01:24.788677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:c24043eb6e584db3165dc2e253fd40709df0fef17753adda2024fe69c5252c6b

Observation ef6d4e3a-24ec-4e7a-a19e-37be794217a0 · outbound

This paper cites No Language Left Behind: Scaling Human-Centered Machine Translation.

DataComp-VLM: Improved Open Datasets for Vision-Language Models No Language Left Behind: Scaling Human-Centered Machine Translation

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-07-02T21:17:24.174216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:d5069cc4b43c1e54851d98bd0116a6cf018bbf48dc39520881824632fb9fed79

Observation 5db8b838-809f-4d58-bc3f-eacb7cf92332 · outbound

This paper cites an unresolved cited work.

DataComp-VLM: Improved Open Datasets for Vision-Language Models Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-07-05T20:01:24.790326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:62c9fb7b05c2d28ea924f977a644cd1c20e68331c84d96f32e791881fbc03eb3

Observation dded3567-3014-47f6-a7d1-af4ba708e4cf · outbound

This paper cites an unresolved cited work.

DataComp-VLM: Improved Open Datasets for Vision-Language Models Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-07-05T20:01:24.829229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:6729d4a34341bf3d772e6afd421bccf11eb39ad3daab8bcfc8b56e68be6506d6

Observation 99c1a564-1d17-4dd2-82c3-b3b2f517d658 · outbound

This paper cites 15M Multimodal Facial Image-Text Dataset.

DataComp-VLM: Improved Open Datasets for Vision-Language Models 15M Multimodal Facial Image-Text Dataset

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:17:24.134239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:87c778bed087208d8931135842cd84b69f32ca61cb66ff851f95b0e6faba7c56

Observation e09ba44f-1768-442f-b750-c94c0d3b9e28 · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

DataComp-VLM: Improved Open Datasets for Vision-Language Models FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-07-02T21:17:24.117417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:a481eef9784fa810d658896c4d6f0441b033b9ef1c76c1ffeaaf3fce6a7a6fe7

Observation 7f37429b-8b2f-4a32-9d5e-86b3c555915c · outbound

This paper cites an unresolved cited work.

DataComp-VLM: Improved Open Datasets for Vision-Language Models Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-07-05T20:01:25.023224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:6bb4d0cb820107221ea31151d710cc90f863327ddd8d76f7a0d79a0cfd44be09

Observation 5225a508-80ea-45b0-8b0c-5523c5b5f6a7 · outbound

This paper cites an unresolved cited work.

DataComp-VLM: Improved Open Datasets for Vision-Language Models Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-07-05T20:01:24.978289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:cd81e6730c5c8211d5515faf24038439cdf2e748e92ac2884ecd7c58e60bf99e

Observation dcb6196b-81d9-4d38-9eb5-5f5d17f1b789 · outbound

This paper cites Deitke, C.

DataComp-VLM: Improved Open Datasets for Vision-Language Models Deitke, C

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T20:01:24.976604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:bbc59f0a78822416fba36f0016c83bf8da8c90ea2213c4dd5f4d3830cfe1d910

Observation 488df6a1-06c1-4ad8-91db-b5c21bc98270 · outbound

This paper cites Nvidia nemotron nano v2 vl.

DataComp-VLM: Improved Open Datasets for Vision-Language Models Nvidia nemotron nano v2 vl

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:17:24.126968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:3e76ab78c90bc5b41dcd3899e8b1ebb18aa0160ff36ff89ea491eb318d088ff4

Observation 3174a4f9-0bbf-4627-9b98-2d542c685c72 · outbound

This paper cites Nemotron-CLIMB: CLustering-based Iterative Data Mixture Bootstrapping for Language Model Pre-training.

DataComp-VLM: Improved Open Datasets for Vision-Language Models Nemotron-CLIMB: CLustering-based Iterative Data Mixture Bootstrapping for Language Model Pre-training

Reference 59

Resolution
verified exact
local_arxiv, observed 2026-07-02T21:17:23.908958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:5cb78dccda9fee625b512de35ad5e09236d87981cfd5ef8a58435d82eed39e84

Observation 9650f10a-f78a-4696-a656-37ff6b1cefe7 · outbound

This paper cites an unresolved cited work.

DataComp-VLM: Improved Open Datasets for Vision-Language Models Unresolved cited work

Reference 60

Resolution
malformed identifier
raw_fallback, observed 2026-07-05T20:01:24.981525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:1ab1dd47d9c36421f6de01671ec379c3f73e8f7d4056aceabc5895898da16e4e

Observation 787ce73d-d622-42d3-91e6-510e764315d8 · outbound

This paper cites Dodge, M.

DataComp-VLM: Improved Open Datasets for Vision-Language Models Dodge, M

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T20:01:24.991988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:37d7196cf604c2e5fa405e87fc2f45af63a9ef0fec09d09d7d5bea6445034991

Observation c5210e19-1062-4daf-8855-be439cea0dfc · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

DataComp-VLM: Improved Open Datasets for Vision-Language Models An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 62

Resolution
verified exact
local_arxiv, observed 2026-07-02T21:17:24.098727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:9ea641128fba9a89d7030e074dcc400c03d6b4bc4506e63274e78c851bc83412

Observation 72c965fe-2ba9-4c4b-9c41-c01cf610aaa7 · outbound

This paper cites Douze, A.

DataComp-VLM: Improved Open Datasets for Vision-Language Models Douze, A

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T20:01:24.839270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:deb34f18aacd344dd6dbe99379eae31b2d465c156ff85bc3b57e52142fe2a822

Observation 065125d0-77a2-43f9-845b-c54055b00c67 · outbound

This paper cites an unresolved cited work.

DataComp-VLM: Improved Open Datasets for Vision-Language Models Unresolved cited work

Reference 64

Resolution
unresolved
raw_fallback, observed 2026-07-05T20:01:24.966297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:46b0f0af785b6e4a73ee03eee75a5ca2dce770e0efe5ccfcdf455649a96cc5ad

Observation 097cf026-92c4-4e89-887c-98796734d68e · outbound

This paper cites Evans, N.

DataComp-VLM: Improved Open Datasets for Vision-Language Models Evans, N

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T20:01:24.770851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:2208d91c03c2bff559aa4b56a5f02f13d1e67d361c26425d0c76f0bfbf88d1f5

Observation 206e2812-4b02-4134-b712-e7f4245c5dca · outbound

This paper cites an unresolved cited work.

DataComp-VLM: Improved Open Datasets for Vision-Language Models Unresolved cited work

Reference 66

Resolution
unresolved
raw_fallback, observed 2026-07-05T20:01:24.772769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:8688ab316df4151216f9ad1d7f834f0e454c7afc02d5a6532fae4bc67fe50556

Observation 3bf946f2-bea8-4367-9c5a-623e43bc7b8b · outbound

This paper cites an unresolved cited work.

DataComp-VLM: Improved Open Datasets for Vision-Language Models Unresolved cited work

Reference 67

Resolution
unresolved
raw_fallback, observed 2026-07-05T20:01:24.953749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:9ed92f2a816b9b321ea391a2aada36c92505cf28dcac429dfb8c968a17c8a4ab

Observation 8698bd1a-56b9-4f3f-a178-ec649158660b · outbound

This paper cites Data Filtering Networks.

DataComp-VLM: Improved Open Datasets for Vision-Language Models Data Filtering Networks

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:17:24.102960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:93d92171fc8a9d6bec2447fdffe3bce7fd52daf64ba1d20dca3acaf2c66ab05d

Observation 1fb8cf54-58f4-4bde-8735-f107a5d33d0f · outbound

This paper cites IFBench: Granular Instruction-Following Evaluation,.

DataComp-VLM: Improved Open Datasets for Vision-Language Models IFBench: Granular Instruction-Following Evaluation,

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:17:23.852571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:33f07579c8cbea9543fbb5f1da7545663521a2419c50d636fe20647caa58ba11

Observation e671bb30-5f82-44ca-a0ea-ac72abea6cb3 · outbound

This paper cites Early Data Exposure Improves Robustness to Subsequent Fine-Tuning.

DataComp-VLM: Improved Open Datasets for Vision-Language Models Early Data Exposure Improves Robustness to Subsequent Fine-Tuning

Reference 70

Resolution
verified exact
local_arxiv, observed 2026-07-02T21:17:24.085681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:59c91fd9f9a107fd5331baf2199c69d46a55d625cf24e67b9e27784c8061c2b0

Observation 478bdd4f-b1cd-4963-8565-356f63378af7 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

DataComp-VLM: Improved Open Datasets for Vision-Language Models MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 71

Resolution
verified exact
local_arxiv, observed 2026-07-02T21:17:24.083029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:a24cec4385c3cd8f30f528040cdba4da0749ab0300c87f1f7fc857d6633f18cc

Observation dbdc4178-2341-4dad-8bee-dfaf3fc61570 · outbound

This paper cites an unresolved cited work.

DataComp-VLM: Improved Open Datasets for Vision-Language Models Unresolved cited work

Reference 72

Resolution
unresolved
raw_fallback, observed 2026-07-05T20:01:25.082520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:7b8b0b4080c3b10e39310762ec37f6ca3d9ecaa399b151f0c58d000f4c737172

Observation 3d7d2e3d-741f-4095-9489-217dccf728fa · outbound

This paper cites an unresolved cited work.

DataComp-VLM: Improved Open Datasets for Vision-Language Models Unresolved cited work

Reference 73

Resolution
unresolved
raw_fallback, observed 2026-07-05T20:01:24.984850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:b43f0c81565681dceab7131639efef5186b9679eafe7db92512febb45b11ccdd

Observation c2cf936d-974f-40b9-a620-7dd49e7ba1a0 · outbound

This paper cites An Empirical Exploration in Quality Filtering of Text Data.

DataComp-VLM: Improved Open Datasets for Vision-Language Models An Empirical Exploration in Quality Filtering of Text Data

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:17:24.091470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:668a600ae52d6a846e3c4bf705067c3b4a1c51c8f79a75b163e2c2965a4cd4f4

Observation fe8eec3c-9204-4b21-937a-7aaecc94be4a · outbound

This paper cites The Pile: An 800GB Dataset of Diverse Text for Language Modeling.

DataComp-VLM: Improved Open Datasets for Vision-Language Models The Pile: An 800GB Dataset of Diverse Text for Language Modeling

Reference 75

Resolution
verified exact
local_arxiv, observed 2026-07-02T21:17:24.108289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:9ea58d485ab919a9168aab4115f4ce96a6514d4f5913d6a5d6fd789bf0abfdb7

Observation 7f36c8f0-1de8-488e-b008-af41948273aa · outbound

This paper cites Gervais, A.

DataComp-VLM: Improved Open Datasets for Vision-Language Models Gervais, A

Reference 76

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T21:17:23.583007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:a03976489cb0ddc9567b1856aae09cacb81fa483bc3e5c5b09be4bb0963ecd29

Observation 76545fc9-79ce-4d5a-a984-df179e950359 · outbound

This paper cites Are Language Models Puzzle Prodigies? Algorithmic Puzzles Unveil Serious Challenges in Multimodal Reasoning.

DataComp-VLM: Improved Open Datasets for Vision-Language Models Are Language Models Puzzle Prodigies? Algorithmic Puzzles Unveil Serious Challenges in Multimodal Reasoning

Reference 77

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:17:24.041612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:ca2166a1e715dcd6b03d11713981fcc1571c4f4637336313170e4d1994bab470

Observation 3b74c3f4-1126-41b3-b0d0-d9191f1757a6 · outbound

This paper cites Ghosh, S.

DataComp-VLM: Improved Open Datasets for Vision-Language Models Ghosh, S

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T20:01:25.043009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:7d94027540fe882c035bbb7cc50dec49ad50bb8ee2955382512d3e4b634c317f

Observation 7fac60a4-3afe-41f0-8141-1365878c0313 · outbound

This paper cites Ghosh, V.

DataComp-VLM: Improved Open Datasets for Vision-Language Models Ghosh, V

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:17:24.042268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:0b7ae6204f04e6140060a932af171b5f6f8d5fe577100af610414bd58a14e724

Observation a132b836-e819-49b8-bb20-d1ed4a61c2bd · outbound

This paper cites Glaive-Code-Assistant, 2023.

DataComp-VLM: Improved Open Datasets for Vision-Language Models Glaive-Code-Assistant, 2023

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T20:01:24.948452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:ab7a8894edf91642c61773e398f297dfabc75ed27298dd4faa2d739e90d02d10

Observation 8e38464e-b58d-4cda-9c82-860092200afa · outbound

This paper cites Goyal, P.

DataComp-VLM: Improved Open Datasets for Vision-Language Models Goyal, P

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T20:01:25.024815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:a077331197d92e8cfdbc57b320c60e6bf8002ab7d0a09bd35088704902634611

Observation df863e2f-403a-4d42-a3b0-56c1fc610f98 · outbound

This paper cites Goyal, T.

DataComp-VLM: Improved Open Datasets for Vision-Language Models Goyal, T

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T20:01:24.778212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:36c6a585fea260450468cfef74e9ce5e1d288cf6d7b778533d3efd4b61866c8d

Observation 16612208-c915-4752-9789-33d4066ac477 · outbound

This paper cites The Llama 3 Herd of Models.

DataComp-VLM: Improved Open Datasets for Vision-Language Models The Llama 3 Herd of Models

Reference 83

Resolution
verified exact
local_arxiv, observed 2026-07-02T21:17:24.044845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:e156fd829325e55cfceaa605e76ca8f6f1620d5c177da74c77678cc58400a3d3

Observation ab05b996-0e13-418e-81e6-69e9b34831f7 · outbound

This paper cites an unresolved cited work.

DataComp-VLM: Improved Open Datasets for Vision-Language Models Unresolved cited work

Reference 84

Resolution
unresolved
raw_fallback, observed 2026-07-05T20:01:24.935664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:058f902b04aa6273569a234260f233e88e0c2bfad5bba949df8c765f28b19c28

Observation 231b05a2-f2f3-48ab-84f6-401741e2eb3b · outbound

This paper cites an unresolved cited work.

DataComp-VLM: Improved Open Datasets for Vision-Language Models Unresolved cited work

Reference 85

Resolution
unresolved
raw_fallback, observed 2026-07-05T20:01:25.044667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:c82c2f1fe8e2f4ae692d65f45252986f466d07b1be5b70d6679e87e0bbae2c6c

Observation fd1c174c-5f95-4e6a-84d8-91bc27a2438f · outbound

This paper cites OpenThoughts: Data Recipes for Reasoning Models.

DataComp-VLM: Improved Open Datasets for Vision-Language Models OpenThoughts: Data Recipes for Reasoning Models

Reference 86

Resolution
verified exact
local_arxiv, observed 2026-07-02T21:17:24.048261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:2d56f518f6aecd8012e214be58704daf141b0233b26b73920ce91a0415890f80

Observation e3428f63-441f-4269-99f9-19fa13808946 · outbound

This paper cites Seed1.5-VL Technical Report.

DataComp-VLM: Improved Open Datasets for Vision-Language Models Seed1.5-VL Technical Report

Reference 87

Resolution
verified exact
local_arxiv, observed 2026-07-02T21:17:24.091463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:48e56e917ee6ad40d4601cfd424f9ecbf1d558fa32e2a81a64ae26c64d8e623c

Observation 66e6f1e6-52e1-40b0-ba10-3b585ea52f8e · outbound

This paper cites an unresolved cited work.

DataComp-VLM: Improved Open Datasets for Vision-Language Models Unresolved cited work

Reference 88

Resolution
unresolved
raw_fallback, observed 2026-07-05T20:01:25.015161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:c6a9c9b959656027661348fdcb73e43225aa52f51a8b6d8e323ecbb42dabfaf0

Observation 75843850-55d1-439a-80dc-688e9e893a65 · outbound

This paper cites an unresolved cited work.

DataComp-VLM: Improved Open Datasets for Vision-Language Models Unresolved cited work

Reference 89

Resolution
unresolved
raw_fallback, observed 2026-07-05T20:01:25.041448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:9fd74f0c4ed6159b6786abca3daeae4d30fade3c1b7af0475fb02dbe43cf4256

Observation c993d1ea-6d58-4fd2-afdb-b45f869e21cc · outbound

This paper cites Gupta, A.

DataComp-VLM: Improved Open Datasets for Vision-Language Models Gupta, A

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T20:01:25.016910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:821b990e8e10e593ffce521069dfce86946504c79eb23639250c7827c6fef545

Observation 386c077e-ac45-47bf-8a86-2dc4cb51c76b · outbound

This paper cites Gurari, Q.

DataComp-VLM: Improved Open Datasets for Vision-Language Models Gurari, Q

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T20:01:25.058272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:456c50a38f86b243a0d7298cdd7af89c993965b8a5385c4ffb1adb54c6dadd30

Observation 7643d279-951d-42b0-8d9b-e6dcafb71e32 · outbound

This paper cites ChartLlama: A Multimodal LLM for Chart Understanding and Generation.

DataComp-VLM: Improved Open Datasets for Vision-Language Models ChartLlama: A Multimodal LLM for Chart Understanding and Generation

Reference 92

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:17:24.078675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:5414b0449d1f50906140966ce736072913f4ad9de21e042d5b691607cffc176d

Observation 6d6c1834-3211-407a-9c86-935bbd2ffb7d · outbound

This paper cites Hanu and Unitary team.

DataComp-VLM: Improved Open Datasets for Vision-Language Models Hanu and Unitary team

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T20:01:25.054745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:dfa9757869aebc30559563a142582771cae787b6a76c2dad25c4d3ebe8d3cebf

Observation e9f6c4f1-71e5-415e-9bc6-f65b012bfcca · outbound

This paper cites WanJuan: A Comprehensive Multimodal Dataset for Advancing English and Chinese Large Models.

DataComp-VLM: Improved Open Datasets for Vision-Language Models WanJuan: A Comprehensive Multimodal Dataset for Advancing English and Chinese Large Models

Reference 94

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:17:24.114264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:35326d816ead49e358a672fced95d833785e95b55486d54e94bae6c05433feba

Observation a75e9af8-31b5-43a3-9e71-a4debd4b2aff · outbound

This paper cites an unresolved cited work.

DataComp-VLM: Improved Open Datasets for Vision-Language Models Unresolved cited work

Reference 95

Resolution
unresolved
raw_fallback, observed 2026-07-05T20:01:24.989980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:3a080e77f8e6830339bb572663fb31eb5608a77a774cff0c74b81a65f807750b

Observation 4ffc4dcf-9c0c-454b-add9-9872b245103d · outbound

This paper cites PathVQA: 30000+ Questions for Medical Visual Question Answering.

DataComp-VLM: Improved Open Datasets for Vision-Language Models PathVQA: 30000+ Questions for Medical Visual Question Answering

Reference 96

Resolution
verified exact
local_arxiv, observed 2026-07-02T21:17:23.952660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:f4f9f9da99ff5621eaed212c0936640c1d212df5b65bafbaae4d25f09e1ccbc3

Observation cb4686b8-c5f2-4d3c-b97c-306c93e5f914 · outbound

This paper cites Signal and Noise: A Framework for Reducing Uncertainty in Language Model Evaluation.

DataComp-VLM: Improved Open Datasets for Vision-Language Models Signal and Noise: A Framework for Reducing Uncertainty in Language Model Evaluation

Reference 97

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:17:24.071487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:e8f90200cf4c86ef5cbdb12f331c9bfc6a34787df40ae163cd0666037399b6ab

Observation aeeecf78-24f6-47d4-83d8-9f20fa2bf562 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

DataComp-VLM: Improved Open Datasets for Vision-Language Models Measuring Massive Multitask Language Understanding

Reference 98

Resolution
verified exact
local_arxiv, observed 2026-07-02T21:17:24.102803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:6ab869a079cfcbccfe561a63c17f2cdacaf9f037bca05ee1ba93352d3101615e

Observation d62847b9-4d8f-487b-97da-08357da10f24 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

DataComp-VLM: Improved Open Datasets for Vision-Language Models Measuring Mathematical Problem Solving With the MATH Dataset

Reference 99

Resolution
verified exact
local_arxiv, observed 2026-07-02T21:17:24.057945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:e1d9632b8456217bef2db22cd53724f2b6deba4ce31dd3d19adb5ebba1f5c84a

Observation 947c227c-b47a-4c5f-9bd1-c7b27ac6ea1f · outbound

This paper cites Scaling Laws and Interpretability of Learning from Repeated Data.

DataComp-VLM: Improved Open Datasets for Vision-Language Models Scaling Laws and Interpretability of Learning from Repeated Data

Reference 100

Resolution
verified exact
local_arxiv, observed 2026-07-02T21:17:24.155516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:3c7ec0976312893d30d3787b659c693168951d0dc987fb70a725418382afeac6

Observation 63b553bb-4aa3-404c-8ce6-4e490ff6c7fa · outbound

This paper cites Hessel, A.

DataComp-VLM: Improved Open Datasets for Vision-Language Models Hessel, A

Reference 101

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T20:01:24.811681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:15c1002d74e118962c9ab0a8a1df9f6251bb3bd9d4e9c1a5de52774a7b55eec7

Pith citing papers

Observation 3fb70424-9216-48a8-b94d-89b097a08d86 · inbound

CAPEval: A Decoupled Caption Evaluation across Understanding and Generation cites this paper.

CAPEval: A Decoupled Caption Evaluation across Understanding and Generation DataComp-VLM: Improved Open Datasets for Vision-Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T04:09:49.764855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:09:49.764855Z digest=sha256:92cf1d0719fc34e51ee179c66708c7ea0dd92ef1683a1d519b95309733755997