Pith. sign in

Paper Citation Record · LEDGER

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training

As of 7 August 2026, this Paper Citation Record lists 100 of 137 outbound references and 41 inbound Pith citation observations for arXiv:2403.09611.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2403.09611 v4

Coverage vector

measured 100 of 137 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-16T04:09:36.019146Z

measured 141 of 141 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 41 of 41 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:03:25.101268Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

100 of 137 outbound references displayed

  • verified exact7
  • verified fuzzy35
  • unresolved15
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch43

External citation measurements

11
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 52984cd2-2a08-4832-8309-08850565750d · outbound

This paper cites GPT-4 Technical Report.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training GPT-4 Technical Report

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T04:09:36.440929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:68a19c03a70b81a69da327ecbc9c124bf98b424280960d465d0924bcae026233

Observation ff5ae10a-e89a-4ee7-a990-5b5971cbec47 · outbound

This paper cites In: ICCV (2019).

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training In: ICCV (2019)

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T04:09:36.544311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:37550e31abff78a094bca8288f24dce7ce2253f9c0e2f81b8b8dbe26774a05ba

Observation 3f98930c-b76d-4ffd-af2b-2078e34f6aab · outbound

This paper cites an unresolved cited work.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-05-16T04:09:36.548097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:0f98c4fa4bcb2febbfc78eda20919f3a1eda6698163f50b1a803062f2c3e0210

Observation 31f07fd6-8f3e-49ba-9eed-7e8a91a09192 · outbound

This paper cites OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T04:09:36.180189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:75baff166c227dce9c631219a0b6ef228f659b05b37fa9352041fd153896fb18

Observation e91d606a-c567-4ecb-9f5e-7ae270836808 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T04:09:36.214525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:4c3a960f7473bdd29b866add7c2de19e5e289f387361805abf8a986d16c725a8

Observation 52e6c210-4f2d-4713-84c1-1bcd47a01b8e · outbound

This paper cites In: EMNLP (2013).

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training In: EMNLP (2013)

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T04:09:36.551875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:f912bc0b432a6384acef542572d2ac1a13288dba0214175b90617324377b2213

Observation 7f1acd2c-a893-421b-9128-92049224e440 · outbound

This paper cites AAAI (2020).

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training AAAI (2020)

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T04:09:36.555688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:1639c9741806ba0c54af1a097488b695c90baacd4a82b4ed765bf67f98fecab0

Observation e079e4cb-7c98-495b-8c24-8982cc7f0483 · outbound

This paper cites Training Diffusion Models with Reinforcement Learning.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training Training Diffusion Models with Reinforcement Learning

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T04:09:36.373850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:8520f24d8ebb481caabda2006c88c3b39089165fe21bb1d2232babb0d5111d25

Observation 970824c6-f955-4db4-9983-7c6969419393 · outbound

This paper cites On the Opportunities and Risks of Foundation Models.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training On the Opportunities and Risks of Foundation Models

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T04:09:36.132763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:7a9d3fc54b2a9efb9f15cb3950df8c2d77b6f1e9ca71f31313b7baf3889c4edf

Observation 37cee271-5c8f-466c-95a1-258b6d9700d2 · outbound

This paper cites NeurIPS (2020).

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training NeurIPS (2020)

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T04:09:36.559111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:7ec8862521c1bb0ef380297d0f7e5a3e03c572afc20e4dc7e63399be84f5a6cb

Observation 671852e6-e722-4d51-bdf2-25cfe45285fd · outbound

This paper cites https://github.com/kakaobrain/coyo-dataset (2022).

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training https://github.com/kakaobrain/coyo-dataset (2022)

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T04:09:36.562695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:5c592af6e00d97c8be78cf08342178b0261bf7dee9cc624ebb50809786fc01c3

Observation 7bdbda1d-41da-4d34-8bf9-82923ec64273 · outbound

This paper cites Honeybee: Locality-enhanced Projector for Multimodal LLM.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training Honeybee: Locality-enhanced Projector for Multimodal LLM

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T04:09:36.228739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:f39629162c4d04632a0b17f226f34b7c54b8fd18cd198c36fd94ed35f522ecb6

Observation 42425d51-6414-454d-84a8-ec8c278ea8ee · outbound

This paper cites In: CVPR (2021).

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training In: CVPR (2021)

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T04:09:36.566166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:20991556eb0119a2f71c2639568a9dc4bb6294edbc83f1c294d777f0208bba6e

Observation a3fc9b4e-5f33-4700-8d05-4bb087fed4f7 · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T04:09:36.498427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:9530ed77f0b298b2165cc974ab358ad1c4b53b01f17cc0ae7a5246d7eda1184c

Observation 8e2c4f3b-0dc4-4d48-9875-ca59ea3711f1 · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T04:09:36.127582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:bdf0cfd49cdb8821b8f0c9c05d131b7bb35f2f3848c6e7242abffd4daa6bfe17

Observation e5e08aa1-5855-456e-9fed-a694872c8e5b · outbound

This paper cites In: ICCV (2023).

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training In: ICCV (2023)

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T04:09:36.569961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:fc6e936a4b702465469883db142c6a99281149a7d3e87080fc426f369e2c060f

Observation 28579ef0-fbde-4206-a8f7-2108b1517492 · outbound

This paper cites PaLI-X: On Scaling up a Multilingual Vision and Language Model.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training PaLI-X: On Scaling up a Multilingual Vision and Language Model

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T14:36:10.299419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:edf6e70d10ffd4163c954d80d6ea3a7f9e6c51e08db6f269560168d3367c5c89

Observation 634ff84b-25a1-48e8-8fe9-f4856ffe6db7 · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 18

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T04:09:36.145523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:ae5d294deac2938842d20eef0e8a58b2494978ee2fbae15b5b7fbb75fdb66f9e

Observation 5addcd36-ca0f-4d0c-913a-87b665f0e296 · outbound

This paper cites JMLR (2023).

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training JMLR (2023)

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T04:09:36.573122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:e1aa585a571eb4fc79bcfec12331189d19d98518e0f29bf108b29d20d64b0fe2

Observation a6f2d595-4bbe-4f80-8f1a-d7393d254b97 · outbound

This paper cites MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T16:35:38.347525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:da96ee44a065a17543a21774e4e1f0a4fcfec5bd82af2a5ebda97ba35c984d82

Observation 08f49448-5684-42c0-b86b-f5d11a8854df · outbound

This paper cites Scaling Instruction-Finetuned Language Models.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training Scaling Instruction-Finetuned Language Models

Reference 21

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T04:09:36.193132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:04d50ecdf34dc8b8e96252e4571048baad97dcc8592d4b18d0b8a70e6c855433

Observation 7ee52c6a-7a32-48d4-8a35-5a7120a86125 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 22

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T04:09:36.199890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:fff43735084e2d86d6ef4de64d544fe95decfb3f7aff17484142a2778ade78ad

Observation 99ab4eca-51ac-4c8a-9a8f-62e3b3f9b773 · outbound

This paper cites DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models

Reference 23

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T04:09:36.207078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:233ee75daf8a89fa465226e24661148583ea6149da0f7a600f4c4b6804d9f676

Observation ead157dd-857d-4ad3-a04f-660a503e8b3c · outbound

This paper cites an unresolved cited work.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-05-16T04:09:36.576986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:97602157a67633cb3534433e23e8ae28d2b09bd999170b7cb2a8da7596c7d041

Observation 2a60623f-18ef-4d72-9ccb-63e5341c61d3 · outbound

This paper cites an unresolved cited work.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-05-16T04:09:36.580427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:351af83ab633fd3508e76534536fa5fb4624fb2eb2bc1a78e99fdba92daf6523

Observation 942439c7-c3f0-438d-a349-3cdbb21074e3 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 26

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T04:09:36.247419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:741b435e19bc5e757eb5abbf7d7dc5d497df539066d0218b66fd923b33506d7d

Observation 6e9e65cb-c432-462a-80a8-2739a352a78c · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 27

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T04:09:36.252729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:5bcd9eca2eef64b5b7df740dbb95c4e09c5768f11e7c80bd2c17026ca7bd7e71

Observation 02ce5c85-46d4-4d83-ab83-ad289e7fc8c3 · outbound

This paper cites PaLM-E: An Embodied Multimodal Language Model.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training PaLM-E: An Embodied Multimodal Language Model

Reference 28

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T04:09:36.320359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:bfaf0bf3e728bfa9d2131edd0c10711ba8026991787134c827139229f9621cc5

Observation 2b6f2cf3-f31f-4742-aaab-f310aba4c423 · outbound

This paper cites In: ICML (2022).

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training In: ICML (2022)

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T04:09:36.583602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:37e7214a5fea6b15d4ebdea749662298148e5fac32c3634d6c2b0135ccf1f9c2

Observation cc51c86e-8447-4528-87ac-473afe76094a · outbound

This paper cites Scalable Pre-training of Large Autoregressive Image Models.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training Scalable Pre-training of Large Autoregressive Image Models

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-16T04:09:36.398124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:9a6e84809934c1a8bed1c5d6a08a92e831d26ef3227193a1e3c29c246649a6b7

Observation f77a8e15-076d-4275-ad1e-8aeae032a19a · outbound

This paper cites Data Filtering Networks.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training Data Filtering Networks

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-16T04:09:36.465728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:2ce7b7a44a342ea526be0dbebeec2b66e82ebea8b3b0839c3d3f3bfc53e4f0bb

Observation 7623c334-3a45-4175-ac6c-ea1bce611f53 · outbound

This paper cites an unresolved cited work.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-05-16T04:09:36.586798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:2bd8d8f19a4b4b1bdeb05a1efca4938d3681fbf372364e7cdaa6097a6ecc1896

Observation 80c1ef7a-c135-48fa-9e2b-aa9d77b07421 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 33

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T04:09:36.522638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:3ab289515c258f6da1b23f751b78a94c28b193df3cfde8b1102766c978fba132

Observation 81b7d2a8-df2a-4a4d-838e-2ac8de554557 · outbound

This paper cites Guiding Instruction-based Image Editing via Multimodal Large Language Models.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training Guiding Instruction-based Image Editing via Multimodal Large Language Models

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T04:09:36.527995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:2b5b812742fe90ed21ebc22c70c1824a14ce08282eb04f197fb544f55a1897e0

Observation 13e48b14-591c-44f2-82bc-19a8764a8f82 · outbound

This paper cites https://doi.org/10.5281/zenodo.10256836,https://zenodo.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training https://doi.org/10.5281/zenodo.10256836,https://zenodo

Reference 35

Resolution
verified exact
doi, observed 2026-05-16T04:09:36.095815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:b2b4a8167c47741f285822195b61cb511efe71c2d45ff93e87da56c388d9c2bf

Observation af8ae0c2-4493-4ead-8c91-db74540d08e0 · outbound

This paper cites SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T04:09:36.109213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:d95e4c20c57e621be51aead524ed0aa3372779331e94d37a268c093aaf68c911

Observation 5c10f305-4288-4827-a656-462842bf09e0 · outbound

This paper cites MultiModal-GPT: A Vision and Language Model for Dialogue with Humans.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training MultiModal-GPT: A Vision and Language Model for Dialogue with Humans

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T04:09:36.115905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:d41db6b5cb0222599ae11cb62cfc56a4407ccc92b355a1bfba8e4a46b10f77f7

Observation 49dccf1b-219c-4baa-9507-872f9a74bfe4 · outbound

This paper cites In: CVPR (2017).

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training In: CVPR (2017)

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T04:09:36.589896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:4d11284001175ac537223ad61945523b4169f7e4e2c3bfaf47d03dc375662043

Observation 39177b3e-2dd3-453d-be19-866453919463 · outbound

This paper cites In: CVPR (2018).

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training In: CVPR (2018)

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T04:09:36.592975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:3747bf9cb44d335e3d42319e5af0f1491cf932b69054aecc5f9a52b18ac77d04

Observation 09e96b13-8f7b-4534-955e-ea6b92dfd300 · outbound

This paper cites In: CVPR (2022).

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training In: CVPR (2022)

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T04:09:36.596132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:1d096637ba7f3b5b29d50e5c1300004b43431e44104f5fbfe8edc85ee907ca7c

Observation 948384f8-c340-4e11-819f-f329b793ebde · outbound

This paper cites In: CVPR (2016).

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training In: CVPR (2016)

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T04:09:36.599739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:76ab3e22821ddeff71354a73f84372c89d021a0bff53b602d61ac02046d0ccfe

Observation 10fc5fcb-2c2c-449c-bf9d-7a35a5b434e0 · outbound

This paper cites Efficient Multimodal Learning from Data-centric Perspective.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training Efficient Multimodal Learning from Data-centric Perspective

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T04:09:36.158828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:a1ea20dc1d787b73356e3aead7a9e659bcc8aa68859a7b8a4931f7fe4d9281f9

Observation 4a7dad1e-d0d9-456e-91c1-11bf8ec428f5 · outbound

This paper cites Scaling Laws for Autoregressive Generative Modeling.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training Scaling Laws for Autoregressive Generative Modeling

Reference 43

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T04:09:36.169325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:4d7148a6ac560f744a55f88123d7d3552cb48bb9faf525901f407283201d5800

Observation e0854ebc-7f93-4679-b5ac-fb21eb035aac · outbound

This paper cites an unresolved cited work.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-05-16T04:09:36.603096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:23829e3adf0a87a4c6baceff5aea86cea4e6c0a1857a0ad0246068e16ce8895e

Observation aaf443ae-937c-4c18-b6b5-7042acc417bc · outbound

This paper cites an unresolved cited work.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-05-16T04:09:36.606308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:e91960245618ce1b7ccc9f2d92810154e241364869db7ba927dd1a889c3aaf3b

Observation d8cb4cf0-c05b-4ebd-b965-5fd44d7a30fc · outbound

This paper cites In: CVPR (2019).

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training In: CVPR (2019)

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T04:09:36.609658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:c7fbc4a9e235663dd2e45f8096dc21cb5271a39a554503a0e6afef6806a3c94c

Observation c3af3a71-f461-4258-b78c-40d68e878ec3 · outbound

This paper cites https://huggingface.co/blog/idefics (2023).

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training https://huggingface.co/blog/idefics (2023)

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T04:09:36.612755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:a840408bf394c92aa0f7f6c824ac2163dd700473202810e6c8eb0122de825f66

Observation 4187dedb-bbe1-475d-9bc6-3132a3ed0cbc · outbound

This paper cites an unresolved cited work.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-05-16T04:09:36.615957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:9fb1daa0dff044201f19c1302a2a69070424cf26b87d48abcf115b90bc845368

Observation 6cd8749b-43f6-4a2f-9ad3-2b64a24af156 · outbound

This paper cites an unresolved cited work.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-05-16T04:09:36.619196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:77eb8a05c8a15b3f055b822724e672fdec1007694164e1d530ee51be155f1c04

Observation eb5610d7-05a4-4afe-91d7-69ad82224a29 · outbound

This paper cites TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension

Reference 50

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T04:09:36.221334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:5e261fbc4f23a6d1fbf30a39706dbe90b52f631be5b03a0c57bbd4c0015835b4

Observation e0595f8f-a6e0-475f-bc56-8139c0607170 · outbound

This paper cites In: CVPR (2018).

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training In: CVPR (2018)

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T04:09:36.622409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:6ad129f9b2a42d10155e12a0c4a383d7f073116513b77064e811ffde70845eae

Observation 51721e93-c87c-456e-9d5b-a5110834c7b1 · outbound

This paper cites In: ECCV (2016).

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training In: ECCV (2016)

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T04:09:36.625838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:ee5aeb5e3239192c3b0316eb77cd2c74c1dbca7ea5a48d899a42bdeccecb46d6

Observation 55fcfc5b-c42b-4508-ab2d-47d917163cbf · outbound

This paper cites In: ECCV (2022).

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training In: ECCV (2022)

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T04:09:36.629426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:346062e17592aa193380c0b580dcab7039c0738f91afc3ebe97ec0ef3e8ac410

Observation 574486fd-fdde-468e-b95a-9d2b89837cd2 · outbound

This paper cites Generating Images with Multimodal Language Models.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training Generating Images with Multimodal Language Models

Reference 54

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T04:09:36.284826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:3f6d489e3e0e6cc19f5458f6fc636d80e8136049a0aa38c75c06f6fa85e5a9aa

Observation 0d4ca5d2-cea4-450d-ae1e-0f88769819bf · outbound

This paper cites In: ICLR (2023) 20 B.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training In: ICLR (2023) 20 B

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T04:09:36.633169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:b5ff072bab9f294ae497ad248eb07229afd97e3db6f35ad35abe98eb64aced1a

Observation 3b836275-5181-4428-b431-2b6e90226b10 · outbound

This paper cites LISA: Reasoning Segmentation via Large Language Model.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training LISA: Reasoning Segmentation via Large Language Model

Reference 56

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T04:09:36.331924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:0a8427b53bd3de3f8ef747a65e2401215557cf55d7c919a6eac70ec9db14478f

Observation e7fd15e9-a78d-4c20-b9bd-1793625f42e7 · outbound

This paper cites VeCLIP: Improving CLIP Training via Visual-enriched Captions.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training VeCLIP: Improving CLIP Training via Visual-enriched Captions

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-16T04:09:36.342638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:c8b28dd789e103900011b9e20181ce5937086dd2b9e9bcde2d23dcdb83c189bd

Observation 1c8ec99a-7469-475b-b139-f4cd71a4b265 · outbound

This paper cites an unresolved cited work.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-05-16T04:09:36.636218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:6bd3a78120486eda44ca3e22935d7a8708343744a6e19e83229280e0fddfeccb

Observation 9e5fe84a-a7cc-4f80-a240-335a5e24ed15 · outbound

This paper cites In: ICLR (2021).

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training In: ICLR (2021)

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T04:09:36.639720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:7a5dff253c69fb551fe8c434b846d9a447976df10f3a4fa59639cf15889d2f2f

Observation 88cc9aa3-66f0-45ff-8787-07577844ec3b · outbound

This paper cites MIMIC-IT: Multi-Modal In-Context Instruction Tuning.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training MIMIC-IT: Multi-Modal In-Context Instruction Tuning

Reference 60

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T04:09:36.435534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:cdfcb8a713ac1a295bb67bee0d0ef5ae4d7e7c71718b7b4363cf9edc76b26268

Observation 38f5fa54-ffdb-4752-abd5-3518a7d1718d · outbound

This paper cites Otter: A Multi-Modal Model with In-Context Instruction Tuning.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training Otter: A Multi-Modal Model with In-Context Instruction Tuning

Reference 61

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T04:09:36.448062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:56d5087647cd360d451734849b02d50ea7408e14b19c1432af6ed4edff7a1fee

Observation 9dbfeace-473f-4087-bcf4-30b28aa6b0bd · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 62

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T04:09:36.453862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:ff7f686cd47b48b860f07a5ffdb544ba088695d51fff467273a6d013e3042e22

Observation 125869bc-ad0a-4f6c-ae49-1192b81d123b · outbound

This paper cites Multimodal Foundation Models: From Specialists to General-Purpose Assistants.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training Multimodal Foundation Models: From Specialists to General-Purpose Assistants

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-16T04:09:36.459955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:29cab2ce9c0e7c56f0c44949db825a4d225eb6f3efb634822ec4ea7581e322af

Observation fa4f27df-5e83-4913-8cb8-e909bf4afdb8 · outbound

This paper cites an unresolved cited work.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training Unresolved cited work

Reference 64

Resolution
unresolved
raw_fallback, observed 2026-05-16T04:09:36.643433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:ad57ee087e5d5a55fb9744e2ae4cf45cc93b1b12738431564be097f2c99ecab3

Observation b2246ba8-9403-43fa-9f09-e147d5e42781 · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 65

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T04:09:36.471641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:83a2892b1c0bee6750c4ebc6d06a07c4cefd8f4567323521e343c6dc263c9932

Observation 0d84ef2d-5fec-424d-922f-c386ee15a067 · outbound

This paper cites M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning

Reference 66

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T04:09:36.478013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:d5f4eff465aeca1316a8991e2b3fc32169ba3898a878a3c4f6c362ac704a80be

Observation 71e9bf56-b60f-4038-b075-f2bbf19b93c0 · outbound

This paper cites VisualBERT: A Simple and Performant Baseline for Vision and Language.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training VisualBERT: A Simple and Performant Baseline for Vision and Language

Reference 67

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T04:09:36.483344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:03b8de84e239fc508a36fea2743d02942e9b833c546e2df03cce865e85b0e68c

Observation b678ba0d-8536-4d27-ba67-4011e718dfa2 · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training Evaluating Object Hallucination in Large Vision-Language Models

Reference 68

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T04:09:36.488296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:af1fcd6ba8b3bf5225e9fc5b58af110ab8fdb527a63af1c310182bd8cf7c2917

Observation 6b69ddb0-f51d-49e8-ad2a-daef9811f7b1 · outbound

This paper cites Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 69

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T04:09:36.493562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:1a6e96806939d99581f30df2f6d73469454778c71d76c2f7801e534592f5f700

Observation c684790d-403a-421b-b6b9-7a70d445709d · outbound

This paper cites an unresolved cited work.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training Unresolved cited work

Reference 70

Resolution
unresolved
raw_fallback, observed 2026-05-16T04:09:36.646929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:31f7dbfe723fc87d7071c1389271cccf7ab3b1786c8b268aff2c258ce9f8c52c

Observation e62c1d3d-15ac-4a15-ab2e-037fa7173fb2 · outbound

This paper cites VILA: On Pre-training for Visual Language Models.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training VILA: On Pre-training for Visual Language Models

Reference 71

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T04:09:36.503622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:ea55d356f8ebe1b22f416d76515534760a8c5b86d6e5c08c9e6ab733fa1a60b9

Observation 4667a88a-6cea-4f9a-9748-108961504995 · outbound

This paper cites Microsoft COCO: Common Objects in Context.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training Microsoft COCO: Common Objects in Context

Reference 72

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T04:09:36.508432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:7c8238e7590f522d9dcc082fa23f799233b8fe4d0745dda9628938357116b1eb

Observation e9750ee8-4067-4275-8a67-a83b2d563b43 · outbound

This paper cites SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:03:27.047113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:8bbaabe1f9b88f96f9ffb68d01375cc2ef9acd427d71706e7821f33c7a7134a7

Observation a7572d3a-e83e-4516-be1e-e996279e2ed1 · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training Improved Baselines with Visual Instruction Tuning

Reference 74

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T04:09:36.518232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:18f3688f614a4653a036d8a448aa03ec2003a296a72f8ad0fe5bbb67d7396159

Observation 05ed3e91-c99a-4fec-a901-904bf5699223 · outbound

This paper cites io/blog/2024-01-30-llava-next/.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training io/blog/2024-01-30-llava-next/

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T04:09:36.650865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:27089cf5169a283cf1db5ec67026cf52f101d4174e6a2725c9eca17b49210e7b

Observation 003de143-d596-4f17-bc3f-430b4edce97b · outbound

This paper cites an unresolved cited work.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training Unresolved cited work

Reference 76

Resolution
unresolved
raw_fallback, observed 2026-05-16T04:09:36.654426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:21cda7c142a335c5e3a76f3b305cfc6b92e6419ff11dcb00700c30decb8601af

Observation 1ffe80b4-f047-46b4-9bbf-997312599f30 · outbound

This paper cites LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents

Reference 77

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T04:09:36.536840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:5de655e56ff6eda665f5a3a7745446ef06d17352ce74bde2c0ae0fce21b6bff4

Observation ccdcdf43-7d9a-4726-9c80-dda1bfa04da9 · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training MMBench: Is Your Multi-modal Model an All-around Player?

Reference 78

Resolution
verified exact
local_arxiv, observed 2026-05-16T04:09:36.532290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:57dec6b567661bec358bc04851f85f7edf219a22aa6907e7cd0a1aef46f5b63f

Observation 5cabe2d8-13c7-4438-96c3-ace57effbacc · outbound

This paper cites NeurIPS (2019).

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training NeurIPS (2019)

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T04:09:36.658025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:94570e9901ae8c914b89c8d9be64d058c63c0b03351efb3cdc19e0e40117bf1f

Observation 73c880c6-bb6b-4a68-a0fa-272319facf30 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 80

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T04:09:36.102982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:fa8d347cf9276708ea4e4c0ba2ce889201e481b32b709c97fc326432bdce4210

Observation 871a3cd4-4e37-40f2-b055-fdc1f36fbf30 · outbound

This paper cites NeurIPS (2022).

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training NeurIPS (2022)

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T04:09:36.661678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:583f8fc0f3fc9703da42d47abc50ac1f778c55578c552c41d3c8e286bd097961

Observation e76a5442-ac1e-4876-8b39-6e7a7df635bf · outbound

This paper cites In: CVPR (2019).

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training In: CVPR (2019)

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T04:09:36.665293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:a0bdcd1356f12d453847522caf9a973ceadda73124dacbe6e691dba3896c76a4

Observation bd20772e-827e-4fe0-8507-3e43805db485 · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 83

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T04:09:36.121556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:7720a2a139f1d9b8a7573f931192db379703ecb77f172bb4aa84eaa2a50e1263

Observation f35af081-a5c1-449f-9c66-9bf0da08ae3b · outbound

This paper cites In: WACV (2022).

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training In: WACV (2022)

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T04:09:36.669063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:5347a2b5bdcd37d6b95646237bb8363d91cf9834c8b3c27168bc66bfb7d31cdc

Observation 30305ea0-6963-4d35-b7a0-f847132a4b3e · outbound

This paper cites In: WACV (2021).

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training In: WACV (2021)

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T04:09:36.672927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:f75fea34f7aeaa48798109d838a6436a0d01f14cb528b8fd28be33603bce1406

Observation 8bf52fcc-1aa3-462a-a583-941e5644c6be · outbound

This paper cites In: ICDAR (2019).

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training In: ICDAR (2019)

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T04:09:36.676624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:01e0361649b7b6e96f1d747ecd90f1d5c42ede41dd3cb6c9320a581bca59847c

Observation 01d8b1b6-8edb-492f-bc97-5338971b6e4d · outbound

This paper cites In: NeurIPS (2022).

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training In: NeurIPS (2022)

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T04:09:36.680572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:e060858e09dd7ce71c3da73ef897bd79293de6630caaa1f301d034becfb84047

Observation 6827c88e-c9f4-4974-9d28-a097dea10b28 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training DINOv2: Learning Robust Visual Features without Supervision

Reference 88

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T04:09:36.151501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:df5e87b7069853b2f9dd5bcb1945163cf37142ecc50ddef7ba8489938144a7e5

Observation bbabfbed-a48e-485b-ad9c-e1e65b7a0f60 · outbound

This paper cites an unresolved cited work.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training Unresolved cited work

Reference 89

Resolution
unresolved
raw_fallback, observed 2026-05-16T04:09:36.684792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:82fa5fd782443fa546307366db2e81f76ed1a157226cbb110b3d92077e3c2640

Observation 931fbc62-baa8-44de-b477-0f568596d534 · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 90

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T04:09:36.163953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:83ba7bdfd00aafb798a3e83b1471c63e693e25e994af631b7f5bb714aec0fd01

Observation 8109ac73-0b16-4720-80b7-f2af40eab767 · outbound

This paper cites In: ICML (2021).

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training In: ICML (2021)

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T04:09:36.688939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:163ceec478babdbed3626fd75f21961b09750311eb361b09d8985be3564f4962

Observation 1bedb279-c508-4f94-b660-122cce42945d · outbound

This paper cites Scaling Language Models: Methods, Analysis & Insights from Training Gopher.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training Scaling Language Models: Methods, Analysis & Insights from Training Gopher

Reference 92

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T04:09:36.174953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:f8ba8c2a8f88f04d7ded00ca3b43f87baf1bde32b5db79287a8e94a1e08cb997

Observation 5123a8ea-01d0-42e5-b3bf-7100642977ba · outbound

This paper cites JMLR (2020).

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training JMLR (2020)

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T04:09:36.692817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:14afad0fdac69d34df12176d227c05b21831f2eb8a21430c92117da24c1d1007

Observation 00ea4575-ceb4-4f8a-9798-008834fa56ae · outbound

This paper cites In: ICCV (2023).

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training In: ICCV (2023)

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T04:09:36.696773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:27a468cf60e8247b231dae269580e3749f98a29547d02bae96018eec0a9c60d5

Observation 1f945584-c92f-4092-a28f-6cda591bcfd7 · outbound

This paper cites In: CVPR (2022).

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training In: CVPR (2022)

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T04:09:36.700458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:20c90b1f379b1629f7706b5414b0ba652a315fe6525ba0ba529498aa413fcfca

Observation 22655654-3f06-49e7-ab3d-77bc325d3960 · outbound

This paper cites In: Beygelzimer, A., Dauphin, Y., Liang, P., Vaughan, J.W.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training In: Beygelzimer, A., Dauphin, Y., Liang, P., Vaughan, J.W

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T04:09:36.704444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:a9aec52b3dc5106e8a803f130ff1acc7025c44a52bc8b20400f5f00722a2875a

Observation f75281d8-deb7-4374-a1fc-6da4e35dfaa4 · outbound

This paper cites an unresolved cited work.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training Unresolved cited work

Reference 97

Resolution
unresolved
raw_fallback, observed 2026-05-16T04:09:36.708663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:f9f6cfee4b48812b355870a8222c4c94462870733f3a6c4897ed235e1e9ce342

Observation 50896e14-6c60-41c3-9b5f-547277fcfeaf · outbound

This paper cites In: ECCV (2022).

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training In: ECCV (2022)

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T04:09:36.712686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:0dd4d391c407d8df183fb48c7a014b233206aae8ab44892609de7de1ade5a171

Observation c87d41e1-e7dd-4f46-99e5-8ae662f85539 · outbound

This paper cites an unresolved cited work.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training Unresolved cited work

Reference 99

Resolution
unresolved
raw_fallback, observed 2026-05-16T04:09:36.716261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:04d06f7123a05d1d156d85ba778e69e76b996d9faf55e7d3578b25c9cc924ecf

Observation d33cb590-3fef-4e27-b114-29e0d2f12b11 · outbound

This paper cites In: ACL (2018).

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training In: ACL (2018)

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T04:09:36.719797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:c998e493225fa1e55a0437993aad6789748439e6927c798b40b2dbba40fd3d9f

Pith citing papers

Observation 14b8c1fe-c569-4462-9952-7a6f42276d1b · inbound

A Survey on Multimodal Large Language Models cites this paper.

A Survey on Multimodal Large Language Models MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-16T04:09:36.761640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T02:56:41.658658Z digest=sha256:cb98ac4031ead4a566a439829e25739fc9b69b842e0bd6ff6dbd3a7d385885a9

Observation 61c632e2-a024-4ecd-969a-440c2fc20f46 · inbound

Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone cites this paper.

Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-16T04:09:36.761640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T20:19:27.255515Z digest=sha256:41ee3b5bc11de17a4aa5935fdc0c740140c9de7902df361aa8c07f5ff047dfed

Observation 74b488da-e850-4f47-a3e4-a542f4999c6a · inbound

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites cites this paper.

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training

Reference 84

Resolution
verified exact
arxiv_id, observed 2026-05-16T04:09:36.761640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T20:58:58.849040Z digest=sha256:a0802c4d9497c5fd207ed00e3a07bf0c5a0f6d35831290f135654a79f43474cc

Observation f1e3ac38-2099-4fc0-a54e-1fd79265e8e8 · inbound

OpenVLA: An Open-Source Vision-Language-Action Model cites this paper.

OpenVLA: An Open-Source Vision-Language-Action Model MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training

Reference 87

Resolution
verified exact
arxiv_id, observed 2026-05-16T04:09:36.761640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T14:46:35.942338Z digest=sha256:73dd57a96bf345467499ed43478260d1c753513c477e247c7e3c2c40dc8fe85e

Observation 0cc19cb7-f707-4a8b-abe7-bba47e7b01d8 · inbound

Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs cites this paper.

Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training

Reference 100

Resolution
verified exact
local_arxiv, observed 2026-05-17T00:05:03.829164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T00:05:03.547664Z digest=sha256:ab3b60d281cacc7bcc20c5341a3789139fd7774220eaa954c2296da73f8d3b07

Observation 0cbfdc07-8249-45ed-9f01-533a10a4fe7a · inbound

PaliGemma: A versatile 3B VLM for transfer cites this paper.

PaliGemma: A versatile 3B VLM for transfer MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training

Reference 96

Resolution
verified exact
arxiv_id, observed 2026-05-16T04:09:36.761640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T13:10:19.972353Z digest=sha256:8c75e16e6deabae813e4eded1d49d5ca17e51a7dd5fdcbf20a2901099c2760e7

Observation 1fda17c4-bc48-4651-9e3e-e1407b6fa0ae · inbound

LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models cites this paper.

LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-16T04:09:36.761640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T06:01:53.730356Z digest=sha256:db6e565b6fbd3890afe626d7c2e19daf3fd3d455592539c98fdc3b0aa9beaf69

Observation 5e3cba1e-776f-47fa-bcee-c681ddb079c6 · inbound

MiniCPM-V: A GPT-4V Level MLLM on Your Phone cites this paper.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-05-16T04:09:36.761640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:a821bdf97a5af14aca83a62fb88f818fc129054c711266b864f98f6dd2444552

Observation b4bc1dab-bd26-4776-9b60-1e105abdc0bb · inbound

LLaVA-OneVision: Easy Visual Task Transfer cites this paper.

LLaVA-OneVision: Easy Visual Task Transfer MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training

Reference 104

Resolution
verified exact
arxiv_id, observed 2026-05-16T04:09:36.761640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:b638bff635a1931aa605468b7e9b7d175e8062e7c587bbd6eba87846ed3d1be4

Observation 2b45c23d-f818-4743-b9c3-15974fff83ba · inbound

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models cites this paper.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training

Reference 178

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T06:20:36.457574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:a1baf4d3ba0f491628cdbadc951542d3bfa57bf5008a7792f67c1b66bb9852f9

Observation c3d25bb7-cf7d-4f64-83e8-2634643781dc · inbound

Show-o: One Single Transformer to Unify Multimodal Understanding and Generation cites this paper.

Show-o: One Single Transformer to Unify Multimodal Understanding and Generation MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-16T04:09:36.761640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T21:03:33.427939Z digest=sha256:6d70c7dd47c1c0193de34878569e99df9102ea3c2f6f5fdd3e182c471e307d75

Observation f9614d6f-21ae-4f82-95a6-fcf0f851adaa · inbound

MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans? cites this paper.

MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans? MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-05-16T07:59:32.843836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T07:59:32.638758Z digest=sha256:5a2686099cc9b9e7730b542f8dd16498634fea212b89b95bb6d9411e40a1cad6

Observation db67f8a3-8083-40ea-b8f4-c41282c2bd7a · inbound

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models cites this paper.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training

Reference 85

Resolution
verified exact
arxiv_id, observed 2026-05-16T04:09:36.761640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:6a7f1820f78c258729fa98a54b6d9b080bdc48fab86d4c0a5c6ce7bda79830b3

Observation 40b0d271-7a01-4da9-a3ec-325d39165583 · inbound

PaliGemma 2: A Family of Versatile VLMs for Transfer cites this paper.

PaliGemma 2: A Family of Versatile VLMs for Transfer MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-05-16T04:09:36.761640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T09:15:07.523565Z digest=sha256:bbd0ed4e61e235f0f2f79c0737cd4dae19273b9099ea4e504520951d2c28e456

Observation 43f2bb38-b30a-4eb6-b9ae-0c82ada3946c · inbound

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling cites this paper.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training

Reference 185

Resolution
verified exact
arxiv_id, observed 2026-05-16T04:09:36.761640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:75ee14f91533d08701d002fadac168acd2922a70e8d1f9988199298edde87d18

Observation 0ee7314c-6c4c-4dfd-99ef-63d973fe5a0a · inbound

MetaMorph: Multimodal Understanding and Generation via Instruction Tuning cites this paper.

MetaMorph: Multimodal Understanding and Generation via Instruction Tuning MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training

Reference 172

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T07:51:13.267056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-17T07:51:12.953777Z digest=sha256:4dbd2516b3823667e57c96e34d8dd990e33b878cc68e33268dd988067f9865fa

Observation 010c6de0-5704-47ef-87d1-90b6876a9286 · inbound

Semantics Disentanglement and Composition for Universal Image Coding with Efficiently LLM Reasoning and Generative Diffusion cites this paper.

Semantics Disentanglement and Composition for Universal Image Coding with Efficiently LLM Reasoning and Generative Diffusion MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-05-23T06:42:39.555305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-23T06:40:03.347536Z digest=sha256:0f585242b5e4d0eeb44d2ee3d2abc892a68b8b990806dfd973a559eb32fb2611

Observation b2c3e8b0-a2cb-48c5-8adb-620909f0f6ec · inbound

SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features cites this paper.

SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-16T04:09:36.761640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:49:22.279848Z digest=sha256:ae8f7ae003eb15cb61c70be094e6334910befe7dd8383e4de103f5953b67ad44

Observation ead99e0d-d6a2-4900-aaea-d29ea3561f95 · inbound

FLARE: Fully Integration of Vision-Language Representations for Deep Cross-Modal Understanding cites this paper.

FLARE: Fully Integration of Vision-Language Representations for Deep Cross-Modal Understanding MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-05-22T19:52:01.883820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T19:49:00.961388Z digest=sha256:d97eee61762be9feb35010db03a833aa0bb94e88dd5b2359e0f2eb4f70fece5d

Observation 0c45f6b8-6c9c-4d34-beed-57b63926474f · inbound

MMaDA: Multimodal Large Diffusion Language Models cites this paper.

MMaDA: Multimodal Large Diffusion Language Models MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-16T04:09:36.761640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:5d2cff0035c64194a9607280341b5f9bab00421e4c38d276e1b7be318fbdff55

Observation b3d02633-9be4-4692-80de-134a95a7d6fe · inbound

Seeing Far and Clearly: Mitigating Hallucinations in MLLMs with Attention Causal Decoding cites this paper.

Seeing Far and Clearly: Mitigating Hallucinations in MLLMs with Attention Causal Decoding MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:25.101268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:03:25.101268Z digest=sha256:3be5f8bca4a45badf5cdce4e65dbec4ba55f70ff80378b8940a66b8e5463303c

Observation 301aff66-ce55-45c8-91de-c03f5c645ec0 · inbound

Multimodal LLM-Guided Semantic Correction in Text-to-Image Diffusion cites this paper.

Multimodal LLM-Guided Semantic Correction in Text-to-Image Diffusion MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:06:49.131004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:06:49.131004Z digest=sha256:df99a66afd092f397cefda79d0cc743bf6d5d7a6932e9725fbb57bd5f083f438

Observation d246c863-0a36-4ae5-a52f-ff926dd5d1d1 · inbound

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought cites this paper.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:23.938819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:23.938819Z digest=sha256:1980026e0950e5c6deb0e0386675296d79b353df3bd4e7ba3521fac610835523

Observation 9e39d9be-e59f-403f-976d-c0213474736c · inbound

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation cites this paper.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:43.703061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:43.703061Z digest=sha256:b7471898ffacb1ee71d4f5f253fa0e5339ec1e747cbf3fcdfd63a66659f03cd9

Observation 6e4003eb-1b66-4a0d-bfdf-f24097734015 · inbound

ReFoCUS: Reinforcement-guided Frame Optimization for Contextual Understanding cites this paper.

ReFoCUS: Reinforcement-guided Frame Optimization for Contextual Understanding MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T11:51:28.012308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:51:28.012308Z digest=sha256:45df54713dc05898ecce7ccbd4efe6b4c27864c68a5dc2df19cc26fdc4ae7dee

Observation b72702d1-76c2-49ca-ae00-13d7e5d4cdd0 · inbound

Rex-Thinker: Grounded Object Referring via Chain-of-Thought Reasoning cites this paper.

Rex-Thinker: Grounded Object Referring via Chain-of-Thought Reasoning MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:14.759449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:14.759449Z digest=sha256:acffd3ed1ab8545680b7feb200082bb1e0bc6eb7db9be0476f7612181f434c40

Observation b38f56ce-fee3-45bf-a545-1ee82014b56b · inbound

Synthetic Visual Genome cites this paper.

Synthetic Visual Genome MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:56.348041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:56.348041Z digest=sha256:23310e9e38d7866f49abd0ffcc6880e5dc198622d512dc08739b6bd0d540d1f6

Observation bd027e17-7925-479b-97f7-7d39f0994751 · inbound

GenRecal: Generation after Recalibration from Large to Small Vision-Language Models cites this paper.

GenRecal: Generation after Recalibration from Large to Small Vision-Language Models MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-06T23:57:24.755181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:57:24.755181Z digest=sha256:26d393685bdca0b36d5115780f30e2f403c898ff5d73038ac44b7f69c8046b76

Observation 2280f4f3-b2b0-4be1-838b-7c667339728c · inbound

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets cites this paper.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T19:45:52.956170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:45:52.956170Z digest=sha256:cd80a2414b75ea946c56bb25c8f2312825cdacb3cd4f73c938b6a3b00617cf03

Observation 26b9cebe-850f-4bff-b75f-82e4460c81a5 · inbound

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor cites this paper.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:35.255074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:52:35.255074Z digest=sha256:2ba0b04e1649e7ee204b4f0b6d7b862775eb1e59f59f13571d8623df96c9ef8f

Observation a5b53bb2-96bb-493c-b041-ddd1971f8bb5 · inbound

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model cites this paper.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:11.164532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:11.164532Z digest=sha256:a315b7440e0282ae28c89497a73fc9ce76d5c08b0ba34b9c23ecbec78d742e60

Observation b7d879fd-ff47-4680-a577-3c7659f9c642 · inbound

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models cites this paper.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training

Reference 135

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:09.056933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:09.056933Z digest=sha256:4a2fc4812f772d3f2aab07c7cadd8fa98ca0c6717de6855fd959ca685824e866

Observation 340c3a01-ac8a-4fda-9dd5-eb288decbba2 · inbound

MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning cites this paper.

MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T05:36:20.191131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:36:20.191131Z digest=sha256:c6cc221394208998d072a191cc7607c009769940a85160d2564dfc016ce5bb17

Observation bb22c86a-fb1b-4d97-8a4e-4a83edbcdedb · inbound

R3G: A Reasoning-Retrieval-Reranking Framework for Vision-Centric Answer Generation cites this paper.

R3G: A Reasoning-Retrieval-Reranking Framework for Vision-Centric Answer Generation MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-16T10:47:45.490874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T10:46:17.411843Z digest=sha256:f6653a3824be20f3b34681e1e520ef210a57664160f6892b114f1830860b8af0

Observation 17f2b8d2-d76c-4729-a0dd-2ed2f6483cb5 · inbound

R3G: A Reasoning-Retrieval-Reranking Framework for Vision-Centric Answer Generation cites this paper.

R3G: A Reasoning-Retrieval-Reranking Framework for Vision-Centric Answer Generation MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T08:13:51.875404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:13:51.875404Z digest=sha256:7f212d78d86cc3e2f3d52407daa3dfeaa7c91206a559e0348cbde615f83bc471

Observation 9e47458c-e6e4-484f-bcf0-2aff0da3f539 · inbound

Back to the Barn with LLAMAs: Evolving Pretrained LLM Backbones in Finetuning Vision Language Models cites this paper.

Back to the Barn with LLAMAs: Evolving Pretrained LLM Backbones in Finetuning Vision Language Models MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-16T04:09:36.761640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T16:07:13.367863Z digest=sha256:30fe9755ba31e36229ea048b2e43946be0c7a0e31df208860dd0dfb8560a70ed

Observation 5f442a76-4533-45d8-b2b6-6159d6599e5b · inbound

MixAtlas: Uncertainty-aware Data Mixture Optimization for Multimodal LLM Midtraining cites this paper.

MixAtlas: Uncertainty-aware Data Mixture Optimization for Multimodal LLM Midtraining MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-16T04:09:36.761640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T20:07:29.544613Z digest=sha256:6c2a6ec77e5d01e7af09faebf2e77152816ab2b13dbecbfea4f9eca6267024e4

Observation df5b6639-7c56-4154-ab02-f94b25de8d52 · inbound

Compared to What? Baselines and Metrics for Counterfactual Prompting cites this paper.

Compared to What? Baselines and Metrics for Counterfactual Prompting MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training

Reference 89

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T04:09:36.761640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-09T19:02:46.991897Z digest=sha256:678b60500693510873b2faea3b86505fbfd4e79a01c77dd86aba114b326160a1

Observation e6c0c0b3-0151-4b69-b23d-4747d0d5f1bd · inbound

20/20 Vision Language Models: A Prescription for Better VLMs through Data Curation Alone cites this paper.

20/20 Vision Language Models: A Prescription for Better VLMs through Data Curation Alone MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-16T04:09:36.761640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T02:52:43.674969Z digest=sha256:6027b654a507f9efa947c5fe937def52430c5e14382694aac63cfa4688e88b0c

Observation 2546657a-2e04-4ad7-84ff-2aa0687b3abb · inbound

20/20 Vision Language Models: A Prescription for Better VLMs through Data Curation Alone cites this paper.

20/20 Vision Language Models: A Prescription for Better VLMs through Data Curation Alone MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-16T04:09:36.761640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T21:28:37.680681Z digest=sha256:c76ec109ab7cf4d9bb182a5bfce931bb0ffb56ec986e5e6f72729b203822b7f7

Observation 3c4cf4c1-bc82-485c-a58d-18c69e8190a3 · inbound

An Exam for Active Observers cites this paper.

An Exam for Active Observers MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T21:12:05.997836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:12:05.997836Z digest=sha256:38876323869b30949cad0eb7d5f42776781d18b59df0c71238e2d70c65d762b2