Pith. sign in

Paper Citation Record · LEDGER

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models

As of 21 August 2026, this Paper Citation Record lists 100 of 104 outbound references and 8 inbound Pith citation observations for arXiv:2502.06788.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.06788 v2

Coverage vector

measured 100 of 104 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T14:25:56.014437Z

measured 108 of 108 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:23:47.528488Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-12T18:51:16.046798Z

Reference resolution

100 of 104 outbound references displayed

  • verified exact0
  • verified fuzzy10
  • unresolved90
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation dca388f7-13b1-43f6-bfb8-943b93ac96ec · outbound

This paper cites Menick, Sebastian Borgeaud, Andy Brock, Aida Nematzadeh, Sa- hand Sharifzadeh, Mikolaj Binkowski, Ricardo Barreira, Oriol Vinyals, Andrew Zisserman, and Kar ´en Simonyan.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Menick, Sebastian Borgeaud, Andy Brock, Aida Nematzadeh, Sa- hand Sharifzadeh, Mikolaj Binkowski, Ricardo Barreira, Oriol Vinyals, Andrew Zisserman, and Kar ´en Simonyan

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.538636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.538636Z digest=sha256:8fbc28f73e9396849931a117b9ebb827d8700f885fa42d35d21d8712b5de46df

Observation 67ade161-fe1e-4ff1-8266-153d126342bc · outbound

This paper cites The claude 3 model family: Opus, sonnet, haiku.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models The claude 3 model family: Opus, sonnet, haiku

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.543179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.543179Z digest=sha256:69ae84d854cf1061f4ae242a7d5519e3350fb2c64a082128cff100f16f2fedc8

Observation ad2f3fe8-dd74-4490-89b1-f0c28c67c4e4 · outbound

This paper cites Qwen Technical Report.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Qwen Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.547154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.547154Z digest=sha256:ef2363c8aa25f7678df15d628c893664383db6c46fd421bff1f1bf328d6cd297

Observation b58348f8-41d7-4aef-afbd-a9c67865d6bc · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.551753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.551753Z digest=sha256:a54759ffe7cd7c9e3e555099d6538c47861919a48c04e68e53207424b80b9c83

Observation ce6090de-bcaf-41a7-b85b-0c90eb67a69b · outbound

This paper cites Vlmo: Unified vision- language pre-training with mixture-of-modality-experts.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Vlmo: Unified vision- language pre-training with mixture-of-modality-experts

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.555696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.555696Z digest=sha256:b38cfe45413a7f596aaf0aaa44281e71e8a2af07a467e354cb74e97600c294d0

Observation c762701b-d418-48a4-bcfe-4c86aab26ea1 · outbound

This paper cites Introducing our multimodal models, 2023.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Introducing our multimodal models, 2023

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.559567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.559567Z digest=sha256:0212df14c17ae8e728d953e0e32754212e49208440f1de08464701e2b98e2fba

Observation 221102a1-281e-48c6-bda6-87361a5eb6d7 · outbound

This paper cites PaliGemma: A versatile 3B VLM for transfer.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models PaliGemma: A versatile 3B VLM for transfer

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.563893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.563893Z digest=sha256:d63aa8fc8facabd5d34c1117647192b308bfd750d86a0f14d6d90726b9cde812

Observation 5432cfb1-ca45-4f38-9fb5-61c215c0d961 · outbound

This paper cites DeepSeek LLM: Scaling Open-Source Language Models with Longtermism.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models DeepSeek LLM: Scaling Open-Source Language Models with Longtermism

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.568135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.568135Z digest=sha256:000563d0bea0da9edce95d6c10c267c8736b23dec5f135e95a1996a496803b26

Observation 8ecb0bf6-542c-4f7b-a0ad-00c64a286e54 · outbound

This paper cites InternLM2 Technical Report.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models InternLM2 Technical Report

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.572341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.572341Z digest=sha256:2031953f37e58bb44d81eb68bfb3b9bb3f5deb19792b27f2f3190571f47a7cd4

Observation 80884fde-78f9-4ef3-86fa-a29d904c4544 · outbound

This paper cites Emerg- ing properties in self-supervised vision transformers.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Emerg- ing properties in self-supervised vision transformers

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.577160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.577160Z digest=sha256:420a959b942d565c7c6c618f4da4a8df27b7256a8c3309bf7c30c14bb351f460

Observation 7eafa05a-2662-4281-8b72-ade0241cb59e · outbound

This paper cites ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.581404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.581404Z digest=sha256:2110fa685a755009c40f43906177d68d039cffbd3ba0896e3ede73bfda805221

Observation 8e6df3dd-dabc-4144-b251-8dbbfc64d53f · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.586228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.586228Z digest=sha256:98958e934e1f0457e4eb65f1fd4c2f5f41a8349085f0fab16ecd6aa28ac55430

Observation d783e8fa-cc1c-4ef8-b4d6-8b90a045d9f8 · outbound

This paper cites SOLO: A Single Transformer for Scalable Vision-Language Modeling.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models SOLO: A Single Transformer for Scalable Vision-Language Modeling

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.590338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.590338Z digest=sha256:317c6bae1d89a7656d9c3d5a7221ae10dec66ba0f4f168ebd4b9927ef38c4a33

Observation 6c10a6ea-6332-4215-ab54-882ef11255a6 · outbound

This paper cites InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.594875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.594875Z digest=sha256:0538c80ae9139ccd55e2ad5314d2d9b3683cab0b25af4b97781be4eb669c78e6

Observation 7bde31ad-577c-4d48-bd9e-090266c3d76a · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.599372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.599372Z digest=sha256:452136fb7a7bf763ff0bcb0a27b6771ad466d83d1fd46f4519f3a43efde8a623

Observation 48b8cc72-30ba-4ce2-9fe5-9d10bffefb69 · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Gonzalez, Ion Stoica, and Eric P

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.604025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.604025Z digest=sha256:dce81dc11b2b85797b03c1d9dd35d161a2981254c12acfcf557ad2994e6a4c3e

Observation 629ec753-3b7b-4d9f-9569-a92d02f31f2c · outbound

This paper cites an unresolved cited work.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.608163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.608163Z digest=sha256:af6b923f7ef543f246a71497d7f6049779f523a37ed9c87805c50dd26711fc61

Observation 88d190b8-39bb-458c-b805-d15a854defa4 · outbound

This paper cites an unresolved cited work.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.612376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.612376Z digest=sha256:fde1b7ca7d65377879be921834d9de05bedd284f2d6c225f8c705b7dc473ab4f

Observation ae6bffc1-ee39-4e33-abcb-314b493cca86 · outbound

This paper cites Unveiling Encoder-Free Vision-Language Models.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Unveiling Encoder-Free Vision-Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.616853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.616853Z digest=sha256:204ef86faba099d80b63e22a418c4cd54370a1827fd6537c01b11997fe1a13b0

Observation 4b42a40c-3e5f-4474-b53c-e0a398311143 · outbound

This paper cites Unipt: Universal parallel tuning for transfer learning with efficient parameter and memory.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Unipt: Universal parallel tuning for transfer learning with efficient parameter and memory

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.626452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.626452Z digest=sha256:5fd905b5e4c68c7bd9edf71708b9d7105c8ad8c39f5b8111f513a9501b3788f5

Observation 8c141db6-6e05-4743-b29e-5de3f755d9b9 · outbound

This paper cites Sherl: Synthesizing high accuracy and efficient memory for resource-limited transfer learning.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Sherl: Synthesizing high accuracy and efficient memory for resource-limited transfer learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.631027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.631027Z digest=sha256:bd151d752b22bdf1039507906c8b95832831c0476f1c4991260e5bd5f9400e69

Observation 8df1a290-7bc9-4313-b547-3fd70e2a5dcd · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models An image is worth 16x16 words: Transformers for image recognition at scale

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.635179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.635179Z digest=sha256:04e7ef339d50010c6b01e44609ab82b52df18f1dbc7550f353132f7474e39043

Observation 8f920d46-2e08-450f-9e60-25503d29d0d0 · outbound

This paper cites Taming transformers for high-resolution image synthesis.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Taming transformers for high-resolution image synthesis

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.639618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.639618Z digest=sha256:0ab8621e02c7962566fa151a1351b608bcaabf4eee56c28bb4da870d0df49240

Observation 73857e29-a43c-47da-9f67-1ce4521106e6 · outbound

This paper cites EV A: exploring the limits of masked visual representation learning at scale.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models EV A: exploring the limits of masked visual representation learning at scale

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.643640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.643640Z digest=sha256:365f9938a6a95033f744ea589dff01bec16d65f40a8542ea74072bbe51617033

Observation ffb0a214-e505-4402-b316-4916b8ea9293 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.648170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.648170Z digest=sha256:0dd538a68e0c3bea3080099ee57f64250d0bac6f1676d62647f179aa83c3b42e

Observation 1d493f05-95dd-42c0-bdb9-5841d61728f6 · outbound

This paper cites Dat- acomp: In search of the next generation of multimodal datasets.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Dat- acomp: In search of the next generation of multimodal datasets

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.652587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.652587Z digest=sha256:c101bd06c5aed2519c852079ffe86e97e0d476d407410168f3806e018fbcb74c

Observation 1a0e6a2d-030f-41ad-bafb-7a18b6d46781 · outbound

This paper cites LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.656727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.656727Z digest=sha256:91c493ad678d6c560cc60a5b03096f076e534a46538c84fc7f562e78759f7c77

Observation 95f8c69b-3394-4235-866e-88c90fa637d6 · outbound

This paper cites Making LLaMA SEE and draw with SEED tokenizer.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Making LLaMA SEE and draw with SEED tokenizer

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.661080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.661080Z digest=sha256:1d12515d5b82a650f6ce4599c0fd905e03d4754d4774555bb70a46b40f94b7c5

Observation 51840bca-4d2a-4d65-ae23-6428312a90f4 · outbound

This paper cites Infinity-MM: Scaling Multimodal Performance with Large-Scale and High-Quality Instruction Data.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Infinity-MM: Scaling Multimodal Performance with Large-Scale and High-Quality Instruction Data

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.665218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.665218Z digest=sha256:960b4b496e4763b67374442eb1332e00d48d8a01b17bc45eaf32771fa743eaa5

Observation ecb5431c-b222-4c9f-85f8-60a5fe5f06a1 · outbound

This paper cites CogAgent: A Visual Language Model for GUI Agents.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models CogAgent: A Visual Language Model for GUI Agents

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.669886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.669886Z digest=sha256:ad0a2e87e2308f1876cc2699c7105e806361caed4e2448406af761e321d1a912

Observation 4ee0b32f-9466-44de-912c-0d4252ff1857 · outbound

This paper cites mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document Understanding.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document Understanding

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.674294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.674294Z digest=sha256:0fe72e45320ad43015b299d6a51d7d9d3fcd28ad698c82a9bb7f453da6c297c4

Observation c7b681be-1ef4-4d70-9983-78e3b19a678b · outbound

This paper cites Hudson and Christopher D.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Hudson and Christopher D

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.678998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.678998Z digest=sha256:1c19926cb322d9c914ba343f3576710f72ce740f81c6985627dc69c27f0879b7

Observation 46e6663a-f404-4645-a07e-9ba0fa342654 · outbound

This paper cites Introducing idefics: An open reproduction of state-of-the-art visual language model.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Introducing idefics: An open reproduction of state-of-the-art visual language model

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.683411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.683411Z digest=sha256:2fa58bad7766f0afecff22f8b9bdd6f9cad66e361c28574484af109e092e334d

Observation 05533886-e4eb-4811-bf0a-ed1473626e14 · outbound

This paper cites A diagram is worth a dozen images.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models A diagram is worth a dozen images

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.688325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.688325Z digest=sha256:f5a25705b6539aed05c4dce075b49ab438d5cad96f9cb72921387f26941b3f99

Observation e9e3624d-5fa5-40f2-8600-4f2a71cbff7c · outbound

This paper cites Kingma and Jimmy Ba.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Kingma and Jimmy Ba

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.692617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.692617Z digest=sha256:78c0a79e4a7ad61278be8a9a5ebd61d9ce668da22d0e9ed42dc42bfb3f6a8563

Observation e9c99db6-6e45-492e-93ce-245dc3ed24cb · outbound

This paper cites Segment Anything.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Segment Anything

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.696864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.696864Z digest=sha256:5cb8815eab4eecca946794bb1c129ff358ca3b918e8e36fe2fdd567d65b202bf

Observation 37d774e5-ce92-4d38-b783-0bad9ad04198 · outbound

This paper cites The Open Images Dataset V4: Unified image classification, object detection, and visual relationship detection at scale.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models The Open Images Dataset V4: Unified image classification, object detection, and visual relationship detection at scale

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.702211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.702211Z digest=sha256:4178574aa6447dd1d8d11393e8ed3880b506fa71c44a46db2cc90c3fdaa7cf7d

Observation f78ef67f-3e12-4473-8176-d60d129373ad · outbound

This paper cites Building and better understanding vision-language models: insights and future directions.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Building and better understanding vision-language models: insights and future directions

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.706584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.706584Z digest=sha256:892e2b2045a4f0c8d6e53e86839e13fbcf997eccb3c7b5a2b36d911ddcfee292

Observation c345cac1-3c26-484d-8599-56a3de933b85 · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.711133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.711133Z digest=sha256:1745e9f40741eb78dce944f461f8a28028dd21b7358b262ddb5f3b4d858848c3

Observation d68e8ec1-3498-43aa-bcbb-d638cf8dffe7 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models LLaVA-OneVision: Easy Visual Task Transfer

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.715605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.715605Z digest=sha256:2685180cb38ca1fac298803a0dfb454ddcfe99ca70e89b90002041f1db24c375

Observation 8659e668-aea5-46f6-af9e-87a2a11fa040 · outbound

This paper cites Aria: An Open Multimodal Native Mixture-of-Experts Model.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.720298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.720298Z digest=sha256:7d533ed2d8c193cae9f61ef93374b883462426517e395fda87adc78e11b65001

Observation 434c5d3b-7449-4547-85b9-5bdcaf8734f3 · outbound

This paper cites an unresolved cited work.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Unresolved cited work

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.725842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.725842Z digest=sha256:3492e786e285d1aa6534ea0c208681d63f0df87116555c798b93249aef45c3b0

Observation 04183272-f961-4012-b78e-14344ab02465 · outbound

This paper cites an unresolved cited work.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Unresolved cited work

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.730026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.730026Z digest=sha256:f1c2a70f334fd0289834234d4b2e241bc9c5b23a5c27eed4d2ac59eaaaf072de

Observation cfa8f363-afbf-4b8f-b7c7-016e390096fb · outbound

This paper cites mc-beit: Multi-choice discretization for image bert pre-training.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models mc-beit: Multi-choice discretization for image bert pre-training

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.734219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.734219Z digest=sha256:551d4217bbdca15c451a94e6451d8db5130cfa45751540c89fc56ea9cc4366d7

Observation 30286c70-f0d9-41bc-bf97-62ed05f26b43 · outbound

This paper cites ImageFolder: Autoregressive Image Generation with Folded Tokens.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models ImageFolder: Autoregressive Image Generation with Folded Tokens

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.738517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.738517Z digest=sha256:902c65e9ebcebf3ee70028d021307bf5f4c8a17b97373ce15057612ade642d11

Observation e5b52559-e7e7-42ee-8c3b-0be5be797755 · outbound

This paper cites DenseFusion-1M: Merging Vision Experts for Comprehensive Multimodal Perception.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models DenseFusion-1M: Merging Vision Experts for Comprehensive Multimodal Perception

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.743104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.743104Z digest=sha256:d9441d6ddcf3e0a883cd0c337950b8bdb734d0bfa658ab20897d626808777ecf

Observation 99ea5c4b-be71-47f1-a41a-abb5dfe8ed82 · outbound

This paper cites Evaluating object hallucination in large vision-language models.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Evaluating object hallucination in large vision-language models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.748054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.748054Z digest=sha256:7c0628dc71c529b70e4ae1c9055dce28499e22285cdb50a1c5269828b42c0f1c

Observation 31365a7f-c765-4263-a232-b30f7ed0088d · outbound

This paper cites Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.752574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.752574Z digest=sha256:12f90274e01eab37e1a009bd4a44a6ccadf6f5ed6232ac996505df7d298ec048

Observation 9c745885-3cd2-4023-9e08-f5804c1b595c · outbound

This paper cites MoMa: Efficient Early-Fusion Pre-training with Mixture of Modality-Aware Experts.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models MoMa: Efficient Early-Fusion Pre-training with Mixture of Modality-Aware Experts

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.756888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.756888Z digest=sha256:84e97df114ab9fea8c6567d381df924059b24a02820c533573690241444dfab2

Observation c0f36955-c614-419b-981f-7ea9a73f4a56 · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Improved Baselines with Visual Instruction Tuning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.761283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.761283Z digest=sha256:046366206081469794873e3bf2eac9072a4367efb8d68c9e7494608ce4ac305d

Observation fd7ef618-850f-410d-8e22-7d8e3e76e45b · outbound

This paper cites Visual instruction tuning.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Visual instruction tuning

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.765855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.765855Z digest=sha256:65d67999c72a4929aa8a446e38dbdf95dac4d4339cf3a686c246a525d328e01b

Observation f033e7b3-34cb-4d64-a059-1ab62a84e988 · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.770240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.770240Z digest=sha256:5e23bd04e3aff09de0f21f29440694887335e0f32dd7b7b66e82e582d502b422

Observation c808b132-449d-4dfc-9c58-196374f34cc6 · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models MMBench: Is Your Multi-modal Model an All-around Player?

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.774516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.774516Z digest=sha256:9712db9fb691f859d5b253616aed0dc09febd6eb8e5b2ff491114745e0452d26

Observation a26125b8-633d-4986-9a9e-0a49ba8c0aed · outbound

This paper cites OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.779074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.779074Z digest=sha256:eca0a8e5e8e8b382a7835be482c73056ae8c65cba6c4b363e3b0e78cd53c795b

Observation c8044e42-29ea-4c47-8829-d402045b0052 · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.783785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.783785Z digest=sha256:01f30896d217ec5a4154cdbdd3b0790a2487fa1e88f183739667657e665b9895

Observation 57650a85-2453-464f-8f88-97e868b7b9cb · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:25:57.351165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T14:25:55.788669Z digest=sha256:41491c18bf0691acb8626c0003f8c4af4bd4b9aa44d7bb195c7ae42d8f669959

Observation 6ce369cc-c52d-4505-888e-59dff9b73e0e · outbound

This paper cites Mono-InternVL: Pushing the Boundaries of Monolithic Multimodal Large Language Models with Endogenous Visual Pre-training.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Mono-InternVL: Pushing the Boundaries of Monolithic Multimodal Large Language Models with Endogenous Visual Pre-training

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.793608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.793608Z digest=sha256:97ec4484c60ce2a097b5f52d65af564eafda230e19d2f902b2d270f2b5192d05

Observation c5785fe5-8416-4bfb-ae01-70935eb28920 · outbound

This paper cites Chartqa: A benchmark for question answering about charts with visual and logical reasoning.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Chartqa: A benchmark for question answering about charts with visual and logical reasoning

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:25:57.335376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T14:25:55.798660Z digest=sha256:c895d869d9f353f8f406f215f886b6f7ae7df7360b623810927c11ffe2f7c7e8

Observation 5e87e700-855e-4fdb-b467-e8950e3bfd15 · outbound

This paper cites GPT-4 Technical Report.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models GPT-4 Technical Report

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.803714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.803714Z digest=sha256:0944d53e256295d5ca915b5aab94a4a4cd0f0dc2842eba881d404c9fff845d2c

Observation cca3d66c-37d2-482a-8e41-0ba9f8c5a2c1 · outbound

This paper cites TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.809314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.809314Z digest=sha256:09396450f80cab51d8c45c9a7af3e2f950bbe22621574f5a5e0f85a5fd678258

Observation f225d90c-9381-4d87-afa2-f12751545a72 · outbound

This paper cites Learning transferable visual models from natural language supervision.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Learning transferable visual models from natural language supervision

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:25:57.319691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T14:25:55.814565Z digest=sha256:958e3226a97c78cf662fb15516401f9d420fe58641cd259e6bac083e7425c571

Observation 75ff3c6e-5086-40d9-8139-5aa60808a861 · outbound

This paper cites Laion-5b: An open large-scale dataset for training next generation image-text models.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Laion-5b: An open large-scale dataset for training next generation image-text models

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:25:57.303982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T14:25:55.819490Z digest=sha256:3186c2508f1e8ab9534f2073926a51d2ad4a9e3ed0af1234f122ba38dd96bfcf

Observation e63c32b6-ad8c-4af5-996e-cc77c8173dbb · outbound

This paper cites Towards VQA models that can read.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Towards VQA models that can read

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.824630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.824630Z digest=sha256:70e6e451270ce39627c519d730a62fff7527fbc7be5f425f1daa6fff3afc21ac

Observation 4c5760b1-99d1-4b68-8f66-7d34184d09b8 · outbound

This paper cites Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.829308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.829308Z digest=sha256:88bb102fa786652360292c6cd07b3b685c585c37129315e6fd980bda779f388d

Observation 93fee7ad-4953-479f-a040-2b380c6b26f9 · outbound

This paper cites Generative Multimodal Models are In-Context Learners.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Generative Multimodal Models are In-Context Learners

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.834228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.834228Z digest=sha256:8ba11532411f14349cedd5c475eb6e01b5f3bff415af68e3fd2cb834252236b3

Observation 291b4ece-9106-4d66-b709-a4847c6ec119 · outbound

This paper cites EVA-CLIP: Improved Training Techniques for CLIP at Scale.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.839354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.839354Z digest=sha256:a146bd7ccca9a2c8bf638a95eac100dae933f0fef35d48cf7fe9c725e7d3871b

Observation 896d60e1-29b6-4d45-a5fe-313b2a6f62cd · outbound

This paper cites Emu: Generative Pretraining in Multimodality.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Emu: Generative Pretraining in Multimodality

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.844863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.844863Z digest=sha256:e744bf04e2d3718052fbb6cc0590c0a03bf0eb26c2be304df21740235d819303

Observation 20477537-a323-4c31-956e-b3057dc27188 · outbound

This paper cites EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.850157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.850157Z digest=sha256:e5f4a584bcc5ed8c7a1a0e751ca8bb9c275db85d53c8470f1181ce178246442c

Observation 6fae62d9-4f4b-4d6d-b363-2743ff019593 · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.855575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.855575Z digest=sha256:f053943b1a17f86148c9a5e6680f5a8181b35e1b564254cb18745c4fa4abb3f8

Observation 5032d9f8-67c8-4a72-b8b8-0bd64789f976 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Gemini: A Family of Highly Capable Multimodal Models

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.861005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.861005Z digest=sha256:a26f77ab5881808d340049b3624f84db242e409d67cf30d1c64afdc4c2e1f547

Observation a34eda86-776c-44fa-868b-a9e1fb2913bd · outbound

This paper cites Internlm: A multilingual language model with progressively enhanced capabilities.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Internlm: A multilingual language model with progressively enhanced capabilities

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.866397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.866397Z digest=sha256:d8135f705bfa76e98bdbaafcc779c281fc93b8340d2fe94e02d29435be6b855b

Observation fe8f12eb-f70a-43da-986b-e18ff6f57368 · outbound

This paper cites Llama 3.2: Revolutionizing edge ai and vision with open, customizable models, 2024.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Llama 3.2: Revolutionizing edge ai and vision with open, customizable models, 2024

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:25:57.268277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T14:25:55.871639Z digest=sha256:1d804c3cb240dae6a152b55897094e5c44a3da86eb74bce7ebc071aa4b33d91b

Observation 9332fce2-4319-4647-8547-2703bc9313cd · outbound

This paper cites Qwen2.5: A party of foundation models, 2024.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Qwen2.5: A party of foundation models, 2024

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:25:57.250633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T14:25:55.876978Z digest=sha256:4b9cc14c8c040ab62e77de428de1e121e2aca6b26cdf95a0e63f9fe4a6e2de64

Observation ca957625-1e55-4d79-9348-2ac58e780253 · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.882208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.882208Z digest=sha256:aaa6df10d572c8f31c65dfe48fe45712d72b3b56158b6a643192a5d572a2ecff

Observation d6f37dae-d47b-4967-bf9c-903653c1beda · outbound

This paper cites Eyes wide shut? exploring the visual shortcomings of multimodal llms.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Eyes wide shut? exploring the visual shortcomings of multimodal llms

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:25:57.233562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T14:25:55.886692Z digest=sha256:3ba219837a97d20c537a3a763541e6b03cb2d6b7695cdba97a8f727d146b010b

Observation 3f85f9b3-f2d7-426a-b4d0-b354b2599d2c · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models LLaMA: Open and Efficient Foundation Language Models

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.891210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.891210Z digest=sha256:5f41fe1a3274d6842734cd1e7caba5baa5b40e73ea50716b36883673e0c6b8ac

Observation 02f7baf5-9b60-4052-bdd7-c7ededa5d4f7 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.895582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.895582Z digest=sha256:d8060e7e215b0d6a7eaf18c630ed385585df59474117f9a6df4fb85b9aa8dc3d

Observation 66893499-3a31-4a5b-a2ee-e32626443428 · outbound

This paper cites Neural discrete representation learning.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Neural discrete representation learning

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.900288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.900288Z digest=sha256:9568f7bd13202e3c36acac412067847a561f65946942bc9b08efa252bc1a270b

Observation cc369af7-71b0-4de6-a412-14689454093b · outbound

This paper cites Gomez, Lukasz Kaiser, and Illia Polosukhin.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Gomez, Lukasz Kaiser, and Illia Polosukhin

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:25:57.205036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T14:25:55.905656Z digest=sha256:941898d127da6f83b7e1dc5b431ff44a159356efb8fa5c06408cff972e9da86f

Observation c6280041-f4b1-4ef6-ae4a-877331f0c2b9 · outbound

This paper cites To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.910395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.910395Z digest=sha256:f0266f9a52a0bf5728ef62b86ccb8f0f25fbd1907bc32bc4ba95d7ea89e5346d

Observation 929ff2b0-581b-4122-91c6-67f012e7033a · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.915568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.915568Z digest=sha256:a50c814d2679ed5c84ebe19963f2d8b030a4f8a2b2e50fdbbf2c59e7e08e9d8c

Observation fdbfd165-2d57-4b32-a820-fefd637ea767 · outbound

This paper cites Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.920632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.920632Z digest=sha256:dc9a5f3d39a51e72c378a4d3856df5e8772aaa355cfc3c1a1df9b965cf2819d2

Observation 706d3749-3922-494f-bafe-2a68d02fd51d · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Emu3: Next-Token Prediction is All You Need

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.925834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.925834Z digest=sha256:a3b347f8302392668babff35af156c6aa9143d9a16c8c110b4fa31a22ca0230d

Observation 5380f8f9-1095-4541-8124-9df08d2f2b60 · outbound

This paper cites Mio: A foundation model on multimodal tokens.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Mio: A foundation model on multimodal tokens

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.931088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.931088Z digest=sha256:6e15d70f395eed708e06533cb0f42f6c37f03d4c99411bec2836a275d2af4a31

Observation df984a12-468b-48e4-a4dd-cb0e61fe49b1 · outbound

This paper cites Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.935935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.935935Z digest=sha256:4384252e0903409941e1fd8871c28d6fc1d3b822369bf3fa9bcd6af34b3edb45

Observation a1f53753-39f3-43f5-8c53-629bcfa2b0ee · outbound

This paper cites VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.941194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.941194Z digest=sha256:1f002fcdc3f5e698c39584aaa68389fb0dc744c7fb5e7b9f166fa6d929516681

Observation 065aea9b-d63f-44d1-9a1c-4ab918e95172 · outbound

This paper cites Grok-1.5 vision preview, 2024.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Grok-1.5 vision preview, 2024

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:25:57.188693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T14:25:55.946792Z digest=sha256:e80a926301b255c806073bda1c167bcc4f628474ed4c374a26a62bf0d38807ca

Observation 81412cb5-a096-4187-92c2-ff89caf42563 · outbound

This paper cites Show-o: One Single Transformer to Unify Multimodal Understanding and Generation.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.951584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.951584Z digest=sha256:6638459e57608c40df775e879b8b28a1fae05e77bbf4921d07b439edfd4d7b95

Observation c23031d0-36ec-4096-8cf9-467741caeef5 · outbound

This paper cites MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.957533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.957533Z digest=sha256:2f5d732809618ee63cc940a31cded940de7bcd902443f93049cc98657ebc8ae6

Observation c73ca218-7d4b-46d1-83c3-fb6010d0fa22 · outbound

This paper cites LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.962910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.962910Z digest=sha256:835cfa42b0adadbc7969b8cda4809ffc9decdefccec6903ac99f4aaa279032cb

Observation 261b581e-efb5-4d7c-b5d9-71b806c7170c · outbound

This paper cites xgen-mm (blip-3): A family of open large multimodal models.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models xgen-mm (blip-3): A family of open large multimodal models

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.968319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.968319Z digest=sha256:39af80cc4460153af09d1a22380bf15780920fa898a33697f412159ab9a7e250

Observation 52d45ce1-dca6-48c6-ace0-54ccde05b893 · outbound

This paper cites Qwen2 Technical Report.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Qwen2 Technical Report

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.973291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.973291Z digest=sha256:53c9ae6ca3472e9bf4d4fe0ea79c5006f0b168c60c95f89fae79e86645a9dd6e

Observation 4cee62dc-3aea-45e0-9475-11a758954187 · outbound

This paper cites The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision).

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision)

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.978546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.978546Z digest=sha256:600f5fbfdd2473863e513f0b564ce46a19905996cd75dcab43a49c8b44b7ae34

Observation 09ba62fb-7ffe-4938-9208-8263578eccf7 · outbound

This paper cites mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.983562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.983562Z digest=sha256:8e6d8621612ea65a3dad12369c8b9edad59c712ce51a7b46f7a1c80f2b8845a1

Observation f231966d-8bb1-44e8-a38e-3b12ef6530d4 · outbound

This paper cites mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.988746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.988746Z digest=sha256:4bc9ed80a9927461c4c82aae534f0b392d2ba51e9aaa7f7aaaebac10978d361e

Observation 33569c35-748e-4081-9cc9-a6372427a5fe · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.993984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.993984Z digest=sha256:8f4cac2bed716d7fe4d6312794fa5e2cbd2a3c728b8321c1a573e66ce13fc001

Observation 69b9462a-35a9-4365-96b6-38dbc6de95b8 · outbound

This paper cites MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.999125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.999125Z digest=sha256:8397dbbe44fab721d18b004601f24960aa3681a3fb850d133572bd929b086a41

Observation 44cdbe71-d721-4ea0-a560-daa005694dc8 · outbound

This paper cites Sigmoid loss for language image pre-training.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Sigmoid loss for language image pre-training

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:25:57.171985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T14:25:56.004175Z digest=sha256:a638424881f160451496c3e4f6525e5ddea5fa2a1d3857f8beab48b57e66a4ac

Observation 4c7259e3-331b-4aaf-866c-d932f3a74305 · outbound

This paper cites AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:56.009237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:56.009237Z digest=sha256:08b5505b999c0e5ec10696fd9b36ffbd60c0685df9da2a4dd66171e2a96d5d71

Observation 6dde0cbd-6c00-4a84-a005-0ae1a1f197c0 · outbound

This paper cites LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:56.014437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:56.014437Z digest=sha256:91448ed071336680a746edd038464e6beb0700aa636d35639d6d69929daf739b

Pith citing papers

Observation 738bf878-8fe3-40a2-9a84-d4a0079b3f59 · inbound

SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics cites this paper.

SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics EVEv2: Improved Baselines for Encoder-Free Vision-Language Models

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:22:37.318223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-11T21:22:36.902119Z digest=sha256:56e5b243d4ed265d855e716bf46c0dff5b4283e8e9d9df67d3686ba2d8af6317

Observation debf2a9a-6d76-491b-8c8b-dd7dab277001 · inbound

Dense360: Dense Understanding from Omnidirectional Panoramas cites this paper.

Dense360: Dense Understanding from Omnidirectional Panoramas EVEv2: Improved Baselines for Encoder-Free Vision-Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:47.528488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:47.528488Z digest=sha256:3384cb2998219e693d187a552e284c5e17733f44848f09515a8355b86b9e14b0

Observation 672d1c36-d965-4122-88c2-c3a0e1057c6a · inbound

Show-o2: Improved Native Unified Multimodal Models cites this paper.

Show-o2: Improved Native Unified Multimodal Models EVEv2: Improved Baselines for Encoder-Free Vision-Language Models

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:51:16.049795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-12T18:51:15.428692Z digest=sha256:7a21b49d5bf9666107fba453f5a639dfef0a628ae399a93d692bc30b2140c247

Observation d54a2bed-4601-40be-852d-3d6a8fecd470 · inbound

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs cites this paper.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs EVEv2: Improved Baselines for Encoder-Free Vision-Language Models

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:09.482217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:09.482217Z digest=sha256:b75fa156ddba5b8d1183837275eaf035eccd407c2f4f397b327df701f03f59f5

Observation b134bd5a-39fd-43cc-9b02-961aba8666fb · inbound

NeoBabel: A Multilingual Open Tower for Visual Generation cites this paper.

NeoBabel: A Multilingual Open Tower for Visual Generation EVEv2: Improved Baselines for Encoder-Free Vision-Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:26.683136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:26.683136Z digest=sha256:f0ad73e63ba0d79de0043cf940b91a37c06f4650ad5816862a53ba773ce40a67

Observation fee704c9-80bf-4ae4-aace-ac300f0c8736 · inbound

Regularizing Subspace Redundancy of Low-Rank Adaptation cites this paper.

Regularizing Subspace Redundancy of Low-Rank Adaptation EVEv2: Improved Baselines for Encoder-Free Vision-Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T13:23:33.793178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:23:33.793178Z digest=sha256:f4d3141cc8010f09a76eddf720a0d5955bd30cdd4a92d9dbfc662b90ae2dbb27

Observation 6313292f-a74b-4422-952c-043a5e25126a · inbound

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts cites this paper.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts EVEv2: Improved Baselines for Encoder-Free Vision-Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T20:04:49.342135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:04:49.342135Z digest=sha256:39277d017294f3edaada4a700147c450c787db850ce63dcf23bbde1b8c420299

Observation 3c7b937c-c260-4d00-b0f2-9190e4cf32e9 · inbound

Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture cites this paper.

Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture EVEv2: Improved Baselines for Encoder-Free Vision-Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T11:40:21.474439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:40:21.474439Z digest=sha256:2cc854a826f4257001496c50bdcc8cd47cb3637de30634cb01c5115262a82197