Pith. sign in

Paper Citation Record · LEDGER

OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning

As of 8 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 4 inbound Pith citation observations for arXiv:2509.01644.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.01644 v1

Coverage vector

measured 62 of 62 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T12:22:35.866665Z

measured 66 of 66 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-01T07:35:07.825460Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T07:35:28.745644Z

Reference resolution

62 of 62 outbound references displayed

  • verified exact1
  • verified fuzzy30
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 04be0f62-5e78-4fb7-a8dd-7e5e6a24ee74 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning Flamingo: a visual language model for few-shot learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T12:22:31.971104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:22:31.971104Z digest=sha256:b08da2662c1f763d87a6bde3a735f41fca7162c1a7717389b45f07eb4e53f107

Observation 9aac679f-1a31-42c2-92f5-b743f973cc60 · outbound

This paper cites Lan- guage models are few-shot learners.

OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning Lan- guage models are few-shot learners

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:22:38.058829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T12:22:32.052551Z digest=sha256:96e14c96113df312107e825cedef7d7891c3c205b51a519dc6db013d5632289e

Observation 5add7096-c187-495c-85b7-1c4837441c51 · outbound

This paper cites Open-llava-next: An open- source implementation of llava-next series for facilitating the large multi-modal model community.

OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning Open-llava-next: An open- source implementation of llava-next series for facilitating the large multi-modal model community

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:22:38.013965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T12:22:32.154735Z digest=sha256:2c6d1c2473666794722b60a201d3ffa4c5bdbae525a8b961853cd68fef01ccd7

Observation abb71275-78ed-4e94-a184-7080bc75fb7f · outbound

This paper cites Generative pre- training from pixels.

OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning Generative pre- training from pixels

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:22:37.975688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T12:22:32.257335Z digest=sha256:f4a2876d29e96c8026a1787d4dd3526a995a75296198f0b6423879066b850412

Observation 313fd4f8-c736-42c0-8bfc-c223068f7d4e · outbound

This paper cites PaLI-X: On Scaling up a Multilingual Vision and Language Model.

OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning PaLI-X: On Scaling up a Multilingual Vision and Language Model

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T12:22:32.355438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:22:32.355438Z digest=sha256:e1517020a91920bf122f930278de8eb2377b6b7596a1559e29a4d376ae59fbae

Observation 5a38f1a5-bc7e-40b7-adb5-408051470dd2 · outbound

This paper cites PaLI: A Jointly-Scaled Multilingual Language-Image Model.

OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning PaLI: A Jointly-Scaled Multilingual Language-Image Model

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T12:22:32.423076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:22:32.423076Z digest=sha256:6d18a37db6a41cf61f100c7f250505f0c819c3a9b404f6f57fb9eed3d6444f28

Observation 2af4d339-31f2-45cc-9cb4-a95d41a1886c · outbound

This paper cites Virtex: Learning visual representations from textual annotations.

OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning Virtex: Learning visual representations from textual annotations

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:22:37.953645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T12:22:32.526318Z digest=sha256:59b37775a48146dd7ea89e144088633e582f594b40b8c6af09562ac82f0e742c

Observation 487eeaa0-cee4-4056-9ee8-d9d6f8a06f5b · outbound

This paper cites Learning Musical Representations for Music Performance Question Answering.

OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning Learning Musical Representations for Music Performance Question Answering

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-05T12:22:36.803786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T12:22:32.582902Z digest=sha256:6df8ed6a9eb0d1b93fcff963052bc8ae6b26e63ae35dd753fdc07e17ba975eeb

Observation b107f77a-1c8a-42a2-9d10-8475833fc3d6 · outbound

This paper cites Temporal Working Memory: Query-Guided Segment Refinement for Enhanced Multimodal Understanding.

OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning Temporal Working Memory: Query-Guided Segment Refinement for Enhanced Multimodal Understanding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T12:22:32.652497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:22:32.652497Z digest=sha256:c3df8d11f5c6b6bab3440f20b59c8ec0451d32ed40f15b87e9af5fe809c899e7

Observation 505638ea-f1f5-4a45-97f1-1463a38f024b · outbound

This paper cites Scalable Pre-training of Large Autoregressive Image Models.

OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning Scalable Pre-training of Large Autoregressive Image Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T12:22:32.780054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:22:32.780054Z digest=sha256:ca3e2c186480478ab537bf5525727b689145b517ff5bf110b3e120586145238a

Observation 63f53512-35cf-4345-918e-a3e7c47c7c9c · outbound

This paper cites Improving clip training with language rewrites.

OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning Improving clip training with language rewrites

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:22:37.917562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T12:22:32.872014Z digest=sha256:c9dc3100a32744f9b3059463720be10734dc9e18cad8eb690233a07480a3aa77

Observation 2644c2e3-a988-4830-a232-90903decb198 · outbound

This paper cites Data Filtering Networks.

OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning Data Filtering Networks

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T12:22:33.199565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:22:33.199565Z digest=sha256:9e0a23be3e910820955c7e168dd0fd696c38830749e36dd486cdca5a0a44facb

Observation 0318e99b-ce9f-42c5-bb61-293b850ed45c · outbound

This paper cites Multimodal Autoregressive Pre-training of Large Vision Encoders.

OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning Multimodal Autoregressive Pre-training of Large Vision Encoders

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T12:22:33.329919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:22:33.329919Z digest=sha256:66f4c7d5a2fae30adf4487fa56d10e9c7df27d2d3b6f55e47599a54eea6eb641

Observation 941b4be7-1e48-4b85-8b2e-757a4e93f2f2 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T12:22:33.521610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:22:33.521610Z digest=sha256:bf9e634b67e8ecc7edf810a771cc4740fe638835086870f089f4c8c798003529

Observation 98faa756-30e0-4b35-a4e4-807594854e0b · outbound

This paper cites DataComp: In search of the next generation of multimodal datasets.

OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning DataComp: In search of the next generation of multimodal datasets

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T12:22:33.659440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:22:33.659440Z digest=sha256:b514ff1f6c769a3608a69b54ad0bd5496c08e4a08b3c491fcb6b3b89a5f21fa5

Observation 47a14ad9-84c8-4574-90ca-4e2cff0b5ee6 · outbound

This paper cites Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers.

OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T12:22:33.772130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:22:33.772130Z digest=sha256:94cef06074abda6a7c258baef4834ce952e6e72eea39f8f09d3aa8bc39960666

Observation 395ade9c-bbc6-4a3a-95c3-67e5ce3beaec · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T12:22:33.923744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:22:33.923744Z digest=sha256:3faee4a16a666dff878990288d900ee3be51a7b6cb1c4ef358e9d62ef738799f

Observation d9bd0af9-2c21-43c1-9fe0-b36231603e0d · outbound

This paper cites Open- clip.

OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning Open- clip

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:22:37.840545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T12:22:34.066643Z digest=sha256:43d3645c0e96dc8c7460a6059c10e92d81410e9bac7b349d304b11aca39ac6a9

Observation dd03223c-6af5-4d78-8340-399ab782ead3 · outbound

This paper cites Scaling up visual and vision-language representation learning with noisy text supervision.

OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning Scaling up visual and vision-language representation learning with noisy text supervision

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:22:37.798252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T12:22:34.191768Z digest=sha256:99a600d8b195cfb736ef92fb4e25f8e3a2485356f835ae86008dbb70ad3ddce6

Observation e0aff7eb-aadf-4f01-9269-4ea65727829c · outbound

This paper cites Learning visual features from large weakly supervised data.

OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning Learning visual features from large weakly supervised data

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:22:37.759099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T12:22:34.328207Z digest=sha256:87c726799d46895e9d52ae5c1bb22576488a1edb04a6381f1d35ff15eff14853

Observation 931daa9a-01c9-4e79-b00e-68952f45e7d7 · outbound

This paper cites Deep visual-semantic align- ments for generating image descriptions.

OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning Deep visual-semantic align- ments for generating image descriptions

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:22:37.720481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T12:22:34.421647Z digest=sha256:6d097771655a4ef2196fd5d9d5b18c2bd6896899f7b404240be6090722452d4a

Observation 19d4ccc2-4a75-4acd-8614-b093f8c1c5e7 · outbound

This paper cites Vilt: Vision- and-language transformer without convolution or region su- pervision.

OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning Vilt: Vision- and-language transformer without convolution or region su- pervision

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:22:37.689855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T12:22:34.591130Z digest=sha256:fdb301dd10796aa27b5eafcfad40e7aec0ac9f0da0846d73209d6c2c4635d7e4

Observation 12c60015-cc86-464e-9daf-49dda71f9333 · outbound

This paper cites Veclip: Improving clip training via visual-enriched captions.

OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning Veclip: Improving clip training via visual-enriched captions

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:22:37.656536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T12:22:34.737816Z digest=sha256:6423fd4c281ce84f3fb489084b8c3e5f5d361d55d465366469a66c1f5abc6e33

Observation dc92feb3-0176-4b9b-bf84-89e1c99efe61 · outbound

This paper cites Learning visual n-grams from web data.

OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning Learning visual n-grams from web data

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:22:37.629877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T12:22:34.867133Z digest=sha256:e0faeb9270083c2beeb339940acc54e491512d74cbf593560a8cc756612db9d4

Observation 14433b13-cfb1-4179-bd76-b63cca80af55 · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T12:22:34.980214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:22:34.980214Z digest=sha256:e5a3f8d6b50df56a942f1b296b39cb30df1ee127aaf3f24feec54b74b4f09894

Observation 32a9aefd-6dd5-494d-9197-e9b69b23443e · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T12:22:35.099382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:22:35.099382Z digest=sha256:f04db69f82708e7dbd4c305ffa92f78e3e4b2db71794b610de5278da88f0300d

Observation be819195-7e08-4fa5-8a42-58e0aca29630 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T12:22:35.212426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:22:35.212426Z digest=sha256:162b5390a0f83ea0d8bf6e31754f1d20cf983846824a2d4dc64e915ce020d53f

Observation c4b6daa7-0aa3-4079-bc33-fc9a67b7ded3 · outbound

This paper cites Align before fuse: Vision and language representation learn- ing with momentum distillation.

OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning Align before fuse: Vision and language representation learn- ing with momentum distillation

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:22:37.550385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T12:22:35.350031Z digest=sha256:db5152f56a6e0169a79645bca79a442b1a1348741c858d4d67af86bbc3375815

Observation 67f44edf-b97d-46a6-bf67-be7c43d53446 · outbound

This paper cites OpenVision: A Fully-Open, Cost-Effective Family of Advanced Vision Encoders for Multimodal Learning.

OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning OpenVision: A Fully-Open, Cost-Effective Family of Advanced Vision Encoders for Multimodal Learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T12:22:35.377071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:22:35.377071Z digest=sha256:d145bdd49244e91bb75774027f435e667c871b6bb4349ecc3894228372f91326

Observation 2124472e-1761-4fa3-88f0-08854232d528 · outbound

This paper cites What If We Recaption Billions of Web Images with LLaMA-3?.

OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning What If We Recaption Billions of Web Images with LLaMA-3?

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T12:22:35.506534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:22:35.506534Z digest=sha256:8baebb3d0836972070edc31f417983c7f99d45415691cfefe503269573356482

Observation 159157b4-97f0-45f4-a65a-0f524f7c7493 · outbound

This paper cites An inverse scal- ing law for clip training.

OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning An inverse scal- ing law for clip training

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:22:37.517305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T12:22:35.536881Z digest=sha256:9442adadfd12943b140c9d3a21ce8ddc68839162b4c6ec113fda61dda14ba0d6

Observation 8ca4bc67-469d-426d-8095-ec448cac8178 · outbound

This paper cites Evaluating object hallucination in large vision-language models.

OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning Evaluating object hallucination in large vision-language models

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:22:37.483786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T12:22:35.565297Z digest=sha256:31f957b00482d24d3055b36937f4cb71df67ad58da9558558df83c662c0518e8

Observation ccc04d01-7477-4b74-a313-2d1d61bf5473 · outbound

This paper cites Improved baselines with visual instruction tuning.

OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning Improved baselines with visual instruction tuning

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:22:37.446501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T12:22:35.578726Z digest=sha256:e05343d61463622a836ec50c1b34c7fb61e0b384c6984c69be2166e40912102c

Observation be78f29c-2021-4df5-adb2-06593a868d3c · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge.

OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning Llava-next: Im- proved reasoning, ocr, and world knowledge

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:22:37.397293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T12:22:35.588360Z digest=sha256:80d9f9719f6333c8c84dd0f9fda50e3a8ae1c5454719f7c17a63b277d03cde20

Observation 1eb0e4c0-8a68-481b-b4f0-9fd423270969 · outbound

This paper cites Visual Instruction Tuning.

OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning Visual Instruction Tuning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T12:22:35.596028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:22:35.596028Z digest=sha256:622782771aca62864c9cce0e68a7c2453d8233e20cb8458e7c5aa6c79927bc6c

Observation bd5172db-a689-4111-ae86-087c771bf7ab · outbound

This paper cites CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions.

OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T12:22:35.606923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:22:35.606923Z digest=sha256:65dcf4a7101a98013881fdae8607460ab91bfbe4f81a4da8b851e3a72ea3dcd4

Observation e47fd057-702d-4d69-8b35-91d7acde64ef · outbound

This paper cites Ocrbench: on the hidden mystery of ocr in large multimodal models.

OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning Ocrbench: on the hidden mystery of ocr in large multimodal models

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:22:37.366993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T12:22:35.619550Z digest=sha256:578ca768cf12005dd58b830ff13052385bbbc7f52380257eefa938738c426979

Observation d66bc4c3-8c74-4396-8636-f1dcaca02fe3 · outbound

This paper cites MLLMs-Augmented Visual-Language Representation Learning.

OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning MLLMs-Augmented Visual-Language Representation Learning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T12:22:35.627900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:22:35.627900Z digest=sha256:8688bc1da616eb9ec3907e0b9ee762c8d33585abdae6a917b10fd759ae8909d9

Observation a4e4d833-649b-466e-a76b-920ad2148b11 · outbound

This paper cites Unified-io 2: Scaling autoregressive multimodal models with vision language audio and action.

OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning Unified-io 2: Scaling autoregressive multimodal models with vision language audio and action

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:22:37.338939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T12:22:35.643604Z digest=sha256:bf0bf1c7d48e24db30e1c42b4dc42a799a77ac09ca7cc2880a66e27d7a575a4c

Observation 01b1815e-e04a-4ccc-92a6-5e6a9a142855 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:22:37.311538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T12:22:35.652763Z digest=sha256:a76785157c735fc1aeed94849060973941e1af077f3bfd98e92b13e6fe87d2d9

Observation d4278d2c-71e8-4bb0-af6d-0ff39448a37f · outbound

This paper cites ChartQA: A benchmark for question answering about charts with visual and logical reasoning.

OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning ChartQA: A benchmark for question answering about charts with visual and logical reasoning

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:22:37.287786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T12:22:35.666074Z digest=sha256:86bb93c1ce4b8f744ae5bfbf3e37c1de9ce72c9d88edfc56ed16d288afb6c9a6

Observation 05f11dd9-6210-4a1f-afb4-27d7e2e95ff0 · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T12:22:35.673312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:22:35.673312Z digest=sha256:b0cccef30e6df7dd5c100c55d9934f62e7db6065a7ab5d3c494c818480698bb1

Observation 3c9f56c5-0976-4beb-8888-4e92d67123fa · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning Learn- ing transferable visual models from natural language super- vision

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-05T12:22:35.679406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:22:35.679406Z digest=sha256:bf714a8dd83b522770f02c7cf2466db2f29673469ce16f9050cfb50cbd3794d0

Observation 5987f3ac-96e9-4bf3-b8c5-8a67bc3d6004 · outbound

This paper cites Improving language understanding by gen- erative pre-training.

OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning Improving language understanding by gen- erative pre-training

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T12:22:35.685613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:22:35.685613Z digest=sha256:198a8e0ae941612f1b4705baa5c6f64bd1c18b6cb8e20a85a48314764d0cfd7c

Observation e2a84847-cf8c-4886-b920-c439921bc361 · outbound

This paper cites Language models are unsu- pervised multitask learners.

OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning Language models are unsu- pervised multitask learners

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:22:37.188232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T12:22:35.693138Z digest=sha256:fad85cce76d7c50b01bd1ce0ec92546f9ffacb87740c957d329d22146ef57550

Observation 1861d3a1-de43-44ad-b365-45865d87b0b4 · outbound

This paper cites Learning visual representations with caption annotations.

OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning Learning visual representations with caption annotations

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:22:37.150797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T12:22:35.703968Z digest=sha256:9e94c0d6e52c2603532cee7cbffcc9a2c1ca43aac25063ff08a80f1dbfc42374

Observation f6ccd962-bd75-4097-a000-c4602d3e3c7b · outbound

This paper cites LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs.

OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-05T12:22:35.710452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:22:35.710452Z digest=sha256:9ee6b8298d480dc2d21f3366ed565cb50e5ed870380e025913a81c9c23d74b4b

Observation b35f7dc4-eb34-448f-9acf-bc9f496caaf1 · outbound

This paper cites Towards vqa models that can read.

OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning Towards vqa models that can read

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-05T12:22:35.718812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:22:35.718812Z digest=sha256:e3c845bd3d4612f520832ef731a806d00a189872494c14d97caefaf23c1beffb

Observation 2d429898-6337-47ec-82bf-e4572a281f2c · outbound

This paper cites EVA-CLIP: Improved Training Techniques for CLIP at Scale.

OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-05T12:22:35.731050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:22:35.731050Z digest=sha256:d9edfedb8ae4675c0c792a2664bde073485ffc5a48c9ad1021f629441cd61e87

Observation b994c06d-2bf8-42f8-a806-75fb8d5c903c · outbound

This paper cites Emu: Generative Pretraining in Multimodality.

OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning Emu: Generative Pretraining in Multimodality

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-05T12:22:35.752713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:22:35.752713Z digest=sha256:b3dedf5ea8ed573eb70cdf9901a0f311613b813f89ba86dd4ae061669020225c

Observation 6f30213f-1104-46d1-b48d-583098ef147b · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-05T12:22:35.760863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:22:35.760863Z digest=sha256:2c2f5b1a2d8841462c704c90d5fd459018c74a39cf941bbadcd3ba6921432419

Observation 9bba2b83-8108-49f8-8411-f070ad9b491f · outbound

This paper cites Image captioners are scalable vision learners too.

OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning Image captioners are scalable vision learners too

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:22:37.111167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T12:22:35.768792Z digest=sha256:7f087198670ba039dde1642902244d335bdd098a03222a1b6fde7cb7086f6c1d

Observation e361bb20-d1f6-4f69-89b4-bea689bb6f11 · outbound

This paper cites Show and tell: A neural image caption gen- erator.

OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning Show and tell: A neural image caption gen- erator

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:22:37.087691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T12:22:35.777479Z digest=sha256:0ce6a0f70e20e7be12f2ec987555fcfd2942da661de5fe7e9f8726517c667999

Observation 5cd19f60-9f85-4e89-884c-0cb945a16a3c · outbound

This paper cites GIT: A Generative Image-to-text Transformer for Vision and Language.

OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning GIT: A Generative Image-to-text Transformer for Vision and Language

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-05T12:22:35.785999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:22:35.785999Z digest=sha256:35b19ef5cb18875089e540485a5bf265f0a44828199cb7ade132d992657edf36

Observation 3d0e87c1-b4d3-4fdc-85cd-1b468b421cd9 · outbound

This paper cites Simvlm: Simple visual language model pretraining with weak supervision.

OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning Simvlm: Simple visual language model pretraining with weak supervision

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:22:37.053871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T12:22:35.796062Z digest=sha256:b1049e4eca1eb53155f1e7d76d823938b171a517bfc942ba252681e06bba3a44

Observation 53b3c1f6-142c-4f3f-af8c-5d3e7362bde7 · outbound

This paper cites VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation.

OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-05T12:22:35.803847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:22:35.803847Z digest=sha256:1839d2874f46dda56ed7c86cfd9eca138dafb0ddf8caa493b7a655ee8264cbe0

Observation ebb02969-488b-488d-8a65-8b1db26bcfe2 · outbound

This paper cites Show-o: One Single Transformer to Unify Multimodal Understanding and Generation.

OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-05T12:22:35.817127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:22:35.817127Z digest=sha256:2b2142632c72a148b9c682315dbcf619b8f470235c6dff7ac79be6adec45c403

Observation 6618a9d2-9863-41de-a2ce-e5512b08acca · outbound

This paper cites Demystifying clip data.

OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning Demystifying clip data

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:22:37.019627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T12:22:35.829539Z digest=sha256:830c9611eba8de9cd1c10d20482f19f25cf74eb046a3ebb54bcfa2f7c840e67c

Observation 0ea1aedb-6682-474a-a017-ff205f8672d1 · outbound

This paper cites Show, attend and tell: Neural image caption gen- eration with visual attention.

OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning Show, attend and tell: Neural image caption gen- eration with visual attention

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:22:36.990932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T12:22:35.836490Z digest=sha256:6741dc6ad455e313051a867c27b6426c4eab44a4eee3c96c9575138220d34cb0

Observation 0f74dd6b-fb83-4269-9b00-d7cd63c16282 · outbound

This paper cites CoCa: Contrastive Captioners are Image-Text Foundation Models.

OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning CoCa: Contrastive Captioners are Image-Text Foundation Models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-05T12:22:35.842656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:22:35.842656Z digest=sha256:320a0d6a21dc105b5285cd01f1fcb379725a8458644a5b79487b7f39a8abc59f

Observation 15110c82-2031-4011-877f-ca696559a30a · outbound

This paper cites Sigmoid loss for language image pre-training.

OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning Sigmoid loss for language image pre-training

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:22:36.957260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T12:22:35.851286Z digest=sha256:566e07a6b2547cb202fc8b555469febdbb7e317be4dd17b12dd5263f6d015f6b

Observation 32d4ff18-a9f5-40b9-baf0-6c7177c5d672 · outbound

This paper cites Dreamlip: Language- image pre-training with long captions.

OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning Dreamlip: Language- image pre-training with long captions

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:22:36.924665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T12:22:35.866665Z digest=sha256:9d5568a46a07836bcf9160632fb02e727c49a75deca289dd6de09c18cf788571

Pith citing papers

Observation d54b4bd1-d104-473b-9408-1017bb394642 · inbound

Music Audio-Visual Question Answering Requires Specialized Multimodal Designs cites this paper.

Music Audio-Visual Question Answering Requires Specialized Multimodal Designs OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-19T14:12:22.099238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T14:11:52.011626Z digest=sha256:a94ae612665b06ce98c79229e952ad320d40677065b976706f676cc8c4fb3539

Observation ceff4272-2fc6-4feb-aacd-0a7e9e4d51ac · inbound

EXAONE 4.5 Technical Report cites this paper.

EXAONE 4.5 Technical Report OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T06:10:57.643157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T17:47:34.692414Z digest=sha256:f7c5f531f3ac2ddac1807346f36f2161ac8315eaac421489cf5e2f920f3acd0c

Observation b2aebc1d-71bf-4d4c-9d10-796e6daf5d92 · inbound

Let ViT Speak: Generative Language-Image Pre-training cites this paper.

Let ViT Speak: Generative Language-Image Pre-training OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning

Reference 46

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T16:01:09.608761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-09T18:56:51.627714Z digest=sha256:4d7ce6042ebfcc593e74e6e85cc23f4d516eafd6ff2a7d3b69a7ba34ef4f1381

Observation 872b850a-b17a-41ea-a8c2-cce065bdfcc7 · inbound

Let ViT Speak: Generative Language-Image Pre-training cites this paper.

Let ViT Speak: Generative Language-Image Pre-training OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning

Reference 46

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T07:35:28.747148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-01T07:35:07.825460Z digest=sha256:a185586ed62dfdec02fb1bf6a6a9a874e456abbc0ab93bf3b81347a0981b550d