Pith. sign in

Paper Citation Record · LEDGER

Multimodal Autoregressive Pre-training of Large Vision Encoders

As of 15 August 2026, this Paper Citation Record lists 100 of 137 outbound references and 20 inbound Pith citation observations for arXiv:2411.14402.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.14402 v1

Coverage vector

measured 100 of 137 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T15:17:22.748659Z

measured 120 of 120 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 20 of 20 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:07:59.768232Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-23T01:25:16.481218Z

Reference resolution

100 of 137 outbound references displayed

  • verified exact0
  • verified fuzzy13
  • unresolved87
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9b4ca4f7-aa21-4a23-b87f-385b1a76dfe2 · outbound

This paper cites GPT-4 Technical Report.

Multimodal Autoregressive Pre-training of Large Vision Encoders GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:22.375349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:22.375349Z digest=sha256:7b4a4ec25fa8bbe5c3be4712a8b09f95c83b1efec303c965ebb37ccf1a96371b

Observation 5bc970ba-21f5-412f-9762-c173c791013b · outbound

This paper cites Nocaps: Novel object cap- tioning at scale.

Multimodal Autoregressive Pre-training of Large Vision Encoders Nocaps: Novel object cap- tioning at scale

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:22.380106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:22.380106Z digest=sha256:7ffa94c4c21fa46b6b360b6f645e62112ba5a1cd2ab5e80705d729bdf7c7e0ab

Observation 31e2e298-fb61-4d80-a9b0-73b813ff38a4 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

Multimodal Autoregressive Pre-training of Large Vision Encoders Flamingo: a visual language model for few-shot learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:22.383937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:22.383937Z digest=sha256:455099dc864a359c73d16e8e6d1845e6fca29f2f47939b9e531fb89506be0719

Observation 8b78b674-1330-43fe-af85-fe247ef76a06 · outbound

This paper cites Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture.

Multimodal Autoregressive Pre-training of Large Vision Encoders Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:22.387847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:22.387847Z digest=sha256:2616c8bf678c849d99ed2275c70e62cd1b54859fd0d8db60d6f41c30aa34c896

Observation 30508ccc-e3ba-41ee-a202-96ea9ac6f0d6 · outbound

This paper cites Qwen Technical Report.

Multimodal Autoregressive Pre-training of Large Vision Encoders Qwen Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:22.391871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:22.391871Z digest=sha256:ad30a80d254319e0d5a6b67c6836944213e8f0faa9f2be9e4a19c97ca7974979

Observation fdae0bd5-2d23-4638-8d52-81c17daf68d3 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Multimodal Autoregressive Pre-training of Large Vision Encoders Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:22.395674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:22.395674Z digest=sha256:3767193c8de7c85e298a0d656894c68eb4133f8b31fc5bf346149936fb0ba2d9

Observation 6ca328a1-9710-41a6-a5d2-fec628540c1f · outbound

This paper cites From detection of individual metastases to classification of lymph node status at the pa- tient level: the camelyon17 challenge.

Multimodal Autoregressive Pre-training of Large Vision Encoders From detection of individual metastases to classification of lymph node status at the pa- tient level: the camelyon17 challenge

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:22.399915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:22.399915Z digest=sha256:a8254d31c489a74274b57fd08aedb8a41e9d69868bf9e64b7ce1a7637e71f59c

Observation a0aaf7d9-e904-46f6-ac68-6ad90c6eac8e · outbound

This paper cites BEiT: Bert pre- training of image transformers.

Multimodal Autoregressive Pre-training of Large Vision Encoders BEiT: Bert pre- training of image transformers

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:22.403826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:22.403826Z digest=sha256:a54c36c99ff2708e7c980ae3aa1446ed62d3e3fda4275f8a5950429e94e9d8db

Observation bcf6dafb-1f3d-476c-ba2e-4b2f8ae9efe9 · outbound

This paper cites Flexivit: One model for all patch sizes.

Multimodal Autoregressive Pre-training of Large Vision Encoders Flexivit: One model for all patch sizes

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:22.407402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:22.407402Z digest=sha256:c4c60b50c523dd24019e4afa5f36ce0f5e8aaba79be4ffc90d65e983a2504bb5

Observation d1e0d4bb-7189-4809-b5e2-fee77737d066 · outbound

This paper cites Food-101 – mining discriminative components with ran- dom forests.

Multimodal Autoregressive Pre-training of Large Vision Encoders Food-101 – mining discriminative components with ran- dom forests

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:22.411184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:22.411184Z digest=sha256:97cb8e477fab9de2de552d7b7958a471d0f12488305e457ea6e47915f7dce0ed

Observation fea36b10-7818-4644-b7e3-81ca5fc6a9a1 · outbound

This paper cites Time series analysis: forecasting and control.

Multimodal Autoregressive Pre-training of Large Vision Encoders Time series analysis: forecasting and control

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:22.415076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:22.415076Z digest=sha256:b21c63c69fff5be202a798b00e4a62798b902b795104c22b2bb1acd0dd06c9a5

Observation 1bfeb3e2-5170-4879-937f-83f4dd8060c3 · outbound

This paper cites Language Models are Few-Shot Learners.

Multimodal Autoregressive Pre-training of Large Vision Encoders Language Models are Few-Shot Learners

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:22.418593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:22.418593Z digest=sha256:1f89c75aa708254897573219a36d819332fd3a0b741fb8bcf59e921bb69df148

Observation faf41b7a-c743-442e-83a2-6ecbf8723a24 · outbound

This paper cites Coyo-700m: Image-text pair dataset, 2022.

Multimodal Autoregressive Pre-training of Large Vision Encoders Coyo-700m: Image-text pair dataset, 2022

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:22.422405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:22.422405Z digest=sha256:d05a9990fca3d263111ac722a5ff1ea09d21692033069f78294f0691a24ccda3

Observation cfe5a846-2afb-449b-83c3-ba8d4b50bd97 · outbound

This paper cites End-to-end object detection with transformers.

Multimodal Autoregressive Pre-training of Large Vision Encoders End-to-end object detection with transformers

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:22.425638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:22.425638Z digest=sha256:c650d512096473bd0624f8cc381c9f87511337a9cf8242586b66363bdd141dbf

Observation 9aae2335-5283-4c6f-9ae1-08d9149961fc · outbound

This paper cites Unsupervised learn- ing of visual features by contrasting cluster assignments.

Multimodal Autoregressive Pre-training of Large Vision Encoders Unsupervised learn- ing of visual features by contrasting cluster assignments

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:22.429185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:22.429185Z digest=sha256:8636bc0666abbe2482b735fe87aa0f01b493e6270f7c9b8ac9836e1606be8506

Observation b02d1d74-2054-4069-a87c-e5af80c729c7 · outbound

This paper cites Emerging properties in self-supervised vision transformers.

Multimodal Autoregressive Pre-training of Large Vision Encoders Emerging properties in self-supervised vision transformers

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:22.432738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:22.432738Z digest=sha256:e32bd10ba42045ae7f71b446800f90b373e27f3f521a22d4da3eacd6acb9e770

Observation bd86e4ca-80fa-4289-b35c-fe7d84ec4339 · outbound

This paper cites A generative approach for wikipedia-scale visual entity recognition.

Multimodal Autoregressive Pre-training of Large Vision Encoders A generative approach for wikipedia-scale visual entity recognition

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:22.436223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:22.436223Z digest=sha256:72aeb78b397724b92f409a03098895fb659953248575aef4b282f38e41fe5b5d

Observation 82341a4f-b06c-4602-9c52-85318551e01e · outbound

This paper cites Mmdetection: Open mmlab detection toolbox and benchmark.

Multimodal Autoregressive Pre-training of Large Vision Encoders Mmdetection: Open mmlab detection toolbox and benchmark

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:22.439932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:22.439932Z digest=sha256:ca799bcfb691e815b88d7cbba38ecffd9d8254d9db9bf79990cf3b07338ca375

Observation 7f693463-45a9-4f47-9598-bf16e1eecb30 · outbound

This paper cites Generative pre- training from pixels.

Multimodal Autoregressive Pre-training of Large Vision Encoders Generative pre- training from pixels

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:22.443385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:22.443385Z digest=sha256:735987a58422157b59b8b4fcd407baac87fa54aa1d33c180dc9dffde15f0c917

Observation 49f7072e-60e7-44f1-bc1e-2b55fb7d9bd3 · outbound

This paper cites A simple framework for contrastive learning of visual representations.

Multimodal Autoregressive Pre-training of Large Vision Encoders A simple framework for contrastive learning of visual representations

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:22.446597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:22.446597Z digest=sha256:e98d65cfa99bd8bcff8da8a75141ae664f0d546bd4b85ab1b7167f23ecaa945c

Observation 98afe1f3-3b4f-4e56-a41f-3f6d42fbaaab · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

Multimodal Autoregressive Pre-training of Large Vision Encoders Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:22.449822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:22.449822Z digest=sha256:8595547e519a2710e8c9a3568146a7a92e1c23020467fdca8498fbbf00d7c3c4

Observation d2ad7346-77b2-48b3-a3da-8878fb5733a2 · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

Multimodal Autoregressive Pre-training of Large Vision Encoders Gonzalez, Ion Stoica, and Eric P

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:22.453804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:22.453804Z digest=sha256:2e539f880d45462051a65ac5aabd52433fc18b38e51b8f25ab680eab01058261

Observation 3d45c89c-538e-4a6a-9985-3d4791db9511 · outbound

This paper cites PaLM: Scaling Language Modeling with Pathways.

Multimodal Autoregressive Pre-training of Large Vision Encoders PaLM: Scaling Language Modeling with Pathways

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:22.457456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:22.457456Z digest=sha256:7f2bd8d28745a1825446b94e15015f22cf7aed5b872adb54308c7ab689ba0f24

Observation dd188d22-d450-478e-b068-16a62335877f · outbound

This paper cites Functional map of the world.

Multimodal Autoregressive Pre-training of Large Vision Encoders Functional map of the world

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:22.461117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:22.461117Z digest=sha256:28dae6a4b6dbb0fe57e0e3e973f944de5d99c2f7f721d604e55d4ab1ee522559

Observation 80c5bd87-eacb-4ac9-8091-3d0ea2648a1f · outbound

This paper cites Cimpoi, S.

Multimodal Autoregressive Pre-training of Large Vision Encoders Cimpoi, S

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:22.465645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:22.465645Z digest=sha256:11e4495957f989837bafc3082bd45e646a0712d70d27df174f8d98e5d1a02eb4

Observation 2aca9d57-3dea-4db5-a53a-f39a78fcc253 · outbound

This paper cites Patch n’pack: Navit, a vision trans- former for any aspect ratio and resolution.

Multimodal Autoregressive Pre-training of Large Vision Encoders Patch n’pack: Navit, a vision trans- former for any aspect ratio and resolution

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:22.469894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:22.469894Z digest=sha256:477060f607e14bc9b0adcbda470f436c6cbb9ee655726d88f117087567cf90c0

Observation 91829b37-0168-497f-ae80-517a6d679241 · outbound

This paper cites Imagenet: A large-scale hierarchical im- age database.

Multimodal Autoregressive Pre-training of Large Vision Encoders Imagenet: A large-scale hierarchical im- age database

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:22.474699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:22.474699Z digest=sha256:6bb5c51df60b848f62bdc6714546f54db6a208a6aa20c835ceb4734d8e2f4dc8

Observation 4eb682a7-1c36-42bf-b02a-acaa29fea8cd · outbound

This paper cites Virtex: Learning visual representations from textual annotations.

Multimodal Autoregressive Pre-training of Large Vision Encoders Virtex: Learning visual representations from textual annotations

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:22.478817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:22.478817Z digest=sha256:45c09b133e60cdbaee841f06c81694157b35460b1e9ec693df30e5cf3bf156d8

Observation 690aec9b-61b4-4662-97f4-54e7a7285180 · outbound

This paper cites Unsu- pervised visual representation learning by context predic- tion.

Multimodal Autoregressive Pre-training of Large Vision Encoders Unsu- pervised visual representation learning by context predic- tion

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:22.482407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:22.482407Z digest=sha256:46ede6f4528f7ab773446238bc314ff4a6bdadbd8e5c939821ef61e972aa376a

Observation ae746e19-b83e-4bf3-8fce-8ce1df98e0bd · outbound

This paper cites An image is worth 16x16 words: Trans- formers for image recognition at scale.

Multimodal Autoregressive Pre-training of Large Vision Encoders An image is worth 16x16 words: Trans- formers for image recognition at scale

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:22.487173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:22.487173Z digest=sha256:0cf41a41cb901715d25ee338ae7c32b324e21c366f78b8e0014eaf72dfdfc558

Observation a8a7aa3d-9257-4e57-a637-4714d1977bdc · outbound

This paper cites The Llama 3 Herd of Models.

Multimodal Autoregressive Pre-training of Large Vision Encoders The Llama 3 Herd of Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:22.494393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:22.494393Z digest=sha256:6b5515018cf74a66d380568e054062ed1f2b03934bdce5224db322d9fe8c25eb

Observation 1992e40d-ba70-456d-b241-0dcf513c88bc · outbound

This paper cites Scalable Pre-training of Large Autoregressive Image Models.

Multimodal Autoregressive Pre-training of Large Vision Encoders Scalable Pre-training of Large Autoregressive Image Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:22.498155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:22.498155Z digest=sha256:64b8645beabf3ff496d74f87fdeb888a6e2ad3f2d2c8a4019662d7b9600787c2

Observation 5fb55291-be1a-4bbe-ba82-cf079e8a24dd · outbound

This paper cites Taming transformers for high-resolution image synthesis.

Multimodal Autoregressive Pre-training of Large Vision Encoders Taming transformers for high-resolution image synthesis

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:22.503099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:22.503099Z digest=sha256:8992137ea63fddca8e000651e6015bf7cf7e8a4e81cde5c6a763ba6a91f246d5

Observation a1765765-aa45-4efa-a01b-f7b35f3a3339 · outbound

This paper cites Data Filtering Networks.

Multimodal Autoregressive Pre-training of Large Vision Encoders Data Filtering Networks

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:22.507418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:22.507418Z digest=sha256:de4d9a95ca8dc39791cf769c0c012388ff2496d06e710ff78cb1712b6343362f

Observation 088deace-bdcd-43f2-bbe0-d76de2778712 · outbound

This paper cites Slowfast networks for video recognition.

Multimodal Autoregressive Pre-training of Large Vision Encoders Slowfast networks for video recognition

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:22.512043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:22.512043Z digest=sha256:a97dcf7007bd61653adf701a0f61ddfe7f449b9df8f3766ceb9ebe1f4ce1befd

Observation bf53a0a5-6d35-4b63-9e11-ce2d0c64349b · outbound

This paper cites Improved baselines for vision-language pre-training.

Multimodal Autoregressive Pre-training of Large Vision Encoders Improved baselines for vision-language pre-training

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:22.515438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:22.515438Z digest=sha256:c18ae4487f08a44acc7b661a5106198ccbad5c7d5dfc248683083be659b71a2c

Observation a0b8ea7e-0c8b-4fb4-a487-a74684970089 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

Multimodal Autoregressive Pre-training of Large Vision Encoders MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:22.519701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:22.519701Z digest=sha256:47f444ad8d4347a45144c6f34efc6351151b314527423063e6814f924248c5cc

Observation d70bbe70-3d40-42b7-9f27-f293dfce0bef · outbound

This paper cites Unsupervised Representation Learning by Predicting Image Rotations.

Multimodal Autoregressive Pre-training of Large Vision Encoders Unsupervised Representation Learning by Predicting Image Rotations

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:22.523051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:22.523051Z digest=sha256:7e2af46436d27cde7d48f5052748f3c31d5e0b1b7ee4ae5687b0eb6c7eb65a6e

Observation 86b772e0-2934-4f93-8c38-6549c91ecee1 · outbound

This paper cites Making the v in vqa matter: El- evating the role of image understanding in visual question answering.

Multimodal Autoregressive Pre-training of Large Vision Encoders Making the v in vqa matter: El- evating the role of image understanding in visual question answering

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:22.526494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:22.526494Z digest=sha256:77151d4c1ae5d12aeb4f315dc21e741527f27102d80673474c355c07dcdb9f7c

Observation 8fa25754-c13e-4cf0-8f5b-1b7b38ad7865 · outbound

This paper cites Making the v in vqa matter: El- evating the role of image understanding in visual question answering.

Multimodal Autoregressive Pre-training of Large Vision Encoders Making the v in vqa matter: El- evating the role of image understanding in visual question answering

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:22.529626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:22.529626Z digest=sha256:87e068399a44e4484fbfba48372a4f2ca32fd20480181adad822defe8d4bb705

Observation 23c64246-7b44-458a-a79c-1226b5b4ce5b · outbound

This paper cites Bootstrap your own latent-a new approach to self-supervised learning.

Multimodal Autoregressive Pre-training of Large Vision Encoders Bootstrap your own latent-a new approach to self-supervised learning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:22.533061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:22.533061Z digest=sha256:e27aab795210bceed4c69de36cac4f03a1977e3592d743ab052778fc270c1301

Observation 3ff49287-8071-47bb-8fb6-68544ef8dff3 · outbound

This paper cites Lvis: A dataset for large vocabulary instance segmentation, 2019.

Multimodal Autoregressive Pre-training of Large Vision Encoders Lvis: A dataset for large vocabulary instance segmentation, 2019

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:22.536199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:22.536199Z digest=sha256:56136d262dc408b39cf9bd88135e6586ec3f3e63c83d0454d9aa33269ab6a8d0

Observation 13d5870a-e29d-4229-8d66-f7f743879f10 · outbound

This paper cites Vizwiz grand challenge: Answering visual questions from blind people.

Multimodal Autoregressive Pre-training of Large Vision Encoders Vizwiz grand challenge: Answering visual questions from blind people

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:22.539365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:22.539365Z digest=sha256:465afa171e3a2e95251ca26aa9110a16c2adaf4da8b70845ffa4f24ce1dc0b17

Observation 12e83030-6b2d-4746-ab7c-d33e8bea2616 · outbound

This paper cites Deep residual learning for image recognition.

Multimodal Autoregressive Pre-training of Large Vision Encoders Deep residual learning for image recognition

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:22.542616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:22.542616Z digest=sha256:44758c585553acc7ccd13e230ab91be3d4366c242f5a34de8cd1bd27062c96a4

Observation 7ee7a35d-744a-4062-99a0-2a58a7fd32ce · outbound

This paper cites Mask r-cnn.

Multimodal Autoregressive Pre-training of Large Vision Encoders Mask r-cnn

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:22.546135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:22.546135Z digest=sha256:2151f9e06ff3d92d3546a12e8d970d5753499cb93b845a0940b47038a2e04b61

Observation 5c28c0b0-9a26-4489-b643-e5c3febe66e2 · outbound

This paper cites Mask r-cnn.

Multimodal Autoregressive Pre-training of Large Vision Encoders Mask r-cnn

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:22.550155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:22.550155Z digest=sha256:733d0ab86495422f822dccf2f880bd37c2d783cbd7aa02ee878c12524a511c34

Observation fd3291c8-42a7-4c0c-9baa-875300758e53 · outbound

This paper cites Masked autoencoders are scal- able vision learners.

Multimodal Autoregressive Pre-training of Large Vision Encoders Masked autoencoders are scal- able vision learners

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:22.553317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:22.553317Z digest=sha256:8fe4200215cf7b17e44048e3cf101b1b5d64667f22816ff18bbfa1c9600c0f77

Observation 80e27a5a-c5da-485f-8275-34714b1e1286 · outbound

This paper cites Eurosat: A novel dataset and deep learn- ing benchmark for land use and land cover classification,.

Multimodal Autoregressive Pre-training of Large Vision Encoders Eurosat: A novel dataset and deep learn- ing benchmark for land use and land cover classification,

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:22.556383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:22.556383Z digest=sha256:50f24e6bc2e235fdade24cf882416e55fd94bcb4b99f865f28885f4a725592e0

Observation eb1403ae-34f0-443f-853c-5cd826108dab · outbound

This paper cites Training Compute-Optimal Large Language Models.

Multimodal Autoregressive Pre-training of Large Vision Encoders Training Compute-Optimal Large Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:22.559990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:22.559990Z digest=sha256:f1b43527b775f971054f017cf35116995c7a0d596cd944733f2012b513c8ef88

Observation 05e366ef-30f2-4008-85a3-79ad81433cf5 · outbound

This paper cites Scaling up vision-language pre-training for image captioning.

Multimodal Autoregressive Pre-training of Large Vision Encoders Scaling up vision-language pre-training for image captioning

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:22.564126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:22.564126Z digest=sha256:77fe6c57fa6c0f4cdc3fafb414de97609e4e8b9011b67e5ed814a27d0f15d052

Observation 63a60023-7453-4440-b278-620ce14c247a · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

Multimodal Autoregressive Pre-training of Large Vision Encoders Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:22.567566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:22.567566Z digest=sha256:ff7743d98cf07b4fc4c6a1af4a04e763e4f17773aa5a4f449b6836f28b5736f4

Observation d39d5962-7672-41ff-babf-2bc86d1b0491 · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

Multimodal Autoregressive Pre-training of Large Vision Encoders Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:22.571585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:22.571585Z digest=sha256:f6ba2e075616477575b5e5dc24e23058f257b7c379a37ab2534c490c62af2379

Observation 0aecd157-d05d-41ad-a0ea-d033fc9b4533 · outbound

This paper cites Scaling up visual and vision-language representation learning with noisy text supervision.

Multimodal Autoregressive Pre-training of Large Vision Encoders Scaling up visual and vision-language representation learning with noisy text supervision

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:22.575427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:22.575427Z digest=sha256:cef0eb65d96b9c917145cbbd4fddca07ade9ac1f169696cdb4eedc5eee86b36b

Observation 5748ad72-33e4-4763-a8aa-98118d65e4bc · outbound

This paper cites Scaling Laws for Neural Language Models.

Multimodal Autoregressive Pre-training of Large Vision Encoders Scaling Laws for Neural Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:22.579013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:22.579013Z digest=sha256:12c72b0042acd16ed1784ef5eee4e846904ed1417c6374ee1d34e3b9e9ade2fb

Observation c1318717-68e9-4439-a9f8-c0befa5cc540 · outbound

This paper cites Deep visual-semantic alignments for generating image descriptions.

Multimodal Autoregressive Pre-training of Large Vision Encoders Deep visual-semantic alignments for generating image descriptions

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:22.582688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:22.582688Z digest=sha256:6c7ba8dd393875d7926ae71adad73dfc3e74c378c8b6aa3222bec422047d1849

Observation 3b351665-a929-4a3d-bba2-0eff118a7108 · outbound

This paper cites Referitgame: Referring to objects in pho- tographs of natural scenes.

Multimodal Autoregressive Pre-training of Large Vision Encoders Referitgame: Referring to objects in pho- tographs of natural scenes

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:22.586992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:22.586992Z digest=sha256:7f47d70c03c61aec20d9ecf3e2aadaf4fa4f736b2b8946fb2070abc3c839f94d

Observation 7e5807bb-5974-4221-afa7-8045de742895 · outbound

This paper cites Big transfer (bit): General visual representation learning.

Multimodal Autoregressive Pre-training of Large Vision Encoders Big transfer (bit): General visual representation learning

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:22.591376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:22.591376Z digest=sha256:691b8902fcd0c10f1e961a9de1e3c61aacef7bb0b67651ac44878e572be621f4

Observation 76009505-4fbb-413b-8663-1e5308d636da · outbound

This paper cites 3d object representations for fine-grained categorization.

Multimodal Autoregressive Pre-training of Large Vision Encoders 3d object representations for fine-grained categorization

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:22.595126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:22.595126Z digest=sha256:a942f6446930935214f4dc88ce0e665b444291a6771a39210a7d22e67142e246

Observation 9cdf6400-8965-4acb-8341-ddc5a06bcbe8 · outbound

This paper cites Learning multiple layers of features from tiny images.

Multimodal Autoregressive Pre-training of Large Vision Encoders Learning multiple layers of features from tiny images

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:22.598503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:22.598503Z digest=sha256:fa739195f5689826e0cfd3793b5fa613877397ab89f2c5433b6e63a26ae378f7

Observation b9f089a3-89e1-425f-b757-0011b3f42c5b · outbound

This paper cites Imagenet classification with deep convolutional neural net- works.

Multimodal Autoregressive Pre-training of Large Vision Encoders Imagenet classification with deep convolutional neural net- works

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:22.602219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:22.602219Z digest=sha256:47d6bff698070e6cb3dfb3ef52fc6cac45a93f3e4a302a4d49eb5c7c4966a70c

Observation bc337774-0b21-459f-bd9e-05cb3c7dac15 · outbound

This paper cites MaMMUT: A Simple Architecture for Joint Learning for MultiModal Tasks.

Multimodal Autoregressive Pre-training of Large Vision Encoders MaMMUT: A Simple Architecture for Joint Learning for MultiModal Tasks

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:22.605601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:22.605601Z digest=sha256:93eae1125760b8b048dfcf22641dc322850f952455afb08ad876960c25d8f796

Observation a833fbd6-1add-4ddf-8647-8e803f910412 · outbound

This paper cites Revisit Large-Scale Image-Caption Data in Pre-training Multimodal Foundation Models.

Multimodal Autoregressive Pre-training of Large Vision Encoders Revisit Large-Scale Image-Caption Data in Pre-training Multimodal Foundation Models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:22.609197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:22.609197Z digest=sha256:148e3b3035d774285c1c36b517b010311ca1b5bb4002261413d7447a25adbf5c

Observation fa4c8a09-bab1-4142-96d4-0199b6d75596 · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

Multimodal Autoregressive Pre-training of Large Vision Encoders SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:22.613773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:22.613773Z digest=sha256:37462d0f2c7a49b32dae3a6b2357020dd421003711fc5e01e25975b25e1868ba

Observation 132390e8-8aae-4c4b-aea4-1ae3368f32fe · outbound

This paper cites Align before fuse: Vision and language representation learning with momentum distillation.

Multimodal Autoregressive Pre-training of Large Vision Encoders Align before fuse: Vision and language representation learning with momentum distillation

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:22.617750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:22.617750Z digest=sha256:1fba435cb6959ce110efafafef5b6cf119d22d114166b6f6a0bf15395150daa8

Observation 560a87bd-409a-4538-b435-258c73db3539 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for uni- fied vision-language understanding and generation.

Multimodal Autoregressive Pre-training of Large Vision Encoders Blip: Bootstrapping language-image pre-training for uni- fied vision-language understanding and generation

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:22.621329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:22.621329Z digest=sha256:04f8a377106f564ec9f789132f0a9638c0456f6cee94557bf3be15ba1479e53e

Observation 8198cf4a-a4ef-4afd-a553-3be1aabad19b · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Multimodal Autoregressive Pre-training of Large Vision Encoders Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:22.624641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:22.624641Z digest=sha256:0dbd17e767b513673229865d388815482602684513fbd1e23c71220eee946049

Observation ca82604f-08b2-4827-b329-88128bad07bf · outbound

This paper cites Exploring plain vision transformer backbones for object de- tection.

Multimodal Autoregressive Pre-training of Large Vision Encoders Exploring plain vision transformer backbones for object de- tection

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:22.627917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:22.627917Z digest=sha256:8cda67b830a779eea34a5b3f2113fdb636c7026848721f64051dbe5aef4b251d

Observation c7f349a3-6580-43ca-8b2e-a5e15848f4fd · outbound

This paper cites Exploring plain vision transformer backbones for object de- tection, 2022.

Multimodal Autoregressive Pre-training of Large Vision Encoders Exploring plain vision transformer backbones for object de- tection, 2022

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:22.631804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:22.631804Z digest=sha256:0c59d119eda7934037f166be3d97f9b172ffb58575644f244c4045d4732d5ae7

Observation 4202a41f-ab2b-4e18-a454-ba210cf316be · outbound

This paper cites Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models.

Multimodal Autoregressive Pre-training of Large Vision Encoders Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:22.635481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:22.635481Z digest=sha256:0c8144aed869046fd1c32420a18dee4f1338709829fa2dac9b0eb4d5643ad4f3

Observation ac9ce47e-8e0b-4bd7-9f98-13c7930fda00 · outbound

This paper cites Microsoft coco: Common objects in context.

Multimodal Autoregressive Pre-training of Large Vision Encoders Microsoft coco: Common objects in context

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:22.640231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:22.640231Z digest=sha256:7068439152267d6928344cff3b63cd09e3148b0d9984359e6660b7ec97681b72

Observation 9afdd6f1-283f-4598-896e-be3439bff8dd · outbound

This paper cites SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models.

Multimodal Autoregressive Pre-training of Large Vision Encoders SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:22.643774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:22.643774Z digest=sha256:879030fdcfb5bdc2cad800d3c8b855ab1c85c285e8a098e6f79116246916df66

Observation f866b314-9cce-43a4-9011-b15f6c55691f · outbound

This paper cites Improved baselines with visual instruction tuning.

Multimodal Autoregressive Pre-training of Large Vision Encoders Improved baselines with visual instruction tuning

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:22.647238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:22.647238Z digest=sha256:607b039c31d366b5ad0b9ee08c898229a59a2ae3c1d9a1101ca3819564c4fde9

Observation d2ae77fa-0722-40cd-9765-4d6ad0daf1a8 · outbound

This paper cites Grounding dino: Mar- rying dino with grounded pre-training for open-set object detection.

Multimodal Autoregressive Pre-training of Large Vision Encoders Grounding dino: Mar- rying dino with grounded pre-training for open-set object detection

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:22.650754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:22.650754Z digest=sha256:1bed7af1a421ffa437c4bc2228deb39dab8a34566806717977865151c97accfa

Observation 0047e8f2-658e-4f8f-b9e9-c3ddba6b01c7 · outbound

This paper cites Sgdr: Stochastic gradient descent with warm restarts.

Multimodal Autoregressive Pre-training of Large Vision Encoders Sgdr: Stochastic gradient descent with warm restarts

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:22.656500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:22.656500Z digest=sha256:b033241754ec9d09b74f6e49642ab9798649181fd23d15b427683590654b1bb4

Observation a4fce8aa-c5aa-4817-be8f-0be5cd4befb5 · outbound

This paper cites Decoupled Weight Decay Regularization.

Multimodal Autoregressive Pre-training of Large Vision Encoders Decoupled Weight Decay Regularization

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:22.660174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:22.660174Z digest=sha256:1a8750a8f16b2bb0978e4ac476715eaaa735206014c7aa4f4cbd1c90ccae7433

Observation f1697871-86ca-4499-a580-a7acba5fca1f · outbound

This paper cites Unified-io 2: Scaling autoregressive mul- timodal models with vision language audio and action.

Multimodal Autoregressive Pre-training of Large Vision Encoders Unified-io 2: Scaling autoregressive mul- timodal models with vision language audio and action

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:22.664742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:22.664742Z digest=sha256:415f4816a61337de88d59b3b46d4cf3a0ea677d5d9a8070528896ae7e2c2723d

Observation 90824477-ae47-423c-bdf4-e46c9fa7e2ec · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

Multimodal Autoregressive Pre-training of Large Vision Encoders Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:22.667961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:22.667961Z digest=sha256:8c622c691bd47294358a4b496ec2e261f1fc419fd8a5c45c7f6a970b985188e7

Observation 52fa79ac-f9ad-4a93-9289-df5ffdd51b15 · outbound

This paper cites Generation and comprehension of unambiguous object descriptions.

Multimodal Autoregressive Pre-training of Large Vision Encoders Generation and comprehension of unambiguous object descriptions

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:22.671099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:22.671099Z digest=sha256:7cc461f80d376c2408ed349e444e9d0727e8ef8235d781d7f63e24c2449bf037

Observation 2ca744f4-295d-42da-b98d-afaa0a98c49f · outbound

This paper cites Ok-vqa: A visual question answering benchmark requiring external knowledge.

Multimodal Autoregressive Pre-training of Large Vision Encoders Ok-vqa: A visual question answering benchmark requiring external knowledge

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:17:23.684365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:17:22.674161Z digest=sha256:4c2b4aa631e6c3e14ebfdcc6391aa90d6a161d83743477fa512df910037d2d78

Observation 0abe84aa-6dee-4f3f-9be3-220987b200e0 · outbound

This paper cites Ok-vqa: A visual question answering benchmark requiring external knowledge.

Multimodal Autoregressive Pre-training of Large Vision Encoders Ok-vqa: A visual question answering benchmark requiring external knowledge

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:17:23.671892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:17:22.678504Z digest=sha256:99abe0f1020608d76144b6c3e864f04018120f2ffc789b5fb061841d5929c8a7

Observation bea6496d-aa2e-4caf-b0d2-3b82603a3a1d · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

Multimodal Autoregressive Pre-training of Large Vision Encoders ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:22.682104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:22.682104Z digest=sha256:e42d752eca907c1284c24b0000b5a333da7e6cda5102b02fc98dbbb2364f3073

Observation f8ecadcb-5000-4dbe-bf0a-f24340b5c8a3 · outbound

This paper cites Docvqa: A dataset for vqa on document images.

Multimodal Autoregressive Pre-training of Large Vision Encoders Docvqa: A dataset for vqa on document images

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:17:23.658332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:17:22.685585Z digest=sha256:6811aa26cf2d02dfc7d3c2664eedc5e997237871f20edc5736a0a61632f8f516

Observation a853ea2a-9a0c-4984-8eac-8852141e4573 · outbound

This paper cites Infographicvqa.

Multimodal Autoregressive Pre-training of Large Vision Encoders Infographicvqa

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:17:23.646818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:17:22.689412Z digest=sha256:b219515b887d660ea3ff188b19ef037ac931de18e748631746b61dc493c6ff4d

Observation 36a14cf9-c4b1-4240-aa1a-b7e7ca2a5d02 · outbound

This paper cites MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training.

Multimodal Autoregressive Pre-training of Large Vision Encoders MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:22.692838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:22.692838Z digest=sha256:6e6b072f88d218abf4c6dc3c6028810397f99991f49ca13f29383328390ef024

Observation 1d64046b-f57e-4bfb-ac4a-b991eab63b52 · outbound

This paper cites Unsupervised learning of visual representations by solving jigsaw puzzles.

Multimodal Autoregressive Pre-training of Large Vision Encoders Unsupervised learning of visual representations by solving jigsaw puzzles

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:22.696091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:22.696091Z digest=sha256:9beaf2d610ae298faef0f3731fc84c2ccbd5fceaa15af0b24405799cfc296fc0

Observation 77b4acc1-b1e3-4955-940a-b30e1a2b54a2 · outbound

This paper cites an unresolved cited work.

Multimodal Autoregressive Pre-training of Large Vision Encoders Unresolved cited work

Reference 87

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:17:23.628987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:17:22.699767Z digest=sha256:d41a54f7c46995be64708edb7686bece90f74a2b7c83bd4d57ca943b27b7d93f

Observation 056ed4f9-b7c3-48f7-8054-b0d53f8977a4 · outbound

This paper cites an unresolved cited work.

Multimodal Autoregressive Pre-training of Large Vision Encoders Unresolved cited work

Reference 88

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:17:23.617133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:17:22.702845Z digest=sha256:75350721092841779415997b5061ee041622599f53336ddf66d6ca9c0bb98a7b

Observation 981aae9a-5948-410d-9a05-2c1bcc9ca78f · outbound

This paper cites Moment matching for multi-source domain adaptation.

Multimodal Autoregressive Pre-training of Large Vision Encoders Moment matching for multi-source domain adaptation

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:17:23.606978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:17:22.705978Z digest=sha256:420c30f0730bdd02f71789986544eaf0b3a20a2de0e3024516efa5c64271f510

Observation 837b0f71-f317-4159-90cc-57ef100438dc · outbound

This paper cites Plummer, Liwei Wang, Christopher M.

Multimodal Autoregressive Pre-training of Large Vision Encoders Plummer, Liwei Wang, Christopher M

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:17:23.596192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:17:22.709790Z digest=sha256:aa4d03cec7962091a5681c28eb988f4e77e1613bd6f541bd315ae1213fe4765c

Observation 70cf87f0-7534-407b-90ba-181f700b07fc · outbound

This paper cites Dataset Decomposition: Faster LLM Training with Variable Sequence Length Curriculum.

Multimodal Autoregressive Pre-training of Large Vision Encoders Dataset Decomposition: Faster LLM Training with Variable Sequence Length Curriculum

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:22.713005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:22.713005Z digest=sha256:06a5c3bbe4eda2c0692e292392fc246220dbd9f4986a947f40147c58c350ca66

Observation f0fd77eb-00f8-4758-a39b-cb163db3bb9b · outbound

This paper cites Improving language understanding by genera- tive pre-training.

Multimodal Autoregressive Pre-training of Large Vision Encoders Improving language understanding by genera- tive pre-training

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:17:23.585017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:17:22.716442Z digest=sha256:60da8f3bb295c74fa263c594c50d0ea71a656d2e8752dd369936c7a5aee6514d

Observation a9206d5f-e916-4bcb-9e10-25881c872060 · outbound

This paper cites Language models are unsupervised multitask learners.

Multimodal Autoregressive Pre-training of Large Vision Encoders Language models are unsupervised multitask learners

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:17:23.574339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:17:22.719677Z digest=sha256:cc6dd4102bb533100fea9fd6031c5944c4f2608fdec997aefa540d30c565f4dd

Observation 73d0b2e4-c04c-459c-9fa9-1cc43ff0898a · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

Multimodal Autoregressive Pre-training of Large Vision Encoders Learn- ing transferable visual models from natural language super- vision

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:17:23.564349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:17:22.722945Z digest=sha256:b76d9506fd03efc941f89843c55c47e9718cea22029e57f4311e0ce9bd4df77e

Observation 31576629-c6c8-42ba-b0e3-b70ab8335efc · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.

Multimodal Autoregressive Pre-training of Large Vision Encoders Exploring the limits of transfer learning with a unified text-to-text transformer

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:17:23.554504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:17:22.726126Z digest=sha256:0fe20e54bec76036d94eea3540655e2049333849a257689b7c4274abaca4a8ca

Observation 95e33792-7b00-4930-85a6-636fcca93cbf · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Multimodal Autoregressive Pre-training of Large Vision Encoders Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:22.730042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:22.730042Z digest=sha256:1c85e1876cab14bc441bd298cd300fac680934d11e187298b1763cf280285fba

Observation f7bc0033-fea6-4e3e-b576-b92d6dd8c2c4 · outbound

This paper cites ImageNet-21K Pretraining for the Masses.

Multimodal Autoregressive Pre-training of Large Vision Encoders ImageNet-21K Pretraining for the Masses

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:22.734482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:22.734482Z digest=sha256:b2b46d357e06e432551c5ab9d452afe992686e95e1b54e1fb148941065f5fedd

Observation a4e7862f-5c31-4469-9a18-e0f51c30bcef · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Multimodal Autoregressive Pre-training of Large Vision Encoders High-resolution image synthesis with latent diffusion models

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:22.738582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:22.738582Z digest=sha256:3a042076efadc03d39ecc921fda4ff82d1597b765c94eac25177e3bcddce5798

Observation 73459927-ae09-4c30-b219-c6542dfc8834 · outbound

This paper cites Learning visual representations with caption annotations.

Multimodal Autoregressive Pre-training of Large Vision Encoders Learning visual representations with caption annotations

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:17:23.537678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:17:22.742118Z digest=sha256:1b0b37dd2b8963542c60241aed32f1176ffa8f37b0934ce3f461fc285ecd14d3

Observation 6832edac-8faf-42c4-8946-8bedddc15daf · outbound

This paper cites Laion- 400m: Open dataset of clip-filtered 400 million image-text pairs.

Multimodal Autoregressive Pre-training of Large Vision Encoders Laion- 400m: Open dataset of clip-filtered 400 million image-text pairs

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:17:23.526452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:17:22.745434Z digest=sha256:0090f4598f5fa15afd24714574997fb73a399bc96530dddf899fe6cb24e9a8fa

Observation a6202d2b-30c8-4fd7-b472-2a0244c38e96 · outbound

This paper cites Objects365: A large-scale, high-quality dataset for object detection.

Multimodal Autoregressive Pre-training of Large Vision Encoders Objects365: A large-scale, high-quality dataset for object detection

Reference 101

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:17:23.516189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:17:22.748659Z digest=sha256:2f9f52f0edeb977e58b87ae66ff63ef22f29e3f7e7e780988eb98049e3487915

Pith citing papers

Observation b9a9bf37-d122-4f11-942b-87fddc10995f · inbound

Analyzing Finetuning Representation Shift for Multimodal LLMs Steering cites this paper.

Analyzing Finetuning Representation Shift for Multimodal LLMs Steering Multimodal Autoregressive Pre-training of Large Vision Encoders

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T22:03:25.012606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:03:25.012606Z digest=sha256:0ccf7dd230df8df28ffe152f2e48d4644c5e84770880e8d5a4b0c96f76a41c5c

Observation 35ae0f1d-efac-48c4-b0fa-6a96058e3f29 · inbound

Evaluating Sample Utility for Efficient Data Selection by Mimicking Model Weights cites this paper.

Evaluating Sample Utility for Efficient Data Selection by Mimicking Model Weights Multimodal Autoregressive Pre-training of Large Vision Encoders

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T21:02:00.842792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:02:00.842792Z digest=sha256:caffcab64b80b094d8171c5ea734a506546373e24adf4dc3bdca57eca0984622

Observation 11613278-2c26-4b96-9ae0-c64f4c515b2a · inbound

Vision-Language Models Do Not Understand Negation cites this paper.

Vision-Language Models Do Not Understand Negation Multimodal Autoregressive Pre-training of Large Vision Encoders

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T20:08:46.054919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:08:46.054919Z digest=sha256:3d412650a5b3601f7f7432e8d3694decd30314ae7df96a665568107285234d7f

Observation 38860a84-a34c-4cd6-9d80-ac34c38cf331 · inbound

Visual RAG: Expanding MLLM visual knowledge without fine-tuning cites this paper.

Visual RAG: Expanding MLLM visual knowledge without fine-tuning Multimodal Autoregressive Pre-training of Large Vision Encoders

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T18:59:22.319238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:59:22.319238Z digest=sha256:f7b7d6f2cb6a52e85b1756de739b77c1840ec26475dc220facc957b15f55d333

Observation fe39133e-42aa-4f9e-83a9-02747cc62934 · inbound

From Visuals to Vocabulary: Establishing Equivalence Between Image and Text Token Through Autoregressive Pre-training in MLLMs cites this paper.

From Visuals to Vocabulary: Establishing Equivalence Between Image and Text Token Through Autoregressive Pre-training in MLLMs Multimodal Autoregressive Pre-training of Large Vision Encoders

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T22:43:59.386454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T22:43:59.386454Z digest=sha256:11930947b6ef458eafe077669ffd2c4747cccd90c08bb32136fd8500a6cf91ef

Observation 8c3bff0d-e2ae-4cf5-a8d7-9c7ecdd809c9 · inbound

SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features cites this paper.

SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features Multimodal Autoregressive Pre-training of Large Vision Encoders

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-10T15:49:22.317639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T15:49:22.279848Z digest=sha256:67a078f2e4870b1627bd102b8f2b5a3b958924f207a236fd683e66f5b351b1cc

Observation db3b232f-c938-47f5-8409-9c61c4a76f1f · inbound

Seeing is Understanding: Unlocking Causal Attention into Modality-Mutual Attention for Multimodal LLMs cites this paper.

Seeing is Understanding: Unlocking Causal Attention into Modality-Mutual Attention for Multimodal LLMs Multimodal Autoregressive Pre-training of Large Vision Encoders

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-23T01:25:16.484883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-23T01:23:01.892612Z digest=sha256:1cde2ddd84c08bff4e6bba92418645ac8f6199e26a0f5eac5a0ce503d966cb14

Observation 1692588c-b51e-4fdc-ae1b-c2cdc1f82b46 · inbound

VISTA: Enhancing Vision-Text Alignment in MLLMs via Cross-Modal Mutual Information Maximization cites this paper.

VISTA: Enhancing Vision-Text Alignment in MLLMs via Cross-Modal Mutual Information Maximization Multimodal Autoregressive Pre-training of Large Vision Encoders

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-15T21:07:59.768232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:07:59.768232Z digest=sha256:bab3fe389bc27be177b0732795735182c15dbeafe0438bc5208dc4629b47e142

Observation c8812366-15d3-45be-8a37-11d69eec5b6e · inbound

Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach cites this paper.

Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach Multimodal Autoregressive Pre-training of Large Vision Encoders

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T15:39:33.091881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:39:33.091881Z digest=sha256:183c9cd6fe7695a846d3492cdc2814b45fc6d681538f3222f207e87517b0c812

Observation 0728f44c-cc8d-4784-b3e7-b2147815bf8c · inbound

Advancements in Medical Image Classification through Fine-Tuning Natural Domain Foundation Models cites this paper.

Advancements in Medical Image Classification through Fine-Tuning Natural Domain Foundation Models Multimodal Autoregressive Pre-training of Large Vision Encoders

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:09:15.985872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:09:15.985872Z digest=sha256:13148b46011005c7784f72f50681f15476e991c2e0faea8af3b7c45dc6ed9d35

Observation 73860be0-67e3-4044-ad80-3203f726ffc8 · inbound

CuRe: Cultural Gaps in the Long Tail of Text-to-Image Systems cites this paper.

CuRe: Cultural Gaps in the Long Tail of Text-to-Image Systems Multimodal Autoregressive Pre-training of Large Vision Encoders

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:16.016313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:16.016313Z digest=sha256:a8aafa4f87a06b2b5c047f11017d38fd6381afb795c5b2c3574014b7f9e9dee4

Observation 8f3a1ec8-383f-4a72-92bd-bacd9751eb5e · inbound

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models cites this paper.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models Multimodal Autoregressive Pre-training of Large Vision Encoders

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:13:02.851822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:b287421a573a0c1b65b4d90a71713a53e41bd81c1977c8b006f15d4b0bb8957a

Observation 0d1cdbad-61e0-4aa1-8757-c176ebea21e6 · inbound

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning cites this paper.

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning Multimodal Autoregressive Pre-training of Large Vision Encoders

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T00:33:50.796334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T00:33:50.471804Z digest=sha256:29957034ef9cad0f3525e20dae98694a4d47f358774a33d798f292623d2f3db7

Observation 420672e1-1066-4026-9620-448ba0bb8c5f · inbound

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement cites this paper.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Multimodal Autoregressive Pre-training of Large Vision Encoders

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:04.264739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:04.264739Z digest=sha256:cb95806a1f91eba380eba8718c117fbb1e4915314453253b9fb9d74b46c62e3e

Observation a26c49c9-2976-4085-b180-0cb9d3af72cb · inbound

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models cites this paper.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Multimodal Autoregressive Pre-training of Large Vision Encoders

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:28.632831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:28.632831Z digest=sha256:d29589a453106b7004e8d554301272357668b46be0692fc68d5916614d5cd6a0

Observation 0318e99b-ce9f-42c5-bb61-293b850ed45c · inbound

OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning cites this paper.

OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning Multimodal Autoregressive Pre-training of Large Vision Encoders

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T12:22:33.329919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:22:33.329919Z digest=sha256:aef651630e619c815dd259ccdd45036b7d27183b266740bce56d38567038481c

Observation 08145ee1-f2ae-4ef8-ac64-74f29d294ab4 · inbound

Hierarchical Pre-Training of Vision Encoders with Large Language Model cites this paper.

Hierarchical Pre-Training of Vision Encoders with Large Language Model Multimodal Autoregressive Pre-training of Large Vision Encoders

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T05:37:50.158868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:37:50.158868Z digest=sha256:cf6134802319a1f0c71f135d7eca8ff7dbbfff545c852f9b68bfbca7537a07f2

Observation bb0a169b-0d0c-4acd-9423-c1c812e585c2 · inbound

Towards Generalizable Deepfake Image Detection with Vision Transformers cites this paper.

Towards Generalizable Deepfake Image Detection with Vision Transformers Multimodal Autoregressive Pre-training of Large Vision Encoders

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T06:41:36.897978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T06:36:56.465272Z digest=sha256:403ff535358d9f230fb6c62d1deee21142713ffd035cf35f0cff06441bcf5fed

Observation 70d6a3f6-3ad7-4717-bdb5-3abddc3e4ec1 · inbound

SigLIP-HD by Fine-to-Coarse Supervision cites this paper.

SigLIP-HD by Fine-to-Coarse Supervision Multimodal Autoregressive Pre-training of Large Vision Encoders

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-13T02:39:58.984385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T02:39:58.984385Z digest=sha256:ee30f9791ca96f8bbbc68b9b6e592c7ddf0b3be4527dee3851ef46beb7051be7

Observation 552d91c6-28e5-46c2-a490-ea569baa9126 · inbound

Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model cites this paper.

Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model Multimodal Autoregressive Pre-training of Large Vision Encoders

Reference 107

Resolution
unresolved
no resolver link, observed 2026-07-31T06:20:14.033388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:20:14.033388Z digest=sha256:f4d286442637064ea983b4314dcdee7ff65af7205606ba87e63c53ba708da470