Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T15:17:22.748659Z
Paper Citation Record · LEDGER
As of 15 August 2026, this Paper Citation Record lists 100 of 137 outbound references and 20 inbound Pith citation observations for arXiv:2411.14402.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T15:17:22.748659Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T21:07:59.768232Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-23T01:25:16.481218Z
100 of 137 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 9b4ca4f7-aa21-4a23-b87f-385b1a76dfe2 · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5bc970ba-21f5-412f-9762-c173c791013b · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders Nocaps: Novel object cap- tioning at scale
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31e2e298-fb61-4d80-a9b0-73b813ff38a4 · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders Flamingo: a visual language model for few-shot learning
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b78b674-1330-43fe-af85-fe247ef76a06 · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30508ccc-e3ba-41ee-a202-96ea9ac6f0d6 · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders Qwen Technical Report
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fdae0bd5-2d23-4638-8d52-81c17daf68d3 · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ca328a1-9710-41a6-a5d2-fec628540c1f · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders From detection of individual metastases to classification of lymph node status at the pa- tient level: the camelyon17 challenge
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0aaf7d9-e904-46f6-ac68-6ad90c6eac8e · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders BEiT: Bert pre- training of image transformers
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bcf6dafb-1f3d-476c-ba2e-4b2f8ae9efe9 · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders Flexivit: One model for all patch sizes
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1e0d4bb-7189-4809-b5e2-fee77737d066 · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders Food-101 – mining discriminative components with ran- dom forests
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fea36b10-7818-4644-b7e3-81ca5fc6a9a1 · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders Time series analysis: forecasting and control
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1bfeb3e2-5170-4879-937f-83f4dd8060c3 · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders Language Models are Few-Shot Learners
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation faf41b7a-c743-442e-83a2-6ecbf8723a24 · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders Coyo-700m: Image-text pair dataset, 2022
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cfe5a846-2afb-449b-83c3-ba8d4b50bd97 · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders End-to-end object detection with transformers
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9aae2335-5283-4c6f-9ae1-08d9149961fc · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders Unsupervised learn- ing of visual features by contrasting cluster assignments
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b02d1d74-2054-4069-a87c-e5af80c729c7 · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders Emerging properties in self-supervised vision transformers
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd86e4ca-80fa-4289-b35c-fe7d84ec4339 · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders A generative approach for wikipedia-scale visual entity recognition
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82341a4f-b06c-4602-9c52-85318551e01e · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders Mmdetection: Open mmlab detection toolbox and benchmark
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f693463-45a9-4f47-9598-bf16e1eecb30 · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders Generative pre- training from pixels
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49f7072e-60e7-44f1-bc1e-2b55fb7d9bd3 · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders A simple framework for contrastive learning of visual representations
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98afe1f3-3b4f-4e56-a41f-3f6d42fbaaab · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders Microsoft COCO Captions: Data Collection and Evaluation Server
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2ad7346-77b2-48b3-a3da-8878fb5733a2 · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders Gonzalez, Ion Stoica, and Eric P
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d45c89c-538e-4a6a-9985-3d4791db9511 · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders PaLM: Scaling Language Modeling with Pathways
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd188d22-d450-478e-b068-16a62335877f · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders Functional map of the world
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80c5bd87-eacb-4ac9-8091-3d0ea2648a1f · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders Cimpoi, S
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2aca9d57-3dea-4db5-a53a-f39a78fcc253 · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders Patch n’pack: Navit, a vision trans- former for any aspect ratio and resolution
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91829b37-0168-497f-ae80-517a6d679241 · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders Imagenet: A large-scale hierarchical im- age database
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4eb682a7-1c36-42bf-b02a-acaa29fea8cd · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders Virtex: Learning visual representations from textual annotations
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 690aec9b-61b4-4662-97f4-54e7a7285180 · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders Unsu- pervised visual representation learning by context predic- tion
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae746e19-b83e-4bf3-8fce-8ce1df98e0bd · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders An image is worth 16x16 words: Trans- formers for image recognition at scale
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8a7aa3d-9257-4e57-a637-4714d1977bdc · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders The Llama 3 Herd of Models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1992e40d-ba70-456d-b241-0dcf513c88bc · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders Scalable Pre-training of Large Autoregressive Image Models
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5fb55291-be1a-4bbe-ba82-cf079e8a24dd · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders Taming transformers for high-resolution image synthesis
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1765765-aa45-4efa-a01b-f7b35f3a3339 · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders Data Filtering Networks
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 088deace-bdcd-43f2-bbe0-d76de2778712 · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders Slowfast networks for video recognition
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf53a0a5-6d35-4b63-9e11-ce2d0c64349b · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders Improved baselines for vision-language pre-training
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0b8ea7e-0c8b-4fb4-a487-a74684970089 · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d70bbe70-3d40-42b7-9f27-f293dfce0bef · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders Unsupervised Representation Learning by Predicting Image Rotations
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86b772e0-2934-4f93-8c38-6549c91ecee1 · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders Making the v in vqa matter: El- evating the role of image understanding in visual question answering
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8fa25754-c13e-4cf0-8f5b-1b7b38ad7865 · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders Making the v in vqa matter: El- evating the role of image understanding in visual question answering
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23c64246-7b44-458a-a79c-1226b5b4ce5b · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders Bootstrap your own latent-a new approach to self-supervised learning
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ff49287-8071-47bb-8fb6-68544ef8dff3 · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders Lvis: A dataset for large vocabulary instance segmentation, 2019
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13d5870a-e29d-4229-8d66-f7f743879f10 · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders Vizwiz grand challenge: Answering visual questions from blind people
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12e83030-6b2d-4746-ab7c-d33e8bea2616 · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders Deep residual learning for image recognition
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ee7a35d-744a-4062-99a0-2a58a7fd32ce · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders Mask r-cnn
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c28c0b0-9a26-4489-b643-e5c3febe66e2 · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders Mask r-cnn
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd3291c8-42a7-4c0c-9baa-875300758e53 · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders Masked autoencoders are scal- able vision learners
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80e27a5a-c5da-485f-8275-34714b1e1286 · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders Eurosat: A novel dataset and deep learn- ing benchmark for land use and land cover classification,
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb1403ae-34f0-443f-853c-5cd826108dab · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders Training Compute-Optimal Large Language Models
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05e366ef-30f2-4008-85a3-79ad81433cf5 · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders Scaling up vision-language pre-training for image captioning
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63a60023-7453-4440-b278-620ce14c247a · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders Gqa: A new dataset for real-world visual reasoning and compositional question answering
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d39d5962-7672-41ff-babf-2bc86d1b0491 · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders Gqa: A new dataset for real-world visual reasoning and compositional question answering
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0aecd157-d05d-41ad-a0ea-d033fc9b4533 · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders Scaling up visual and vision-language representation learning with noisy text supervision
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5748ad72-33e4-4763-a8aa-98118d65e4bc · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders Scaling Laws for Neural Language Models
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1318717-68e9-4439-a9f8-c0befa5cc540 · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders Deep visual-semantic alignments for generating image descriptions
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b351665-a929-4a3d-bba2-0eff118a7108 · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders Referitgame: Referring to objects in pho- tographs of natural scenes
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e5807bb-5974-4221-afa7-8045de742895 · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders Big transfer (bit): General visual representation learning
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76009505-4fbb-413b-8663-1e5308d636da · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders 3d object representations for fine-grained categorization
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9cdf6400-8965-4acb-8341-ddc5a06bcbe8 · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders Learning multiple layers of features from tiny images
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9f089a3-89e1-425f-b757-0011b3f42c5b · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders Imagenet classification with deep convolutional neural net- works
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc337774-0b21-459f-bd9e-05cb3c7dac15 · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders MaMMUT: A Simple Architecture for Joint Learning for MultiModal Tasks
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a833fbd6-1add-4ddf-8647-8e803f910412 · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders Revisit Large-Scale Image-Caption Data in Pre-training Multimodal Foundation Models
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa4c8a09-bab1-4142-96d4-0199b6d75596 · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 132390e8-8aae-4c4b-aea4-1ae3368f32fe · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders Align before fuse: Vision and language representation learning with momentum distillation
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 560a87bd-409a-4538-b435-258c73db3539 · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders Blip: Bootstrapping language-image pre-training for uni- fied vision-language understanding and generation
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8198cf4a-a4ef-4afd-a553-3be1aabad19b · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca82604f-08b2-4827-b329-88128bad07bf · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders Exploring plain vision transformer backbones for object de- tection
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7f349a3-6580-43ca-8b2e-a5e15848f4fd · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders Exploring plain vision transformer backbones for object de- tection, 2022
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4202a41f-ab2b-4e18-a454-ba210cf316be · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac9ce47e-8e0b-4bd7-9f98-13c7930fda00 · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders Microsoft coco: Common objects in context
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9afdd6f1-283f-4598-896e-be3439bff8dd · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f866b314-9cce-43a4-9011-b15f6c55691f · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders Improved baselines with visual instruction tuning
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2ae77fa-0722-40cd-9765-4d6ad0daf1a8 · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders Grounding dino: Mar- rying dino with grounded pre-training for open-set object detection
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0047e8f2-658e-4f8f-b9e9-c3ddba6b01c7 · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders Sgdr: Stochastic gradient descent with warm restarts
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4fce8aa-c5aa-4817-be8f-0be5cd4befb5 · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders Decoupled Weight Decay Regularization
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1697871-86ca-4499-a580-a7acba5fca1f · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders Unified-io 2: Scaling autoregressive mul- timodal models with vision language audio and action
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90824477-ae47-423c-bdf4-e46c9fa7e2ec · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders Learn to explain: Multimodal reasoning via thought chains for science question answering
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52fa79ac-f9ad-4a93-9289-df5ffdd51b15 · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders Generation and comprehension of unambiguous object descriptions
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ca744f4-295d-42da-b98d-afaa0a98c49f · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders Ok-vqa: A visual question answering benchmark requiring external knowledge
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 0abe84aa-6dee-4f3f-9be3-220987b200e0 · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders Ok-vqa: A visual question answering benchmark requiring external knowledge
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation bea6496d-aa2e-4caf-b0d2-3b82603a3a1d · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8ecadcb-5000-4dbe-bf0a-f24340b5c8a3 · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders Docvqa: A dataset for vqa on document images
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a853ea2a-9a0c-4984-8eac-8852141e4573 · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders Infographicvqa
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 36a14cf9-c4b1-4240-aa1a-b7e7ca2a5d02 · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d64046b-f57e-4bfb-ac4a-b991eab63b52 · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders Unsupervised learning of visual representations by solving jigsaw puzzles
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77b4acc1-b1e3-4955-940a-b30e1a2b54a2 · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders Unresolved cited work
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 056ed4f9-b7c3-48f7-8054-b0d53f8977a4 · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders Unresolved cited work
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 981aae9a-5948-410d-9a05-2c1bcc9ca78f · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders Moment matching for multi-source domain adaptation
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 837b0f71-f317-4159-90cc-57ef100438dc · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders Plummer, Liwei Wang, Christopher M
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 70cf87f0-7534-407b-90ba-181f700b07fc · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders Dataset Decomposition: Faster LLM Training with Variable Sequence Length Curriculum
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0fd77eb-00f8-4758-a39b-cb163db3bb9b · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders Improving language understanding by genera- tive pre-training
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a9206d5f-e916-4bcb-9e10-25881c872060 · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders Language models are unsupervised multitask learners
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 73d0b2e4-c04c-459c-9fa9-1cc43ff0898a · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders Learn- ing transferable visual models from natural language super- vision
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 31576629-c6c8-42ba-b0e3-b70ab8335efc · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders Exploring the limits of transfer learning with a unified text-to-text transformer
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 95e33792-7b00-4930-85a6-636fcca93cbf · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7bc0033-fea6-4e3e-b576-b92d6dd8c2c4 · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders ImageNet-21K Pretraining for the Masses
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4e7862f-5c31-4469-9a18-e0f51c30bcef · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders High-resolution image synthesis with latent diffusion models
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73459927-ae09-4c30-b219-c6542dfc8834 · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders Learning visual representations with caption annotations
Reference 99
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 6832edac-8faf-42c4-8946-8bedddc15daf · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders Laion- 400m: Open dataset of clip-filtered 400 million image-text pairs
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a6202d2b-30c8-4fd7-b472-2a0244c38e96 · outbound
Multimodal Autoregressive Pre-training of Large Vision Encoders Objects365: A large-scale, high-quality dataset for object detection
Reference 101
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation b9a9bf37-d122-4f11-942b-87fddc10995f · inbound
Analyzing Finetuning Representation Shift for Multimodal LLMs Steering Multimodal Autoregressive Pre-training of Large Vision Encoders
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35ae0f1d-efac-48c4-b0fa-6a96058e3f29 · inbound
Evaluating Sample Utility for Efficient Data Selection by Mimicking Model Weights Multimodal Autoregressive Pre-training of Large Vision Encoders
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11613278-2c26-4b96-9ae0-c64f4c515b2a · inbound
Vision-Language Models Do Not Understand Negation Multimodal Autoregressive Pre-training of Large Vision Encoders
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38860a84-a34c-4cd6-9d80-ac34c38cf331 · inbound
Visual RAG: Expanding MLLM visual knowledge without fine-tuning Multimodal Autoregressive Pre-training of Large Vision Encoders
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe39133e-42aa-4f9e-83a9-02747cc62934 · inbound
From Visuals to Vocabulary: Establishing Equivalence Between Image and Text Token Through Autoregressive Pre-training in MLLMs Multimodal Autoregressive Pre-training of Large Vision Encoders
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c3bff0d-e2ae-4cf5-a8d7-9c7ecdd809c9 · inbound
SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features Multimodal Autoregressive Pre-training of Large Vision Encoders
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation db3b232f-c938-47f5-8409-9c61c4a76f1f · inbound
Seeing is Understanding: Unlocking Causal Attention into Modality-Mutual Attention for Multimodal LLMs Multimodal Autoregressive Pre-training of Large Vision Encoders
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 1692588c-b51e-4fdc-ae1b-c2cdc1f82b46 · inbound
VISTA: Enhancing Vision-Text Alignment in MLLMs via Cross-Modal Mutual Information Maximization Multimodal Autoregressive Pre-training of Large Vision Encoders
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8812366-15d3-45be-8a37-11d69eec5b6e · inbound
Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach Multimodal Autoregressive Pre-training of Large Vision Encoders
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0728f44c-cc8d-4784-b3e7-b2147815bf8c · inbound
Advancements in Medical Image Classification through Fine-Tuning Natural Domain Foundation Models Multimodal Autoregressive Pre-training of Large Vision Encoders
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73860be0-67e3-4044-ad80-3203f726ffc8 · inbound
CuRe: Cultural Gaps in the Long Tail of Text-to-Image Systems Multimodal Autoregressive Pre-training of Large Vision Encoders
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f3a1ec8-383f-4a72-92bd-bacd9751eb5e · inbound
AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models Multimodal Autoregressive Pre-training of Large Vision Encoders
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 0d1cdbad-61e0-4aa1-8757-c176ebea21e6 · inbound
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning Multimodal Autoregressive Pre-training of Large Vision Encoders
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 420672e1-1066-4026-9620-448ba0bb8c5f · inbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Multimodal Autoregressive Pre-training of Large Vision Encoders
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a26c49c9-2976-4085-b180-0cb9d3af72cb · inbound
LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Multimodal Autoregressive Pre-training of Large Vision Encoders
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0318e99b-ce9f-42c5-bb61-293b850ed45c · inbound
OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning Multimodal Autoregressive Pre-training of Large Vision Encoders
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08145ee1-f2ae-4ef8-ac64-74f29d294ab4 · inbound
Hierarchical Pre-Training of Vision Encoders with Large Language Model Multimodal Autoregressive Pre-training of Large Vision Encoders
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb0a169b-0d0c-4acd-9423-c1c812e585c2 · inbound
Towards Generalizable Deepfake Image Detection with Vision Transformers Multimodal Autoregressive Pre-training of Large Vision Encoders
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 70d6a3f6-3ad7-4717-bdb5-3abddc3e4ec1 · inbound
SigLIP-HD by Fine-to-Coarse Supervision Multimodal Autoregressive Pre-training of Large Vision Encoders
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 552d91c6-28e5-46c2-a490-ea569baa9126 · inbound
Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model Multimodal Autoregressive Pre-training of Large Vision Encoders
Reference 107
Source-reported events for the cited work
Unavailable: canonical work link unavailable.