Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T17:21:04.547305Z
Paper Citation Record · LEDGER
As of 20 August 2026, this Paper Citation Record lists 100 of 104 outbound references and 22 inbound Pith citation observations for arXiv:2501.12327.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T17:21:04.547305Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T11:48:03.500362Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-08T00:04:22.542063Z
100 of 104 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 02ef5ae7-22ec-402d-8993-4ca00f815ea1 · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Albergo and Eric Vanden-Eijnden
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7091db3-b8a5-4832-b0ac-8719fe9ba569 · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 687c7d63-90b1-4a60-b1ea-791b5c9e0375 · outbound
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78110544-edc4-4cf0-a5c4-ebec50a60318 · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Analytic- dpm: an analytic estimate of the optimal reverse variance in diffusion probabilistic models, 2022
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ec4130f-6976-4b60-ac31-064da350ec4b · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c19ee446-7ef1-4332-9b85-34e2226e1571 · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model HALC: Object Hallucination Reduction via Adaptive Focal-Contrast Decoding
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eaf80934-d9d7-4a17-8708-d5024abaf709 · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e00e5a9b-bcb2-493e-b511-a2ae577baac3 · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c819938f-f686-4df2-bff3-4272318f25c8 · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Instructblip: Towards general- purpose vision-language models with instruction tuning,
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e567638-7001-4bb0-a5a9-4c0278796200 · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Deepseek-v3 technical report, 2024
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc5976e8-5f6c-4489-8476-096b526b951e · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Imagenet: A large-scale hierarchical image database
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ccaabffa-b2e6-4969-a60d-0cf3c5f7e961 · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Cogview: Mastering text-to- image generation via transformers, 2021
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2ba1748-3173-4c43-9221-9c2e8c359b3c · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Dreamllm: Synergistic multimodal compre- hension and creation, 2024
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8187f5d8-4eae-4649-b2f8-30c8dfab353a · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Taming transformers for high-resolution image synthesis, 2021
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b7f01ba-b9f6-4e40-a94a-3baf2bb28b9f · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Fluid: Scaling Autoregressive Text-to-image Generative Models with Continuous Tokens
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b049a85-1fee-4821-b33f-9a53edb8b170 · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Mme: A comprehen- sive evaluation benchmark for multimodal large language models, 2024
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7db3e256-48c4-40de-8a5e-8aa20d838f68 · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Making llama see and draw with seed tokenizer, 2023
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7e95b52-af29-49cd-8716-f2adb530d859 · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 395f0c5e-97e6-4b70-9bcb-f6804c9c849e · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Making the v in vqa matter: Ele- vating the role of image understanding in visual question answering, 2017
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad7bb450-fe95-49f0-bc2c-078aa624bb27 · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de348b17-ce73-4b05-a90d-647049fd6fb4 · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Stangl, Anhong Guo, Chi Lin, Kristen Grauman, Jiebo Luo, and Jeffrey P
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd96de1b-a8e3-458e-94ae-45c05faaf459 · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Infinity: Scaling bitwise autoregressive modeling for high-resolution image synthesis, 2024
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83199dc8-7144-4c07-b62f-6f31345dea6e · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Scaling Laws for Autoregressive Generative Modeling
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a913b5d-76b8-4ff4-9ac7-5456d24a75dd · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Denoising diffu- sion probabilistic models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e14c0d7f-ef16-40b3-893b-5cb64c3118df · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Denoising diffu- sion probabilistic models, 2020
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c710cb36-5d00-4c49-99b7-f9ba42f719a9 · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Fleet, Mohammad Norouzi, and Tim Salimans
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05c48a21-3f0f-415d-b190-9e083acc0075 · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Gqa: A new dataset for real-world visual reasoning and compositional question answering
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 077c2209-9771-4993-9318-d9fca4d48144 · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Hudson and Christopher D
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1473683f-a911-4071-a4b0-79ac04e25f79 · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Scaling Laws for Neural Language Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a59168d2-cd0d-417f-a551-f25d6a825290 · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Referitgame: Referring to objects in pho- tographs of natural scenes
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation edafc5d2-ee27-42f5-a95d-57ac40f5108d · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Visual genome: Connecting language and vision using crowdsourced dense image annotations
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 76c23acd-d042-4fa4-92df-e026d85caa07 · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 6b1f51eb-97d4-4ad2-8abc-b868ad943db6 · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Unresolved cited work
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation a2674656-240c-47d0-8726-57db70b158b6 · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Datasets: imagenet-1k-vl-enriched
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation b1af209a-5da4-4f7b-9133-5f254b5b1a4a · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Deep learning
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation ebbe9e8b-7c63-4326-914a-90bf4f6025bd · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Autoregressive image generation using residual quantization, 2022
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation c5d0fb22-0895-45e8-8094-17af05f79de6 · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Mitigating Object Hallucinations in Large Vision-Language Models through Visual Contrastive Decoding
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c690a647-abe8-4335-87f0-ebd032a52f8a · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Seed-bench: Benchmarking multimodal llms with generative comprehension, 2023
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 5840bd26-27e1-476a-acc6-65ef61c90cda · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Llava-next: What else influences visual instruction tuning beyond data?, 2024
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation a75f6018-7d76-4c07-888b-d6a32297b8e0 · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Llava-next: Stronger llms supercharge multimodal capabilities in the wild, 2024
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 94b0743b-231f-45f1-946c-d4c93240bdb2 · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Llava-onevision: Easy visual task transfer, 2024
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 58bda4e4-2289-47ce-ad53-f14da0f948d4 · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Llava-next: Tackling multi-image, video, and 3d in large multimodal models,
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation d0023b5b-6601-469d-a346-fd2e97af8db9 · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Blip- 2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 53ab4899-7181-43f8-a309-a5e25a2ed6b7 · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Blip- 2: Bootstrapping language-image pre-training with frozen image encoders and large language models, 2023
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 3b873459-04b5-4fb5-b0f4-0d25d8fa23a4 · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Evaluating object hallucination in large vision- language models
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 2c4a206b-008a-4440-ad1f-e8d0e1bc9402 · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Dual diffusion for unified image generation and understanding, 2024
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 7f25221c-d942-4d8b-919a-cd855aff610c · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Improved Baselines with Visual Instruction Tuning
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bfdfb097-5980-4de6-83ca-b0cf21c628b3 · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Visual Instruction Tuning
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 878d5f9d-8086-4274-92d0-89b11770d923 · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Visual instruction tuning
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 401fee1b-c74a-4080-9fdf-84cc84033a6c · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Improved baselines with visual instruction tuning
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 1e21cf32-bc91-4ed1-bd31-3ecdcce2692c · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Improved baselines with visual instruction tuning, 2024
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 821f2ff4-a50b-4d46-86e5-e27b39403060 · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 367b11ed-4739-47f0-852c-cdacab181abc · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model World model on million-length video and language with ringattention
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation ee285e19-6127-4fdf-be4f-100c99da170e · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Mmbench: Is your multi-modal model an all-around player?, 2024
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 93a6883f-d62b-40bd-b416-7dbb7e0339c3 · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Dpm-solver: A fast ode solver for diffusion probabilistic model sampling in around 10 steps, 2022
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 0232c3dd-4465-4a54-a616-ba7197d81649 · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Learn to explain: Multimodal reasoning via thought chains for science question answering, 2022
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 84008c93-d13a-40c1-97d6-d1ef4d9b0602 · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Unified multi-modal latent diffusion for joint subject and text conditional image generation, 2023
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 4f9e81fe-ab55-47af-8ea5-e797451337d0 · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Generation and comprehension of unambiguous object descriptions
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed88223d-20e4-4c11-bdf6-8c549e1d9350 · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Ok-vqa: A visual question answering benchmark requiring external knowledge, 2019
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 5f1c9815-c403-4bae-99fb-b01ac8f9e695 · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Ok-vqa: A visual question answering benchmark requiring external knowledge
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation d3395c2d-6c74-44b1-aeea-bb2680354427 · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Ocr-vqa: Visual question answering by reading text in images
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2949aca0-29ff-4e41-ab65-d58979c19940 · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Improved denoising diffusion probabilistic models
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6adfc238-4cb5-4864-b2ab-f49da80dc63b · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Sdxl: Improving latent diffusion models for high-resolution image synthesis, 2023
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4f27e65-f17b-4304-a135-edd8c59b979e · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Du, Zehuan Yuan, and Xin- glong Wu
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 3cf0f681-042d-4c4c-9b7a-5ab69a9c262a · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Improving language understanding by generative pre-training
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e957856-07e5-46f2-9882-ad0d28dbbb64 · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Learning transferable visual models from natural language supervision, 2021
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5dedcf49-fcdc-44c6-a8a6-a2a7723e72b1 · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Zero-shot text-to-image generation, 2021
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 672f3e67-c3be-4d75-a9e2-eabc8f8a6676 · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Hierarchical text-conditional image genera- tion with clip latents, 2022
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 504fc21f-a55a-4ce4-8205-1b4882b92f90 · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model High-resolution image synthesis with latent diffusion models, 2021
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 64eed9d4-bce1-4264-95ed-7368ce130d2f · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model A-okvqa: A bench- mark for visual question answering using world knowledge
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 3b30c682-79a9-4642-8646-a0cc7866efa5 · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model https://sharegpt.com/, 2023
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 4f535246-2cb0-49a0-b2bc-4420abf03743 · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Textcaps: a dataset for image captioning with reading comprehension
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 78a7022f-34f2-440a-af68-a4048fe072ce · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Towards vqa models that can read, 2019
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation c884e091-49c7-49de-bc2f-8a41d6eb0565 · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Denois- ing diffusion implicit models, 2022
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f336a288-88b0-4b96-9958-bab1ceaad257 · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Generative modeling by estimating gradients of the data distribution, 2020
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation f557c42d-75b3-414b-8d52-9bd3628b6b76 · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9534376e-6416-4d9e-a1db-112c70e201eb · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Autoregressive model beats diffusion: Llama for scalable image generation, 2024
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation f1b697cc-4b73-4fd7-a086-d4ba312325a1 · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Emu: Generative pretraining in multimodality, 2024
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 4284d52c-fe79-427f-9813-d5d51dfcddac · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Hart: Efficient visual generation with hybrid autoregressive transformer
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 33cd3e92-1192-4ad6-ac2f-e43e1264750d · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Any-to-any generation via composable diffusion, 2023
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation b7482b26-fc8c-480e-8e16-18dabf8f7f66 · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Chameleon: Mixed-Modal Early-Fusion Foundation Models
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b09cea6d-cd21-4e3a-b502-604c4227c340 · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Chameleon: Mixed-modal early-fusion foundation models, 2024
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 0eb9c253-a750-4dfe-8ce0-e714ac4f1155 · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Gemini: A family of highly capable multi- modal models, 2024
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation c46ec641-224c-4622-a994-0dbfbee052f9 · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Visual autoregressive modeling: Scalable image generation via next-scale prediction
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation afa5e832-8fdd-474a-b72b-2ffe528815ed · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model LLaMA: Open and Efficient Foundation Language Models
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 575930a0-480b-4283-8db9-9c67ae203c2f · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 223ddf3d-9771-4eec-aaca-9a5db3eccc20 · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Emu3: Next-Token Prediction is All You Need
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ccf4fa23-a0f7-4caf-98be-6f1a4a9cc6b0 · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Janus: Decoupling visual encoding for unified multimodal understanding and genera- tion, 2024
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 07804346-6d60-4570-81a4-816bb00d4b59 · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Liquid: Language Models are Scalable and Unified Multi-modal Generators
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2919fe17-42dd-4e1d-a91c-f17cf84ad08b · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model NExT-GPT: Any-to-Any Multimodal LLM
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b285d06-f445-495c-965d-fcd465fe69d6 · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Show-o: One Single Transformer to Unify Multimodal Understanding and Generation
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4acae73c-0414-4deb-bcfe-5f87fece7fba · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Show-o: One single transformer to unify multimodal understanding and generation, 2024
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 3192d077-b9ca-406b-99f0-7d193a3837d9 · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model X-vila: Cross-modality align- ment for large language model, 2024
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 7f74507e-3009-4f7d-be6e-dc1309ffbd58 · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c9753cf-45c7-4c39-bbb5-5c5c724de964 · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation a86774a0-f579-4365-b430-b82d493d0519 · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Woodpecker: Hallucination Correction for Multimodal Large Language Models
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0a5f092-939f-4f8f-904d-ecfc93be4057 · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Scaling autoregressive models for content-rich text-to-image generation, 2022
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e66f95f-5fb4-4782-b7b8-9f09cc37c952 · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Mmmu: A massive multi-discipline multimodal understand- ing and reasoning benchmark for expert agi, 2024
Reference 98
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 86b1424d-1762-47fc-8a10-54433d2ceff1 · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Lmms- eval: Reality check on the evaluation of large multimodal models, 2024
Reference 99
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 016c92f5-1ba3-4d55-b361-2ff766f765d6 · outbound
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Var-clip: Text-to-image generator with visual auto-regressive modeling, 2024
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation ce6f86c3-8ae3-464c-859c-31cbe8c6bfc5 · inbound
A Survey on Vision-Language-Action Models for Embodied AI VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model
Reference 141
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation eeeb5c56-95d2-4e74-ba1a-1502d5f79c9d · inbound
Do we really have to filter out random noise in pre-training data for language models? VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18734500-4d57-4696-84bd-10f7449a3a0c · inbound
Generative Multimodal Pretraining with Discrete Diffusion Timestep Tokens VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model
Reference 100
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3678c49c-a694-4464-8743-21f60760af38 · inbound
Open Set Domain Adaptation with Vision-language models via Gradient-aware Separation VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6b68b22-0eb4-4371-9e7a-5af03bd4b398 · inbound
MMaDA: Multimodal Large Diffusion Language Models VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 20f41297-a7fb-4c90-beca-1ca735b99dca · inbound
Draw ALL Your Imagine: A Holistic Benchmark and Agent Framework for Complex Instruction-based Image Generation VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 712f8acc-f2d9-4f5e-a1d2-08fc799d6f78 · inbound
Not All Tokens and Heads Are Equally Important: Dual-Level Attention Intervention for Hallucination Mitigation VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84a7f1c1-7cc2-4716-a4de-cae359319d48 · inbound
ComplexBench-Edit: Benchmarking Complex Instruction-Driven Image Editing via Compositional Dependencies VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd3cdd43-3777-40cd-aaf1-1e35af05d4e1 · inbound
Instella-T2I: Pushing the Limits of 1D Discrete Latent Space Image Generation VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 226e02d8-a9c5-477b-b69c-0990c34ccad1 · inbound
From Static Inference to Dynamic Interaction: A Survey of Streaming Large Language Models VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 07a519a3-d3da-46d2-b5f8-00102b18b19a · inbound
Demystifying Video Reasoning VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e988bc26-b644-4d96-aa89-0e5d9a344453 · inbound
Demystifying Video Reasoning VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf8003dd-c19c-4a18-adff-8a871acc3f23 · inbound
WinTok: A Win-Win Hybrid Tokenizer via Decomposing Visual Understanding and Generation with Transferable Tokens VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model
Reference 112
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 996a1a94-c94d-40c6-8ce7-64ed0331975c · inbound
Semantic Generative Tuning for Unified Multimodal Models VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 061a9d94-8747-45e3-a2a1-d89cc9b3bf21 · inbound
Semantic Generative Tuning for Unified Multimodal Models VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 39ee1ca7-63ea-49ae-acd4-15ccf867488d · inbound
HACK++: Towards More Effective Head-Aware Key-Value Compression for Efficient Visual Autoregressive Modeling VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 8263a698-25c9-4f07-839f-f4a8d6190a4d · inbound
Ask, Solve, Generate: Self-Evolving Unified Multimodal Understanding and Generation via Self-Consistency Rewards VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation e4e52c5d-0910-439e-b9d3-a27e11303402 · inbound
MEPA: Multi-Scale Representation Alignment for Visual Autoregressive Modeling with Mixture of Experts VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 158619c6-8924-48c6-82e8-3e8a88ce896a · inbound
Unified Audio Intelligence Without Regressing on Text Intelligence VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model
Reference 268
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation c4f225b9-d26d-4fc4-bedc-b1cc4b7b8c79 · inbound
Unified Audio Intelligence Without Regressing on Text Intelligence VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model
Reference 268
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2483e1c5-8177-4b4e-be7f-ee2e17676c4b · inbound
UniGen-AR: Unifying Visual Generation with Auto-Regressive Modeling VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e391844b-f528-4ed1-bb93-78e823d1782e · inbound
SynVAR: Synergizing Spatial and Semantic Alignment in Visual Autoregressive Model VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.