Pith. sign in

Paper Citation Record · LEDGER

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model

As of 20 August 2026, this Paper Citation Record lists 100 of 104 outbound references and 22 inbound Pith citation observations for arXiv:2501.12327.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.12327 v1

Coverage vector

measured 100 of 104 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T17:21:04.547305Z

measured 122 of 122 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 22 of 22 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:48:03.500362Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T00:04:22.542063Z

Reference resolution

100 of 104 outbound references displayed

  • verified exact0
  • verified fuzzy46
  • unresolved54
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 02ef5ae7-22ec-402d-8993-4ca00f815ea1 · outbound

This paper cites Albergo and Eric Vanden-Eijnden.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Albergo and Eric Vanden-Eijnden

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.055182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.055182Z digest=sha256:1af1513810f1303cfdac95343b65dfbe62ad04a9be58ab7ff5e049906e02e032

Observation a7091db3-b8a5-4832-b0ac-8719fe9ba569 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.060897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.060897Z digest=sha256:2d768992405b75617f4272d6fa553b3f1ddb58d4de24ab584f6300af79eed1ae

Observation 687c7d63-90b1-4a60-b1ea-791b5c9e0375 · outbound

This paper cites messages.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model messages

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.066545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.066545Z digest=sha256:f3b7bdf5dcd96724c9f57044b93265fdc3993d0aca50f194017221e6c0224a2d

Observation 78110544-edc4-4cf0-a5c4-ebec50a60318 · outbound

This paper cites Analytic- dpm: an analytic estimate of the optimal reverse variance in diffusion probabilistic models, 2022.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Analytic- dpm: an analytic estimate of the optimal reverse variance in diffusion probabilistic models, 2022

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.071756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.071756Z digest=sha256:87a27a8aed7c5029a3e0e5f1022c211be2dc29a51eac87d2e58539792626f6e9

Observation 1ec4130f-6976-4b60-ac31-064da350ec4b · outbound

This paper cites MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.076616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.076616Z digest=sha256:83b5d7eb01b4654474949834712849dcabfa323a2cd87174d9117e112fd16ed8

Observation c19ee446-7ef1-4332-9b85-34e2226e1571 · outbound

This paper cites HALC: Object Hallucination Reduction via Adaptive Focal-Contrast Decoding.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model HALC: Object Hallucination Reduction via Adaptive Focal-Contrast Decoding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.081816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.081816Z digest=sha256:a1dae290a84e02be4a92e70488c89dba57415edb4fabf83650ad9a53998bbd38

Observation eaf80934-d9d7-4a17-8708-d5024abaf709 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.087133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.087133Z digest=sha256:7ebf64f46ee26bd8e17d8cfa2d44eccc640fc324d1481edbe9336ae21a80fc17

Observation e00e5a9b-bcb2-493e-b511-a2ae577baac3 · outbound

This paper cites Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.092153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.092153Z digest=sha256:f50f0fe4d16a8f2fe6284911a9442989312b6d9787022c7bda48771bd1a6ab3c

Observation c819938f-f686-4df2-bff3-4272318f25c8 · outbound

This paper cites Instructblip: Towards general- purpose vision-language models with instruction tuning,.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Instructblip: Towards general- purpose vision-language models with instruction tuning,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.097084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.097084Z digest=sha256:122ccca223842330cacd6d3e4f562addd5fac0184ab9e7678728056feaa2248e

Observation 7e567638-7001-4bb0-a5a9-4c0278796200 · outbound

This paper cites Deepseek-v3 technical report, 2024.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Deepseek-v3 technical report, 2024

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.107444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.107444Z digest=sha256:b96111410aa199d58b6371a37ac27bfee6607b755c04079b164ee89aa28b7044

Observation cc5976e8-5f6c-4489-8476-096b526b951e · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Imagenet: A large-scale hierarchical image database

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.113109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.113109Z digest=sha256:e01e16346a970247ecdd2b2abe16576b08bdaab8f70b3c312c7c66396eca6d91

Observation ccaabffa-b2e6-4969-a60d-0cf3c5f7e961 · outbound

This paper cites Cogview: Mastering text-to- image generation via transformers, 2021.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Cogview: Mastering text-to- image generation via transformers, 2021

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.117969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.117969Z digest=sha256:59eb4ec95dfec2ba4b6f2de647c0f46afabb6db53f78f1663df3ca85417d0917

Observation e2ba1748-3173-4c43-9221-9c2e8c359b3c · outbound

This paper cites Dreamllm: Synergistic multimodal compre- hension and creation, 2024.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Dreamllm: Synergistic multimodal compre- hension and creation, 2024

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.122884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.122884Z digest=sha256:e461d5a926b6837efdb890711a5bfe9dc746e681945570ccda0594e952bd29cf

Observation 8187f5d8-4eae-4649-b2f8-30c8dfab353a · outbound

This paper cites Taming transformers for high-resolution image synthesis, 2021.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Taming transformers for high-resolution image synthesis, 2021

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.127322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.127322Z digest=sha256:584c3fd5a33dbe27bea1335abe5a89253f037330c61ba512124c88d301b78c51

Observation 6b7f01ba-b9f6-4e40-a94a-3baf2bb28b9f · outbound

This paper cites Fluid: Scaling Autoregressive Text-to-image Generative Models with Continuous Tokens.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Fluid: Scaling Autoregressive Text-to-image Generative Models with Continuous Tokens

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.132171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.132171Z digest=sha256:f1e0f06583c78f743aa18902d10aea304cfcbd80d0169f72b541acbb1c4b6505

Observation 4b049a85-1fee-4821-b33f-9a53edb8b170 · outbound

This paper cites Mme: A comprehen- sive evaluation benchmark for multimodal large language models, 2024.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Mme: A comprehen- sive evaluation benchmark for multimodal large language models, 2024

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.137915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.137915Z digest=sha256:54ab4a9c78d1c41cffc221d8e3a2e807e7151006eb1e7a03e02bd87266481032

Observation 7db3e256-48c4-40de-8a5e-8aa20d838f68 · outbound

This paper cites Making llama see and draw with seed tokenizer, 2023.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Making llama see and draw with seed tokenizer, 2023

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.142561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.142561Z digest=sha256:c437a49eaee410258f4e9980aa7c133505d6f4672124732a537a43368b71f7fa

Observation c7e95b52-af29-49cd-8716-f2adb530d859 · outbound

This paper cites SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.147572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.147572Z digest=sha256:cc627b1172dffea670399a97ed7919496abf856374fab10d4a4a1c0986b6e90c

Observation 395f0c5e-97e6-4b70-9bcb-f6804c9c849e · outbound

This paper cites Making the v in vqa matter: Ele- vating the role of image understanding in visual question answering, 2017.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Making the v in vqa matter: Ele- vating the role of image understanding in visual question answering, 2017

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.152825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.152825Z digest=sha256:cd358d9ed9b69c939ef43cc0030f5bcb4b98dc656964e4f9cff4cf12c0d7a42f

Observation ad7bb450-fe95-49f0-bc2c-078aa624bb27 · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.157578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.157578Z digest=sha256:4fc8fffb1973b00f8c67cbb74fd6548c503661efb06b91a9d196e84b71d582ac

Observation de348b17-ce73-4b05-a90d-647049fd6fb4 · outbound

This paper cites Stangl, Anhong Guo, Chi Lin, Kristen Grauman, Jiebo Luo, and Jeffrey P.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Stangl, Anhong Guo, Chi Lin, Kristen Grauman, Jiebo Luo, and Jeffrey P

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.162915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.162915Z digest=sha256:b01dd98a4faaaf96c37561f44d3c687a021ea515e72a0afc61b0785f9d590065

Observation fd96de1b-a8e3-458e-94ae-45c05faaf459 · outbound

This paper cites Infinity: Scaling bitwise autoregressive modeling for high-resolution image synthesis, 2024.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Infinity: Scaling bitwise autoregressive modeling for high-resolution image synthesis, 2024

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.167558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.167558Z digest=sha256:ebf068d5e1dbeea2cb7af10b34c323021ac415247e3d80f17adc897cf29bb72b

Observation 83199dc8-7144-4c07-b62f-6f31345dea6e · outbound

This paper cites Scaling Laws for Autoregressive Generative Modeling.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Scaling Laws for Autoregressive Generative Modeling

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.172479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.172479Z digest=sha256:519e76b9deda33b5e4190ce5874b728df6a8ad43d388f0d69e27e7d9ffae8f9b

Observation 6a913b5d-76b8-4ff4-9ac7-5456d24a75dd · outbound

This paper cites Denoising diffu- sion probabilistic models.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Denoising diffu- sion probabilistic models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.177268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.177268Z digest=sha256:a38e5d237fed4c49286b698bef4533eaf90a582a423f6d90f4bd46333210a9df

Observation e14c0d7f-ef16-40b3-893b-5cb64c3118df · outbound

This paper cites Denoising diffu- sion probabilistic models, 2020.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Denoising diffu- sion probabilistic models, 2020

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.181964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.181964Z digest=sha256:ea8e16ac8d7dc496c684f1c0b785a500fe2e914e9cb62eb370857913c1f0cafb

Observation c710cb36-5d00-4c49-99b7-f9ba42f719a9 · outbound

This paper cites Fleet, Mohammad Norouzi, and Tim Salimans.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Fleet, Mohammad Norouzi, and Tim Salimans

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.186599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.186599Z digest=sha256:9e01c708831ae887adee69b86f8bb9822da80766e51acb2fd5aa650c1fa63f83

Observation 05c48a21-3f0f-415d-b190-9e083acc0075 · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.191427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.191427Z digest=sha256:92fe9729a87f369b637d4af003c53bedb373654b043873146be516292ea00c0b

Observation 077c2209-9771-4993-9318-d9fca4d48144 · outbound

This paper cites Hudson and Christopher D.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Hudson and Christopher D

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.196480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.196480Z digest=sha256:c2f1c12812cef1dde38c2201ac68dd0df07155e2a514c817f38aeaf2f09ca0ff

Observation 1473683f-a911-4071-a4b0-79ac04e25f79 · outbound

This paper cites Scaling Laws for Neural Language Models.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Scaling Laws for Neural Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.201224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.201224Z digest=sha256:d83a8a50626fce7ec2670696187a189242e0d0393b977d997fba211d5f98d81c

Observation a59168d2-cd0d-417f-a551-f25d6a825290 · outbound

This paper cites Referitgame: Referring to objects in pho- tographs of natural scenes.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Referitgame: Referring to objects in pho- tographs of natural scenes

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.206167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.206167Z digest=sha256:8809ba60dafe0c5074ee8a2f091fbd93881a3e978839695a4407268c8d5b8fe7

Observation edafc5d2-ee27-42f5-a95d-57ac40f5108d · outbound

This paper cites Visual genome: Connecting language and vision using crowdsourced dense image annotations.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Visual genome: Connecting language and vision using crowdsourced dense image annotations

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:21:05.838318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:21:04.211116Z digest=sha256:8a1e0082d57c69786a496e53e20f87b6f9b746f4fd1788b99ff015c43b78942f

Observation 76c23acd-d042-4fa4-92df-e026d85caa07 · outbound

This paper cites The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:21:05.823150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:21:04.216098Z digest=sha256:1cc7f2851f207bd061ee33364a50d6e71a8d673330b8ebc7220bf563c176f6b9

Observation 6b1f51eb-97d4-4ad2-8abc-b868ad943db6 · outbound

This paper cites an unresolved cited work.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-10T17:21:05.807189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:21:04.220768Z digest=sha256:43009b98270a46a585ec9b733e01189611d4478529a2da2829700698b25ae30b

Observation a2674656-240c-47d0-8726-57db70b158b6 · outbound

This paper cites Datasets: imagenet-1k-vl-enriched.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Datasets: imagenet-1k-vl-enriched

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:21:05.792194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:21:04.225733Z digest=sha256:dbb3a41687fba685bbfaca33134db5d714aebb190426a25912ccce564018fa09

Observation b1af209a-5da4-4f7b-9133-5f254b5b1a4a · outbound

This paper cites Deep learning.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Deep learning

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:21:05.777349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:21:04.230513Z digest=sha256:36293abae50cb9281556fdc6821e97a597c629403fd0aad3b0974e4718adae17

Observation ebbe9e8b-7c63-4326-914a-90bf4f6025bd · outbound

This paper cites Autoregressive image generation using residual quantization, 2022.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Autoregressive image generation using residual quantization, 2022

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:21:05.761859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:21:04.235485Z digest=sha256:491d017d9d83d38757cc98b11907a2dd49fe434c537fbe0ab17c65ad8b5335f4

Observation c5d0fb22-0895-45e8-8094-17af05f79de6 · outbound

This paper cites Mitigating Object Hallucinations in Large Vision-Language Models through Visual Contrastive Decoding.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Mitigating Object Hallucinations in Large Vision-Language Models through Visual Contrastive Decoding

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.240341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.240341Z digest=sha256:6d43b4d1950c25f3451700eb89b3dc6063266b767bf40f7e1e99f455dd083880

Observation c690a647-abe8-4335-87f0-ebd032a52f8a · outbound

This paper cites Seed-bench: Benchmarking multimodal llms with generative comprehension, 2023.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Seed-bench: Benchmarking multimodal llms with generative comprehension, 2023

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:21:05.746758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:21:04.245178Z digest=sha256:9c011577ed562c012e938b0710878f66d057777f5746da829a80e5e06e0d40a6

Observation 5840bd26-27e1-476a-acc6-65ef61c90cda · outbound

This paper cites Llava-next: What else influences visual instruction tuning beyond data?, 2024.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Llava-next: What else influences visual instruction tuning beyond data?, 2024

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:21:05.731951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:21:04.250131Z digest=sha256:495a44903537a4327e22c0ac8c180471524f6d4958c4405dc9fbf63531b48709

Observation a75f6018-7d76-4c07-888b-d6a32297b8e0 · outbound

This paper cites Llava-next: Stronger llms supercharge multimodal capabilities in the wild, 2024.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Llava-next: Stronger llms supercharge multimodal capabilities in the wild, 2024

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:21:05.716817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:21:04.254643Z digest=sha256:55d4641a13fd751fcb97aef9d73a3cdb3db88eb4728c5f7e6866d1e879f75610

Observation 94b0743b-231f-45f1-946c-d4c93240bdb2 · outbound

This paper cites Llava-onevision: Easy visual task transfer, 2024.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Llava-onevision: Easy visual task transfer, 2024

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:21:05.702250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:21:04.259415Z digest=sha256:19c13ff9422860d98a84034460da9e6d65dcacd7ab0ca7c6cc6dcb81cfd3fc36

Observation 58bda4e4-2289-47ce-ad53-f14da0f948d4 · outbound

This paper cites Llava-next: Tackling multi-image, video, and 3d in large multimodal models,.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Llava-next: Tackling multi-image, video, and 3d in large multimodal models,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:21:05.687367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:21:04.264372Z digest=sha256:718f646b07fd36a1efac2c437a59f6298cb3b2236528afced61b1f05c40c9144

Observation d0023b5b-6601-469d-a346-fd2e97af8db9 · outbound

This paper cites Blip- 2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Blip- 2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:21:05.672425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:21:04.269031Z digest=sha256:fafc7284e053db69c5616441ea940045af64dc27517779a9f0725cd60a4f7429

Observation 53ab4899-7181-43f8-a309-a5e25a2ed6b7 · outbound

This paper cites Blip- 2: Bootstrapping language-image pre-training with frozen image encoders and large language models, 2023.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Blip- 2: Bootstrapping language-image pre-training with frozen image encoders and large language models, 2023

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:21:05.656693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:21:04.273858Z digest=sha256:566e1119e0abb80cab7a6f89e63b7704c366dfc9084ff72ca7dc1d37ce019aa6

Observation 3b873459-04b5-4fb5-b0f4-0d25d8fa23a4 · outbound

This paper cites Evaluating object hallucination in large vision- language models.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Evaluating object hallucination in large vision- language models

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:21:05.640969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:21:04.278654Z digest=sha256:ee08cde596602d17b4d5a1be91f769cc503d3bee141d18a56940e4f00f06416d

Observation 2c4a206b-008a-4440-ad1f-e8d0e1bc9402 · outbound

This paper cites Dual diffusion for unified image generation and understanding, 2024.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Dual diffusion for unified image generation and understanding, 2024

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:21:05.625246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:21:04.283580Z digest=sha256:cd744411edf7781ed3176e7d7c431a12af50b369e6f242268e11308d31c8b11f

Observation 7f25221c-d942-4d8b-919a-cd855aff610c · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Improved Baselines with Visual Instruction Tuning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.288423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.288423Z digest=sha256:c01acfb7f31c92fc9c21c9a2332b620bbd0a136769c85358fd49b2712ad77f86

Observation bfdfb097-5980-4de6-83ca-b0cf21c628b3 · outbound

This paper cites Visual Instruction Tuning.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Visual Instruction Tuning

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.293521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.293521Z digest=sha256:5e3e3bdf71948eee64331c3c95626f892f50ac1bf6e80b501ef5456b174b27d9

Observation 878d5f9d-8086-4274-92d0-89b11770d923 · outbound

This paper cites Visual instruction tuning.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Visual instruction tuning

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:21:05.609283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:21:04.298785Z digest=sha256:61b2d5493a0c7bc63cc3bedbd7273d4725f7e2192f77e4e8fedf9f5910c730a2

Observation 401fee1b-c74a-4080-9fdf-84cc84033a6c · outbound

This paper cites Improved baselines with visual instruction tuning.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Improved baselines with visual instruction tuning

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:21:05.593616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:21:04.304374Z digest=sha256:db6320aff84c7a80f64ecf826009515b87ac50325ee8c20ee1b8ced04a28158d

Observation 1e21cf32-bc91-4ed1-bd31-3ecdcce2692c · outbound

This paper cites Improved baselines with visual instruction tuning, 2024.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Improved baselines with visual instruction tuning, 2024

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:21:05.578610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:21:04.309619Z digest=sha256:541f6c4501b2cb242804baa373447a3b7bcec90a1bc1324699a9246efdf16bca

Observation 821f2ff4-a50b-4d46-86e5-e27b39403060 · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:21:05.562380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:21:04.314855Z digest=sha256:5ef070f11005099be79088cc898a3123a72d47647014764815d9d82b4042738f

Observation 367b11ed-4739-47f0-852c-cdacab181abc · outbound

This paper cites World model on million-length video and language with ringattention.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model World model on million-length video and language with ringattention

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:21:05.545290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:21:04.319618Z digest=sha256:75df9940e2899b9cf6bcc3dd373933dfd0344693e3fa513d860a61133571659a

Observation ee285e19-6127-4fdf-be4f-100c99da170e · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player?, 2024.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Mmbench: Is your multi-modal model an all-around player?, 2024

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:21:05.529584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:21:04.324310Z digest=sha256:437fbcffe335179c47ecef8b9c21756a5935aca0dccc8be5e3dd11841ac6e2b8

Observation 93a6883f-d62b-40bd-b416-7dbb7e0339c3 · outbound

This paper cites Dpm-solver: A fast ode solver for diffusion probabilistic model sampling in around 10 steps, 2022.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Dpm-solver: A fast ode solver for diffusion probabilistic model sampling in around 10 steps, 2022

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:21:05.511965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:21:04.328973Z digest=sha256:4eb30fb0ac2c2d95a36cf2c1c4e39b1494711532f5b3f236992d33cebaa4d5bd

Observation 0232c3dd-4465-4a54-a616-ba7197d81649 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering, 2022.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Learn to explain: Multimodal reasoning via thought chains for science question answering, 2022

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:21:05.495366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:21:04.333752Z digest=sha256:cd9d90ac8f9015d7edb2774f702ff9e9e177ca8ff51c0bc527481ed0dd124559

Observation 84008c93-d13a-40c1-97d6-d1ef4d9b0602 · outbound

This paper cites Unified multi-modal latent diffusion for joint subject and text conditional image generation, 2023.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Unified multi-modal latent diffusion for joint subject and text conditional image generation, 2023

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:21:05.478887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:21:04.338929Z digest=sha256:09a9ec8779b3fa0b44be05ab159bb3d42374fcb44bf854d3223180fab7b97f89

Observation 4f9e81fe-ab55-47af-8ea5-e797451337d0 · outbound

This paper cites Generation and comprehension of unambiguous object descriptions.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Generation and comprehension of unambiguous object descriptions

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.343869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.343869Z digest=sha256:8d964016328911179b5134103b61fa806bac91e3183498cadeda4815fc241fed

Observation ed88223d-20e4-4c11-bdf6-8c549e1d9350 · outbound

This paper cites Ok-vqa: A visual question answering benchmark requiring external knowledge, 2019.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Ok-vqa: A visual question answering benchmark requiring external knowledge, 2019

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:21:05.452449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:21:04.348553Z digest=sha256:41b5f3c905594b3f3fd93beaa0b7329d022d7905e2e17bf579a684601cd3de88

Observation 5f1c9815-c403-4bae-99fb-b01ac8f9e695 · outbound

This paper cites Ok-vqa: A visual question answering benchmark requiring external knowledge.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Ok-vqa: A visual question answering benchmark requiring external knowledge

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:21:05.435954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:21:04.353053Z digest=sha256:bdfdd43447d9850a1c9eb2c3a8bf3eefe79d9ebe24bc9c0053c619e6c8a9bf3f

Observation d3395c2d-6c74-44b1-aeea-bb2680354427 · outbound

This paper cites Ocr-vqa: Visual question answering by reading text in images.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Ocr-vqa: Visual question answering by reading text in images

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.357580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.357580Z digest=sha256:e7e4748cf7050ce5d9dafc64a00f9227992a6258fa64d77b1721fbbc43ab589b

Observation 2949aca0-29ff-4e41-ab65-d58979c19940 · outbound

This paper cites Improved denoising diffusion probabilistic models.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Improved denoising diffusion probabilistic models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.362267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.362267Z digest=sha256:3ab45aac4ed13d89a9f51e214267ff28b17a05726e6813862023ce30a8be6930

Observation 6adfc238-4cb5-4864-b2ab-f49da80dc63b · outbound

This paper cites Sdxl: Improving latent diffusion models for high-resolution image synthesis, 2023.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Sdxl: Improving latent diffusion models for high-resolution image synthesis, 2023

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.367146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.367146Z digest=sha256:dc5ea376c138e502ccac9da7821241f05be95c175cd049de3c5cccb59d11ff0d

Observation d4f27e65-f17b-4304-a135-edd8c59b979e · outbound

This paper cites Du, Zehuan Yuan, and Xin- glong Wu.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Du, Zehuan Yuan, and Xin- glong Wu

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:21:05.387857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:21:04.371864Z digest=sha256:0f6ab89faad823e46b83e0175fce922e1a6efea65125f7d0482ccc84cc54c69a

Observation 3cf0f681-042d-4c4c-9b7a-5ab69a9c262a · outbound

This paper cites Improving language understanding by generative pre-training.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Improving language understanding by generative pre-training

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.376519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.376519Z digest=sha256:278ca506b854cd792b89c9ebaf8248d0b136b558d77d28fdd500eaefd0e95ac5

Observation 5e957856-07e5-46f2-9882-ad0d28dbbb64 · outbound

This paper cites Learning transferable visual models from natural language supervision, 2021.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Learning transferable visual models from natural language supervision, 2021

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.381413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.381413Z digest=sha256:9878dffc2c4cc9615e771b6b2dc8ee9f6b9ab0be58ff23472e7fb4ba65e09079

Observation 5dedcf49-fcdc-44c6-a8a6-a2a7723e72b1 · outbound

This paper cites Zero-shot text-to-image generation, 2021.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Zero-shot text-to-image generation, 2021

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.386393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.386393Z digest=sha256:fa21219296a558129af7e9f489ad29e08a12e40b6bcbc2d755b1d8606b6abba2

Observation 672f3e67-c3be-4d75-a9e2-eabc8f8a6676 · outbound

This paper cites Hierarchical text-conditional image genera- tion with clip latents, 2022.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Hierarchical text-conditional image genera- tion with clip latents, 2022

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:21:05.341071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:21:04.391482Z digest=sha256:a05c0d345102f5cca42b1f2a231215005ede72fda626bd9c67537fa0159465f6

Observation 504fc21f-a55a-4ce4-8205-1b4882b92f90 · outbound

This paper cites High-resolution image synthesis with latent diffusion models, 2021.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model High-resolution image synthesis with latent diffusion models, 2021

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:21:05.324461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:21:04.396305Z digest=sha256:d4be61aec95dfc34d3e88eda53e768fd6af3090ffd55e34e0ba2f86e3f1a55da

Observation 64eed9d4-bce1-4264-95ed-7368ce130d2f · outbound

This paper cites A-okvqa: A bench- mark for visual question answering using world knowledge.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model A-okvqa: A bench- mark for visual question answering using world knowledge

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:21:05.308472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:21:04.401232Z digest=sha256:d1189560be10b24ce944ccff1ba133891ce1771e5a8034fcb4aa144cf7f81181

Observation 3b30c682-79a9-4642-8646-a0cc7866efa5 · outbound

This paper cites https://sharegpt.com/, 2023.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model https://sharegpt.com/, 2023

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:21:05.292565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:21:04.406214Z digest=sha256:f9eb802521d1765be2b9ef769e630473c6944dd1ebc82429ac7010862ed11dfd

Observation 4f535246-2cb0-49a0-b2bc-4420abf03743 · outbound

This paper cites Textcaps: a dataset for image captioning with reading comprehension.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Textcaps: a dataset for image captioning with reading comprehension

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:21:05.276429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:21:04.411103Z digest=sha256:255da053ce48afde42bf99aa2628334522787daeffcd47f9e78da3287f1cc777

Observation 78a7022f-34f2-440a-af68-a4048fe072ce · outbound

This paper cites Towards vqa models that can read, 2019.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Towards vqa models that can read, 2019

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:21:05.259803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:21:04.416070Z digest=sha256:551b5ee4dc9e26cfb44048a0a7642176d61b2a1297f982506fb8cf4befb40f79

Observation c884e091-49c7-49de-bc2f-8a41d6eb0565 · outbound

This paper cites Denois- ing diffusion implicit models, 2022.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Denois- ing diffusion implicit models, 2022

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.421214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.421214Z digest=sha256:5d5142d30ccdda083025c48ee81334cc8aeda9cc1bb6e6af52f32493c945b35d

Observation f336a288-88b0-4b96-9958-bab1ceaad257 · outbound

This paper cites Generative modeling by estimating gradients of the data distribution, 2020.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Generative modeling by estimating gradients of the data distribution, 2020

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:21:05.233418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:21:04.425900Z digest=sha256:d954de7b5d55b25e149c0393ec7376139d834dd10b583f356d3d262ac842210a

Observation f557c42d-75b3-414b-8d52-9bd3628b6b76 · outbound

This paper cites Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.430791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.430791Z digest=sha256:adf9731cb0adb200046027915732d255bd0452a1f1d8e41a5fa599a15d87d8b4

Observation 9534376e-6416-4d9e-a1db-112c70e201eb · outbound

This paper cites Autoregressive model beats diffusion: Llama for scalable image generation, 2024.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Autoregressive model beats diffusion: Llama for scalable image generation, 2024

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:21:05.218037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:21:04.435876Z digest=sha256:06139b89eddfea9bff5e8aa9352eee7390622f0acbaa746807ecdc2dd3cf6df2

Observation f1b697cc-4b73-4fd7-a086-d4ba312325a1 · outbound

This paper cites Emu: Generative pretraining in multimodality, 2024.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Emu: Generative pretraining in multimodality, 2024

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:21:05.202213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:21:04.440563Z digest=sha256:dfc1ba8cf21171036c8e81cb0631efb2097ae0186a956b5c01d1b41707ea07cf

Observation 4284d52c-fe79-427f-9813-d5d51dfcddac · outbound

This paper cites Hart: Efficient visual generation with hybrid autoregressive transformer.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Hart: Efficient visual generation with hybrid autoregressive transformer

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:21:05.185718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:21:04.444991Z digest=sha256:71d140a53c0677dbe70d3446b902f13f1665446fcea327fe0478a65d753cab33

Observation 33cd3e92-1192-4ad6-ac2f-e43e1264750d · outbound

This paper cites Any-to-any generation via composable diffusion, 2023.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Any-to-any generation via composable diffusion, 2023

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:21:05.169008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:21:04.449643Z digest=sha256:77bf9c5e45373cfadd0482417b02f2a5186b84853c5d90902afa70473b417587

Observation b7482b26-fc8c-480e-8e16-18dabf8f7f66 · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.454118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.454118Z digest=sha256:43872af94e4ffffcfd62a706e17513c4eb19be0bf06b9ebd92b8006e94f6c247

Observation b09cea6d-cd21-4e3a-b502-604c4227c340 · outbound

This paper cites Chameleon: Mixed-modal early-fusion foundation models, 2024.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Chameleon: Mixed-modal early-fusion foundation models, 2024

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:21:05.153496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:21:04.459369Z digest=sha256:d28fa6b9168019c344c1b81465ba6643401207a37f15e0366bacc93a99914e3f

Observation 0eb9c253-a750-4dfe-8ce0-e714ac4f1155 · outbound

This paper cites Gemini: A family of highly capable multi- modal models, 2024.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Gemini: A family of highly capable multi- modal models, 2024

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:21:05.136580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:21:04.464229Z digest=sha256:c92da751b88370a9fe795bee8b1ccecf93034374f12d32b430c5a9ca7428597f

Observation c46ec641-224c-4622-a994-0dbfbee052f9 · outbound

This paper cites Visual autoregressive modeling: Scalable image generation via next-scale prediction.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Visual autoregressive modeling: Scalable image generation via next-scale prediction

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:21:05.117973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:21:04.468923Z digest=sha256:55212a518cb50fe9a0612af56348fa3c2b362538576f8fe3be0f7215fe88a046

Observation afa5e832-8fdd-474a-b72b-2ffe528815ed · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model LLaMA: Open and Efficient Foundation Language Models

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.473439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.473439Z digest=sha256:5e3f090c088d574e5e82d9ba6caeaadaa8adde4ead9d2972be725f8c35e97025

Observation 575930a0-480b-4283-8db9-9c67ae203c2f · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.477802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.477802Z digest=sha256:a2ee5df6e30707672b3299804b433063412f2aa1f0d6237869d760b7f0358280

Observation 223ddf3d-9771-4eec-aaca-9a5db3eccc20 · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Emu3: Next-Token Prediction is All You Need

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.482440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.482440Z digest=sha256:e5b1b11fd9acba48175a2663290d30d3529ec354d55165533ac7402fc8d7f72e

Observation ccf4fa23-a0f7-4caf-98be-6f1a4a9cc6b0 · outbound

This paper cites Janus: Decoupling visual encoding for unified multimodal understanding and genera- tion, 2024.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Janus: Decoupling visual encoding for unified multimodal understanding and genera- tion, 2024

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:21:05.099022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:21:04.487152Z digest=sha256:b0a0cc99358c7d21adb758012186b4c532f389052063c1de804d88517f8e840a

Observation 07804346-6d60-4570-81a4-816bb00d4b59 · outbound

This paper cites Liquid: Language Models are Scalable and Unified Multi-modal Generators.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Liquid: Language Models are Scalable and Unified Multi-modal Generators

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.491869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.491869Z digest=sha256:ad4a70d137831574d00c20df28da9e1d4345d843fdd26d897128ba489566723d

Observation 2919fe17-42dd-4e1d-a91c-f17cf84ad08b · outbound

This paper cites NExT-GPT: Any-to-Any Multimodal LLM.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model NExT-GPT: Any-to-Any Multimodal LLM

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.497553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.497553Z digest=sha256:599d7d1d87fd132c2f92ba62b3be4368f1103e6806a5bf6401cf57a68d31079b

Observation 1b285d06-f445-495c-965d-fcd465fe69d6 · outbound

This paper cites Show-o: One Single Transformer to Unify Multimodal Understanding and Generation.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.503740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.503740Z digest=sha256:a84c42ba016cb4cdd1f0a6fb046c08f9b83fe7aa5eaceb4ac0683d46c190df90

Observation 4acae73c-0414-4deb-bcfe-5f87fece7fba · outbound

This paper cites Show-o: One single transformer to unify multimodal understanding and generation, 2024.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Show-o: One single transformer to unify multimodal understanding and generation, 2024

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:21:05.081298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:21:04.508878Z digest=sha256:a98063972a019b667132e416eb22d2004a43b680663b5cd19ddf597ebbb4e15e

Observation 3192d077-b9ca-406b-99f0-7d193a3837d9 · outbound

This paper cites X-vila: Cross-modality align- ment for large language model, 2024.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model X-vila: Cross-modality align- ment for large language model, 2024

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:21:05.065607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:21:04.513698Z digest=sha256:a1ed882167cddd75ec754a9ccf42e56c93feec0fa8b113eff1591f3cd332983d

Observation 7f74507e-3009-4f7d-be6e-dc1309ffbd58 · outbound

This paper cites mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.518425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.518425Z digest=sha256:7c2181971486e199ed6e540cd672058c9d862c8e6de38a3710ed8f382ee7f0c4

Observation 2c9753cf-45c7-4c39-bbb5-5c5c724de964 · outbound

This paper cites mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:21:05.048940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:21:04.523483Z digest=sha256:f6b0df0a75767c020e86e714f62fae546ef35a515431e13a7ebde60e62f71e45

Observation a86774a0-f579-4365-b430-b82d493d0519 · outbound

This paper cites Woodpecker: Hallucination Correction for Multimodal Large Language Models.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Woodpecker: Hallucination Correction for Multimodal Large Language Models

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.528220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.528220Z digest=sha256:b1eecbdcc247d321d326bcabd941e9ea7f03b5edc25ada1216b1294cfb802bd0

Observation b0a5f092-939f-4f8f-904d-ecfc93be4057 · outbound

This paper cites Scaling autoregressive models for content-rich text-to-image generation, 2022.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Scaling autoregressive models for content-rich text-to-image generation, 2022

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.533079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.533079Z digest=sha256:1745dc680a388e0e515f2fd7ccb1ec9d83e923605acd93624554da3be0e56959

Observation 8e66f95f-5fb4-4782-b7b8-9f09cc37c952 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understand- ing and reasoning benchmark for expert agi, 2024.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Mmmu: A massive multi-discipline multimodal understand- ing and reasoning benchmark for expert agi, 2024

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:21:05.016487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:21:04.537979Z digest=sha256:6d7770c2760f148f85b3695b17fd8ff27d2bb3997d7646b391dde3b0a93fcd93

Observation 86b1424d-1762-47fc-8a10-54433d2ceff1 · outbound

This paper cites Lmms- eval: Reality check on the evaluation of large multimodal models, 2024.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Lmms- eval: Reality check on the evaluation of large multimodal models, 2024

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.542716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.542716Z digest=sha256:ae79d54c6344b4652a8ab589d9e82a1c7ddaa66e202fef5a89bbdc3550dcb66e

Observation 016c92f5-1ba3-4d55-b361-2ff766f765d6 · outbound

This paper cites Var-clip: Text-to-image generator with visual auto-regressive modeling, 2024.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Var-clip: Text-to-image generator with visual auto-regressive modeling, 2024

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:21:04.983358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T17:21:04.547305Z digest=sha256:21295fac0b42cace4d4d373db2c0304a6053fb88e80b97065891b924f39e8696

Pith citing papers

Observation ce6f86c3-8ae3-464c-859c-31cbe8c6bfc5 · inbound

A Survey on Vision-Language-Action Models for Embodied AI cites this paper.

A Survey on Vision-Language-Action Models for Embodied AI VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model

Reference 141

Resolution
verified exact
arxiv_id, observed 2026-05-24T01:25:54.412504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-24T01:25:10.150459Z digest=sha256:cf133b418a28a63d54767f332ad833a844ce79b87760bbb94b4468f93ac7a2ea

Observation eeeb5c56-95d2-4e74-ba1a-1502d5f79c9d · inbound

Do we really have to filter out random noise in pre-training data for language models? cites this paper.

Do we really have to filter out random noise in pre-training data for language models? VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T15:04:29.442501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:04:29.442501Z digest=sha256:b42db0b359af8c56f8b6b862fbb43ba102edc8b764f0773695953dc36dfed63d

Observation 18734500-4d57-4696-84bd-10f7449a3a0c · inbound

Generative Multimodal Pretraining with Discrete Diffusion Timestep Tokens cites this paper.

Generative Multimodal Pretraining with Discrete Diffusion Timestep Tokens VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-16T11:48:03.500362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:48:03.500362Z digest=sha256:87c7c6dcb81f955690ab78c553fe7c7b8447a8b16aa493e663be9ef5e73abad4

Observation 3678c49c-a694-4464-8743-21f60760af38 · inbound

Open Set Domain Adaptation with Vision-language models via Gradient-aware Separation cites this paper.

Open Set Domain Adaptation with Vision-language models via Gradient-aware Separation VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:59.260983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:59.260983Z digest=sha256:7d2d5a257080753950729ab9a496f4eec90822fd84c2e3618974ec9d34ad64d4

Observation e6b68b22-0eb4-4371-9e7a-5af03bd4b398 · inbound

MMaDA: Multimodal Large Diffusion Language Models cites this paper.

MMaDA: Multimodal Large Diffusion Language Models VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:50:59.753839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:c96dca4fae1c1710be709eb6cada0b2d2098750c080670c4b7dfc8433fb7e434

Observation 20f41297-a7fb-4c90-beca-1ca735b99dca · inbound

Draw ALL Your Imagine: A Holistic Benchmark and Agent Framework for Complex Instruction-based Image Generation cites this paper.

Draw ALL Your Imagine: A Holistic Benchmark and Agent Framework for Complex Instruction-based Image Generation VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:18:35.569184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:18:35.569184Z digest=sha256:324cbf234b4e07dbb6ed53789e0177d99bdd8b2dd8740ef0fd0a8fdee49c4802

Observation 712f8acc-f2d9-4f5e-a1d2-08fc799d6f78 · inbound

Not All Tokens and Heads Are Equally Important: Dual-Level Attention Intervention for Hallucination Mitigation cites this paper.

Not All Tokens and Heads Are Equally Important: Dual-Level Attention Intervention for Hallucination Mitigation VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:26.930717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:53:26.930717Z digest=sha256:ce8c290eb9855ca9ab054f97674df115819d57d8aac6473f192db1eeec392b86

Observation 84a7f1c1-7cc2-4716-a4de-cae359319d48 · inbound

ComplexBench-Edit: Benchmarking Complex Instruction-Driven Image Editing via Compositional Dependencies cites this paper.

ComplexBench-Edit: Benchmarking Complex Instruction-Driven Image Editing via Compositional Dependencies VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T00:40:22.561926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:40:22.561926Z digest=sha256:6b142c86d82af775ec1d9d2ea1f0c2f44001582b47647537a8d629727991e6f3

Observation cd3cdd43-3777-40cd-aaf1-1e35af05d4e1 · inbound

Instella-T2I: Pushing the Limits of 1D Discrete Latent Space Image Generation cites this paper.

Instella-T2I: Pushing the Limits of 1D Discrete Latent Space Image Generation VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T22:43:04.966173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:43:04.966173Z digest=sha256:1fc93b9c2552639139ae00a023baa4ea8269e70626481a3cecece5fffcdcc020

Observation 226e02d8-a9c5-477b-b69c-0990c34ccad1 · inbound

From Static Inference to Dynamic Interaction: A Survey of Streaming Large Language Models cites this paper.

From Static Inference to Dynamic Interaction: A Survey of Streaming Large Language Models VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-15T16:20:09.981818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T16:16:15.819622Z digest=sha256:7554cfc845a0ac415a791dd264b9a3629cf669dbe56717a6ad1be38fd0d5eb54

Observation 07a519a3-d3da-46d2-b5f8-00102b18b19a · inbound

Demystifying Video Reasoning cites this paper.

Demystifying Video Reasoning VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model

Reference 80

Resolution
unresolved
no resolver link, observed 2026-07-13T23:27:11.006580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:27:11.006580Z digest=sha256:e94cc0eb5188e5d1c9e23a0c3d127d375474b814171d40870421217432e402ac

Observation e988bc26-b644-4d96-aa89-0e5d9a344453 · inbound

Demystifying Video Reasoning cites this paper.

Demystifying Video Reasoning VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-03T02:34:01.603385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:34:01.603385Z digest=sha256:2b5b2323cb195053f7470c055273a6adbcae3c575057949853f3628b88412618

Observation bf8003dd-c19c-4a18-adff-8a871acc3f23 · inbound

WinTok: A Win-Win Hybrid Tokenizer via Decomposing Visual Understanding and Generation with Transferable Tokens cites this paper.

WinTok: A Win-Win Hybrid Tokenizer via Decomposing Visual Understanding and Generation with Transferable Tokens VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model

Reference 112

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T12:08:15.816753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T12:04:19.761430Z digest=sha256:4144dbbcd6fe86b516705cdc3373ff5cc8744722b3b16c2c7483d7fbbf71a269

Observation 996a1a94-c94d-40c6-8ce7-64ed0331975c · inbound

Semantic Generative Tuning for Unified Multimodal Models cites this paper.

Semantic Generative Tuning for Unified Multimodal Models VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model

Reference 88

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T11:33:14.285652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T11:32:24.007847Z digest=sha256:02fb434cecdbbb094ad865caaa9ba4b875186660d95f4e6571ce03a91a01d31c

Observation 061a9d94-8747-45e3-a2a1-d89cc9b3bf21 · inbound

Semantic Generative Tuning for Unified Multimodal Models cites this paper.

Semantic Generative Tuning for Unified Multimodal Models VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model

Reference 88

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T18:35:00.311433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:1ce27c530b7a415ba592693338d2e6da51339bf0f9f79bc8bff496a02521603b

Observation 39ee1ca7-63ea-49ae-acd4-15ccf867488d · inbound

HACK++: Towards More Effective Head-Aware Key-Value Compression for Efficient Visual Autoregressive Modeling cites this paper.

HACK++: Towards More Effective Head-Aware Key-Value Compression for Efficient Visual Autoregressive Modeling VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:27:24.385865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T19:46:43.514413Z digest=sha256:57dc960fc0c74beb0057c295f7789c994ad41bfc5a34f5c7bf05634e71342c10

Observation 8263a698-25c9-4f07-839f-f4a8d6190a4d · inbound

Ask, Solve, Generate: Self-Evolving Unified Multimodal Understanding and Generation via Self-Consistency Rewards cites this paper.

Ask, Solve, Generate: Self-Evolving Unified Multimodal Understanding and Generation via Self-Consistency Rewards VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-07-04T13:39:51.492070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-26T04:58:15.891214Z digest=sha256:293f94252aa29602309c4ac1bea1dff40ff369471b158e9150c068b7529385ca

Observation e4e52c5d-0910-439e-b9d3-a27e11303402 · inbound

MEPA: Multi-Scale Representation Alignment for Visual Autoregressive Modeling with Mixture of Experts cites this paper.

MEPA: Multi-Scale Representation Alignment for Visual Autoregressive Modeling with Mixture of Experts VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model

Reference 65

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T15:17:07.191767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-02T15:14:36.946247Z digest=sha256:25cf19ae2b1b5dcea662c336854ddfe4b0d56f2003ea0fec82383d63bad52ced

Observation 158619c6-8924-48c6-82e8-3e8a88ce896a · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model

Reference 268

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T00:04:22.543433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-07T23:59:38.702609Z digest=sha256:12771c9dfeabcfee4cb651b803c81ac71a06eed73ae762d475ffd52f463c1bc3

Observation c4f225b9-d26d-4fc4-bedc-b1cc4b7b8c79 · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model

Reference 268

Resolution
unresolved
no resolver link, observed 2026-07-11T07:46:49.059192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T07:46:49.059192Z digest=sha256:880701396089cc497df169741fe6158bb9cebcad9788e9ccadf95e72c883d6d1

Observation 2483e1c5-8177-4b4e-be7f-ee2e17676c4b · inbound

UniGen-AR: Unifying Visual Generation with Auto-Regressive Modeling cites this paper.

UniGen-AR: Unifying Visual Generation with Auto-Regressive Modeling VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model

Reference 96

Resolution
unresolved
no resolver link, observed 2026-07-31T22:25:12.526450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T22:25:12.526450Z digest=sha256:afea645eaef88842d0686c42798f77d9260564044c03fa0b65a4379a7870cdc7

Observation e391844b-f528-4ed1-bb93-78e823d1782e · inbound

SynVAR: Synergizing Spatial and Semantic Alignment in Visual Autoregressive Model cites this paper.

SynVAR: Synergizing Spatial and Semantic Alignment in Visual Autoregressive Model VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T00:44:46.078357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:44:46.078357Z digest=sha256:9450e1bb2cd6f90746a42a88085ff5b24e393534d5483bc6d086576c53168623