Pith. sign in

Paper Citation Record · LEDGER

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation

As of 18 August 2026, this Paper Citation Record lists 100 of 103 outbound references and 17 inbound Pith citation observations for arXiv:2502.05178.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.05178 v1

Coverage vector

measured 100 of 103 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T20:04:02.460265Z

measured 117 of 117 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:09:11.045334Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T20:28:55.819781Z

Reference resolution

100 of 103 outbound references displayed

  • verified exact0
  • verified fuzzy43
  • unresolved57
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8bd6c170-fe81-48bd-a650-84f219af87a7 · outbound

This paper cites GPT-4 Technical Report.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.010635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.010635Z digest=sha256:814953ff06c3b78a52f31a6372f8a94ab96337a8e9b1293775738855f5f3869c

Observation 6a4c86db-f4e0-42f8-8f55-699fa5ff417e · outbound

This paper cites Soft-to-hard vector quantization for end-to-end learn- ing compressible representations.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Soft-to-hard vector quantization for end-to-end learn- ing compressible representations

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.016595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.016595Z digest=sha256:394773d18a7779986f452f6c13f754a15252f0685edc8addc9b7d4640e889148

Observation 70005a9b-b630-48a4-b360-ed96d1e4ba27 · outbound

This paper cites Beit: Bert pre-training of image transformers.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Beit: Bert pre-training of image transformers

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.021553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.021553Z digest=sha256:e780acdc4d0112e73f0542553fd762bda2b455ff642e4d5277da67f03a7abfe3

Observation a550fff0-826a-4228-8e97-170b988d81e5 · outbound

This paper cites Fuyu-8b: A multimodal architecture for ai agents, 2023.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Fuyu-8b: A multimodal architecture for ai agents, 2023

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.025789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.025789Z digest=sha256:d835a32d8e20893f9d153028f63e4946583674f94f54160aeb698349826950df

Observation fd6f13ba-d675-417d-8964-02ab70d2b707 · outbound

This paper cites Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.030015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.030015Z digest=sha256:a585ac2182d53ab53fdd492cab037ec79f866fe1f48208bda60f7e6075aef992

Observation 7d77e1e7-729b-43e5-a76e-dca36ec3a450 · outbound

This paper cites PaliGemma: A versatile 3B VLM for transfer.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation PaliGemma: A versatile 3B VLM for transfer

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.035083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.035083Z digest=sha256:a37a99bed299ac93c19245807bd9c4b7707047e80833fbcda968dd22e353bccd

Observation b1754986-a731-4cee-8330-f1097faf8ba9 · outbound

This paper cites Piqa: Reasoning about physical commonsense in nat- ural language.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Piqa: Reasoning about physical commonsense in nat- ural language

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.040713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.040713Z digest=sha256:af93f78cbef4a44c5009acd022a659725cdff98d1285030eb046cf8da8937f53

Observation 454e073c-559f-424a-9afa-6a2da8bf9107 · outbound

This paper cites Maskgit: Masked generative image transformer.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Maskgit: Masked generative image transformer

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.045345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.045345Z digest=sha256:7258e0608f8946229a5a4ca7ba8b659481956b955dcefc2d89181b1b89612977

Observation 2d31b2e6-2585-473e-bac5-f793d67e00a4 · outbound

This paper cites Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.050027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.050027Z digest=sha256:d9cfaea500a04c2d9afa11a3a05e76b3956752cdec72b5a79b973faac36e4ed9

Observation cf430773-b946-4491-9606-1a949a34a93e · outbound

This paper cites Training Deep Nets with Sublinear Memory Cost.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Training Deep Nets with Sublinear Memory Cost

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.054853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.054853Z digest=sha256:8e06dd012de8f7e709793b8d484fc7a7e3d8c39945ce2eea55d6d160d579e048

Observation 2e71b303-585c-422d-aefa-e640109b3891 · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.059660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.059660Z digest=sha256:dbf37c49b19cd96a59b27a5ffe4c8122169fdfc1d42e63da250f11faa4f55e44

Observation 9d0ea132-5568-4f63-817c-4766f2b741b3 · outbound

This paper cites Gradnorm: Gradient normalization for adaptive loss balancing in deep multitask networks.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Gradnorm: Gradient normalization for adaptive loss balancing in deep multitask networks

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.063916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.063916Z digest=sha256:0106c910a9f0ff99559449c4c411fba0eb8ac880867c078e29f0ca1e2ff66a5d

Observation a05f44f9-6282-4efa-863f-973f8f67adcc · outbound

This paper cites Reproducible scaling laws for contrastive language-image learning.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Reproducible scaling laws for contrastive language-image learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.067982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.067982Z digest=sha256:53549a8afc935dc6e6dab69e224ad92a6b2ad491c2dc87f4918b6047d4c7fff1

Observation f59c9454-ea65-42bb-a951-49cd6d2377b2 · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Gonzalez, Ion Stoica, and Eric P

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.072545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.072545Z digest=sha256:6256a35cbc8c0f06d3dcbba5b724655556f29abbb84eafe1cca689f794765350

Observation 4298893b-9776-4bd9-9e09-8931c84393dc · outbound

This paper cites Scaling instruction- finetuned language models.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Scaling instruction- finetuned language models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.076761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.076761Z digest=sha256:ead760e21d3df2a93c7e5f2967d4c754d811d2087208f1d16ccfc2172a0d0f28

Observation 303ad003-cea0-494f-ba3e-2810c68f09b8 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.081513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.081513Z digest=sha256:ef7b24b23f2f8c3cd73d371bb64b4bb79020912c3271333278941b697ce57c41

Observation ab67c416-a7ab-4037-99ad-b0399c4fac87 · outbound

This paper cites Flashattention: Fast and memory-efficient exact attention with io-awareness.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Flashattention: Fast and memory-efficient exact attention with io-awareness

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.087307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.087307Z digest=sha256:e4089a28d3499874db4639d062908ce08ab67be8abf544d04785b7159ccac418

Observation 68008d26-52c3-4387-8c0d-5669bcfcc032 · outbound

This paper cites Vision transformers need registers.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Vision transformers need registers

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.091863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.091863Z digest=sha256:c21154784e49de7a6917a6da0c568269e0924f0ac23c3afdb362f658cca4e526

Observation 474d4554-538f-4321-886f-968502382f03 · outbound

This paper cites Scaling vision transformers to 22 billion pa- rameters.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Scaling vision transformers to 22 billion pa- rameters

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.097385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.097385Z digest=sha256:01754bb19f392b0b56de926fe9e7d546c121fbfb7d14fec1607db376c68061e1

Observation cf644a5d-2fe0-4b4f-abb8-bf8c8b01521f · outbound

This paper cites Imagenet: A large-scale hierarchical im- age database.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Imagenet: A large-scale hierarchical im- age database

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.101775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.101775Z digest=sha256:41aae5f143cefee5088261c8c2fb24942576908846690aa4ad1298786c52639c

Observation aebfa514-5cbb-4c2f-99a4-f716ca22b36a · outbound

This paper cites Unveiling encoder-free vision-language models.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Unveiling encoder-free vision-language models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.105628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.105628Z digest=sha256:a1e219e0c03fe8d0f8f9097bf421de3859a2987fbfff2301d70204d6c0ab80f5

Observation fefafe59-b9a5-4b4a-995f-44bb4b20097e · outbound

This paper cites The Llama 3 Herd of Models.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation The Llama 3 Herd of Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.109316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.109316Z digest=sha256:9bf49e759023ed723f5eef204b90faa3595ae199ae69ade386b5a091fd29d9dd

Observation e96fd0f2-3d88-4bb5-8d7f-860cde74123a · outbound

This paper cites Scal- able pre-training of large autoregressive image models.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Scal- able pre-training of large autoregressive image models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.113967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.113967Z digest=sha256:79d826f597a16da47cd6c600ba74eb05cefaed72aac6c7063fd633bad39e76ce

Observation d5922cbe-cd1d-4439-8f5c-3877db231ce5 · outbound

This paper cites Taming transformers for high-resolution image synthesis.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Taming transformers for high-resolution image synthesis

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.118734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.118734Z digest=sha256:200ce9c34fad71d9fa92d93df8e416df9749970825938ab1c8c089d01fe2972e

Observation e2648a11-c0fd-4ed2-9ba1-92e466db8701 · outbound

This paper cites Eva-02: A visual representa- tion for neon genesis.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Eva-02: A visual representa- tion for neon genesis

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.123315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.123315Z digest=sha256:1dbcfb3f860c6e3423858a327c348afae15b3284bcd96a45ecbd279d189a478a

Observation 8aae7597-f829-435a-8876-3f2c0b6a72e1 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.128254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.128254Z digest=sha256:dca72be19a9ef66d6bce10cad138605e57d147f19c86d56a7556cee8cbddc46b

Observation 68643c2f-f220-4327-9c53-59953585761a · outbound

This paper cites Dat- acomp: In search of the next generation of multimodal datasets.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Dat- acomp: In search of the next generation of multimodal datasets

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.132772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.132772Z digest=sha256:a90f2279f56040614d838e33c17b523b32029324b15fb993d8fe719883b901f6

Observation ac7db086-f8f4-453a-a351-6da544f5c8b1 · outbound

This paper cites SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.136804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.136804Z digest=sha256:c564eb6f150c047f14a2c706ded80ef3701852f4621061934c19158e2b25fcd4

Observation c849945b-773c-451c-b9e5-9512d611fba2 · outbound

This paper cites Geneval: An object-focused framework for evaluating text- to-image alignment.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Geneval: An object-focused framework for evaluating text- to-image alignment

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.141742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.141742Z digest=sha256:185f0b1510966591c3aaefc5202231e606db486249d03496b0b67c92b190188a

Observation b1d319c2-3025-4ddf-a27e-47606c75b218 · outbound

This paper cites Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.146184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.146184Z digest=sha256:05137fe8865c95a075db7e3229fdfe8017ef8be58568dfb8a31e4ff2fff092d1

Observation c8701ceb-b64e-4d31-9c9a-1f3dea395f7e · outbound

This paper cites Making the v in vqa matter: El- evating the role of image understanding in visual question answering.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Making the v in vqa matter: El- evating the role of image understanding in visual question answering

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.150055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.150055Z digest=sha256:44c9c8b1bdf6b110e42b3ba2d776398f7e2f18472bdc8597fc46c2df20e5c348

Observation d270c97a-4c6c-4981-98bd-4a958a0aa107 · outbound

This paper cites Noise-contrastive estimation: A new estimation principle for unnormalized statistical models.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Noise-contrastive estimation: A new estimation principle for unnormalized statistical models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.153789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.153789Z digest=sha256:0b17dc168ecde8c5f85f0f336a2f1b631712b698f93c9b6533af6f6add41a2fb

Observation f9794115-36d8-4401-896c-bc4332b5b033 · outbound

This paper cites Masked autoencoders are scal- able vision learners.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Masked autoencoders are scal- able vision learners

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.157535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.157535Z digest=sha256:bca97d8568af4dfd9723d5e952e023539687c4143142da6b8d80bacd45048ee7

Observation 69f1c8aa-6680-435a-b8b6-7313cc4c3e6c · outbound

This paper cites CLIPScore: A Reference-free Evaluation Metric for Image Captioning.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation CLIPScore: A Reference-free Evaluation Metric for Image Captioning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.162100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.162100Z digest=sha256:822f41e5ce2bc28e53a25c3b62b853377da6e8042c10f99bb4dc2886aecec837

Observation 2d757e65-fedb-44df-bad2-8b04f82e4657 · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equi- librium.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Gans trained by a two time-scale update rule converge to a local nash equi- librium

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:04:03.630265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-08T20:04:02.166745Z digest=sha256:fd280769e98f5a255859e2faf3c58c51d9ab9e59b5d0086662a21b796b71b4de

Observation 734b2f13-c2ce-4f58-b9ff-f29389b70955 · outbound

This paper cites ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.171586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.171586Z digest=sha256:4377f059def8d5eb7f3b929cc29f397db07a0af841095360b03d13a40d50a080

Observation 3492adc1-c3f7-4d19-9c2f-c4d58fa7e5fd · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:04:03.616127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-08T20:04:02.176426Z digest=sha256:1737ab4b210f841ddc233233c9cc46bd28ac3da692bdeceaac0049d69d61605a

Observation 5f215137-e47e-4192-8f11-6eec88dd893b · outbound

This paper cites Coincidence, categorization, and consolidation: Learning to recognize sounds with minimal supervision.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Coincidence, categorization, and consolidation: Learning to recognize sounds with minimal supervision

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:04:03.602250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-08T20:04:02.180404Z digest=sha256:92831215ea3215ecb7fdd182143868efc736a075adc3e554136cb4e706d6a60b

Observation 14eb0cb4-5d4f-4b17-94a7-bc8f0199e4a7 · outbound

This paper cites Unified language-vision pretraining with dy- namic discrete visual tokenization.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Unified language-vision pretraining with dy- namic discrete visual tokenization

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:04:03.587534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-08T20:04:02.185578Z digest=sha256:c24423f2e50be5a81ef07d45bfd24eca591f762c21858811e16ac3dce547d2c5

Observation 79a1eb1c-072e-424c-90bb-900fcc3bbcb1 · outbound

This paper cites Analyzing and improving the image quality of stylegan.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Analyzing and improving the image quality of stylegan

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:04:03.573043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-08T20:04:02.190018Z digest=sha256:92ff7610ed88b79075e15ec93d78f0052da6fd97f5614332877767bad96a9507

Observation d6e18f26-832d-464c-9ee6-9e7bbca7cb06 · outbound

This paper cites Segment anything.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Segment anything

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:04:03.558335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-08T20:04:02.194325Z digest=sha256:3ff552bc8b8329be391e1bc96bb312a459654d1ac6a6a33d4aae7ec763a2ded8

Observation d3bd40fc-aa3a-4bb7-9e9d-1789d70b079d · outbound

This paper cites Sentencepiece: A sim- ple and language independent subword tokenizer and deto- kenizer for neural text processing.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Sentencepiece: A sim- ple and language independent subword tokenizer and deto- kenizer for neural text processing

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:04:03.540802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-08T20:04:02.198689Z digest=sha256:fa419de3980b5f72c1ff7eb8046754996e32529dfa4bcaa6d0da665e76457929

Observation a4077dbf-4bc7-4ede-b7ec-9613a7a755ea · outbound

This paper cites What matters when building vision-language models?.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation What matters when building vision-language models?

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.202500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.202500Z digest=sha256:8b152c48cc972a42b2610554f3bf577c901beff4adf67df3df34966401459ebb

Observation 2214d419-98f5-4334-8895-303afed07f3a · outbound

This paper cites Autoregressive image generation using residual quantization.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Autoregressive image generation using residual quantization

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:04:03.524016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-08T20:04:02.206911Z digest=sha256:334c338f224b894ceb91be9ffa88e26092a8552b633dd34df359cbc9127ea3b0

Observation 4a9100d2-60fc-4203-a762-1182eb0bc4dc · outbound

This paper cites DataComp-LM: In search of the next generation of training sets for language models.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation DataComp-LM: In search of the next generation of training sets for language models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.211697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.211697Z digest=sha256:733524f500f5761fcde485b30ffb837aa8616ad50eead5023fb6a7ad0dedb3fe

Observation 8eb81f88-1952-4d37-8eb9-4025db60a17f · outbound

This paper cites Mage: Masked generative encoder to unify representation learning and image synthe- sis.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Mage: Masked generative encoder to unify representation learning and image synthe- sis

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:04:03.509256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-08T20:04:02.216578Z digest=sha256:9af9d3faabfeceeaadbaea699f255bfca1cc781dec947d37e24f496527598cdf

Observation 20778747-fe02-45bb-9223-4d91232e0709 · outbound

This paper cites Evaluating object hallucination in large vision-language models.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Evaluating object hallucination in large vision-language models

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:04:03.494385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-08T20:04:02.220361Z digest=sha256:72eae23a3a7e09f29757e22a3bc7023161b27efd7ebdd89c49dcb5edcf86a808

Observation 189cdc51-2d4a-4c52-9488-4c49d257c126 · outbound

This paper cites Microsoft coco: Common objects in context.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Microsoft coco: Common objects in context

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:04:03.480719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-08T20:04:02.224653Z digest=sha256:827e010fef5b1eff3d64e616872e58778cf5885ecf9e30f5d9400aa2e9f7d55d

Observation 0952fad6-6b39-4b4f-9d53-f48ea7c7fc11 · outbound

This paper cites Visual instruction tuning.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Visual instruction tuning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.228660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.228660Z digest=sha256:ee6e0fc98ea4525aec91c4b89aad7200d20bd361aa45b17421198f2aac1d131c

Observation db0fce50-4a05-4b4c-bb37-dfcabc1c592f · outbound

This paper cites Language quantized autoencoders: Towards unsupervised text-image alignment.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Language quantized autoencoders: Towards unsupervised text-image alignment

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:04:03.457846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-08T20:04:02.232849Z digest=sha256:3af3acd77ca8ac635101c2b26d33fa9db03334347c6232dfbed743ce36a0a8ae

Observation 5595dd76-13a6-46ac-8d75-33a5649c7c66 · outbound

This paper cites Improved baselines with visual instruction tuning.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Improved baselines with visual instruction tuning

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:04:03.442955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-08T20:04:02.237764Z digest=sha256:4b0e81ad68a2a135cb50d1411a0684327890d164a34c1f2b5855ee37804856ee

Observation ab91549c-e08a-40e2-984e-525243af81ec · outbound

This paper cites Unified-io: A uni- fied model for vision, language, and multi-modal tasks.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Unified-io: A uni- fied model for vision, language, and multi-modal tasks

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:04:03.427785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-08T20:04:02.242074Z digest=sha256:f2e892c629ecbd215a0329790368703bb6f5e01c2c1d74b4b634f7d6c58d6691

Observation fec7a35a-afad-492b-8035-5bd024506d7b · outbound

This paper cites Unified-io 2: Scaling autoregressive mul- timodal models with vision language audio and action.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Unified-io 2: Scaling autoregressive mul- timodal models with vision language audio and action

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:04:03.414279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-08T20:04:02.246605Z digest=sha256:0ae744afb0407f60783fd05ed181484ec53225ebc5bcbf3058bd487b98b44c1f

Observation 7e77d0d2-b0b7-446e-9b91-8b20da7b51cb · outbound

This paper cites Open-MAGVIT2: An Open-Source Project Toward Democratizing Auto-regressive Visual Generation.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Open-MAGVIT2: An Open-Source Project Toward Democratizing Auto-regressive Visual Generation

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.250674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.250674Z digest=sha256:074126187a1b0e6552d53b1487c4cb43cff7c7ea89d8cdb762fa3f9949cd258b

Observation 252c49b0-6672-4642-99e3-a5b1ec6bec32 · outbound

This paper cites Finite scalar quantization: Vq-vae made simple.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Finite scalar quantization: Vq-vae made simple

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:04:03.400187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-08T20:04:02.255181Z digest=sha256:66f26545fae6c348e232cfdf6314ce401d66c43c6dfb8eb25fd112fe74827172

Observation 4a13c14d-5261-4a09-a6af-fd51658da579 · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Representation Learning with Contrastive Predictive Coding

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.260427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.260427Z digest=sha256:c52807fe4fc8bf493454722fd1bb9c857e84fa4aa559541a666b0ff932c82690

Observation 957365a5-9e76-4c88-85eb-300f411f3848 · outbound

This paper cites BEiT v2: Masked Image Modeling with Vector-Quantized Visual Tokenizers.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation BEiT v2: Masked Image Modeling with Vector-Quantized Visual Tokenizers

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.265135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.265135Z digest=sha256:39807192b3d9dc59fb4dcba6af0c0d182c154bfc65fc11917cc533192e162b51

Observation 579222c2-f70e-45c7-b5fe-90f53bb5ea27 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.270583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.270583Z digest=sha256:1fcfed33f813906e588fa7cc0f411b89be35124ba2bc7e8cd25fc6df2109310f

Observation ef1bc0ab-2f53-4a0c-8a65-83d6094d99fc · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Learn- ing transferable visual models from natural language super- vision

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:04:03.386055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-08T20:04:02.274919Z digest=sha256:ea21b1b0bad539128fe7e9a023d30fb8cd792fd6a6cafaad5d9a8c7fe51e6c46

Observation c9d1c10e-48f9-4711-9b7f-197257a99594 · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Exploring the limits of transfer learning with a unified text-to-text transformer

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:04:03.371521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-08T20:04:02.278871Z digest=sha256:575415fdc967be6e46ce941eb089ac0aff04ad834ca31033c395f43b026c80b6

Observation 70c991cc-d2ef-4978-a734-efc4e8326a6f · outbound

This paper cites Zero: Memory optimizations toward training trillion parameter models.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Zero: Memory optimizations toward training trillion parameter models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.282886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.282886Z digest=sha256:f50ad80a4732a485ce55cd0e2f761e15e60a755ab4c042f17e9878fe0019d435

Observation cf5a9212-2786-4c1b-92f0-4d52f8935b82 · outbound

This paper cites Zero-shot text-to-image generation.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Zero-shot text-to-image generation

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:04:03.347768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-08T20:04:02.287680Z digest=sha256:8901a59bbe73fb9a0cbac8b76058568f0e8657d01ca86805e917ab0d5473a6b7

Observation d7a5a524-7b93-477f-a79e-5bf44a845d72 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation High-resolution image synthesis with latent diffusion models

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:04:03.332837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-08T20:04:02.292062Z digest=sha256:47a03a4bc6d65546721de8308fc428674a25afef75c851c93665dd1f1b5b8c35

Observation d8bea48c-add4-462d-a853-fe13f497a8bb · outbound

This paper cites Winogrande: An adversarial winograd schema challenge at scale.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Winogrande: An adversarial winograd schema challenge at scale

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.297126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.297126Z digest=sha256:4208404b7820180c47975ad1021f8125c36a9cac47c0aa688d2f3695aa362325

Observation 8f40caa5-dcec-44ef-96dd-85f5620d0a7d · outbound

This paper cites SocialIQA: Commonsense reasoning about social interactions.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation SocialIQA: Commonsense reasoning about social interactions

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:04:03.310139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-08T20:04:02.302005Z digest=sha256:ad25f82895b304090791cb3f2737d20f5c6362d3acba4a7d3abc8be0d0ccb535

Observation 3ac211ea-e85a-4996-bc70-9f34a5f8aedd · outbound

This paper cites SBER-MoVQGAN or a new effective image encoder for generative models.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation SBER-MoVQGAN or a new effective image encoder for generative models

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:04:03.296907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-08T20:04:02.306771Z digest=sha256:5be5bb2c9dc2e25dccf4fed3575b06fc4ebdf010943ffccd6b93d28f22a81ce2

Observation a7e14180-f7fe-4cfc-8cdd-9b55d184716d · outbound

This paper cites Laion coco: 600m synthetic captions from laion2b-en.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Laion coco: 600m synthetic captions from laion2b-en

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:04:03.270760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-08T20:04:02.315752Z digest=sha256:76d46d2f7d4f4b8d984cb73856044a9d8b55996ac601647b2f5b85aa3e107e85

Observation 6348c906-7c74-42d8-b82f-6f110828263c · outbound

This paper cites Japanese and Ko- rean voice search.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Japanese and Ko- rean voice search

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:04:03.256999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-08T20:04:02.320176Z digest=sha256:47bbfc38935adee7e385cb7ed3ab0c187eeaf27232022b316a9cc668356da088

Observation a9cb6e2f-c90b-4a91-8bfb-2fda825d10cf · outbound

This paper cites Multi-task learning as multi-objective optimization.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Multi-task learning as multi-objective optimization

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:04:03.243667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-08T20:04:02.324692Z digest=sha256:2b0b8c03c02bf3f226324fdb38ce7a82c9dc0f43e4d2d35ed05f3d55f51b4d66

Observation 5b528a2d-eeca-40fc-a17b-4f30eeaa8852 · outbound

This paper cites Neu- ral machine translation of rare words with subword units.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Neu- ral machine translation of rare words with subword units

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:04:03.230277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-08T20:04:02.328915Z digest=sha256:9e797abb35bcd0651a6cd438afc428c2372797276efca47faf796d4067991734

Observation 763974f5-2b29-4718-9563-692d710ec268 · outbound

This paper cites Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.333188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.333188Z digest=sha256:daf273d1433f87ee2a31278ef622b7d60e186e0099ac7ea1fc410cdcc258a748

Observation a79b9ac9-7ebd-480d-99b5-c170ebc32383 · outbound

This paper cites Very deep con- volutional networks for large-scale image recognition.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Very deep con- volutional networks for large-scale image recognition

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:04:03.216760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-08T20:04:02.338045Z digest=sha256:1e1b5b509429fb0f9d4d38b44ee5d0dc93208997da0316f96f19c7f0cea11f06

Observation 3ed30314-8840-45f0-a088-3f3659833ef6 · outbound

This paper cites Towards vqa models that can read.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Towards vqa models that can read

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:04:03.202775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-08T20:04:02.342447Z digest=sha256:6d322862798a5d6d37fd891035a4a54548baefe88a006d1c2e4f1aa3466a7e29

Observation 1f0bda81-6ac2-4fb6-86f0-acb4f34fcbff · outbound

This paper cites Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.346378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.346378Z digest=sha256:d2d11cd41c1c09a5657e4bdb8f23556821709a81e5c7dde36d0a3731a15dd12a

Observation f9e9d59d-a110-4b33-a961-24993950a4dd · outbound

This paper cites EVA-CLIP: Improved Training Techniques for CLIP at Scale.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.350771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.350771Z digest=sha256:b8b165bbe4a1c7f3601db9a452b79966d08a3df11d4094e3b468d91fd492f349

Observation 6c28f588-8868-4027-bf90-ff5b54a965e5 · outbound

This paper cites Rethinking the inception ar- chitecture for computer vision.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Rethinking the inception ar- chitecture for computer vision

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:04:03.188991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-08T20:04:02.355141Z digest=sha256:e469a10aea803d5648ce64307a558786a90423e1a460267fb3ae63664e3a6983

Observation a51555ff-f1f5-4c9c-84e1-ec308258bc67 · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.359195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.359195Z digest=sha256:6361b2789e749ea5f481cfc12549dd8236ea2d898b50e2902765dfc82c35fa03

Observation cf41c74a-f4c7-4a83-92e8-5d0bf36ab6e2 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Gemini: A Family of Highly Capable Multimodal Models

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.363640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.363640Z digest=sha256:118964485b78c35ecf8b1372cdef8dd3a14610c3be94a6b5331c9ccba9378c3b

Observation 5b9250a5-9f06-4cca-ac87-fe9296e98b1f · outbound

This paper cites Zerocap: Zero-shot image-to-text generation for visual- semantic arithmetic.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Zerocap: Zero-shot image-to-text generation for visual- semantic arithmetic

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:04:03.175294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-08T20:04:02.368347Z digest=sha256:8578a14473cc5725e0782e07b9b3988e33f69906c7e1d4c03491553606891322

Observation 7d387ad0-554e-4504-b412-870f38858b29 · outbound

This paper cites Visual autoregressive modeling: Scalable image generation via next-scale prediction.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Visual autoregressive modeling: Scalable image generation via next-scale prediction

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:04:03.161572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-08T20:04:02.372743Z digest=sha256:2f00e562a1a1991eacf2a6ad524e7f6b379800d6a289d7927edb22e8427cfee9

Observation 5120663f-9535-485f-887a-87118f195826 · outbound

This paper cites Cambrian- 1: A fully open, vision-centric exploration of multimodal llms.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Cambrian- 1: A fully open, vision-centric exploration of multimodal llms

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:04:03.146572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-08T20:04:02.377487Z digest=sha256:e54686430770713330b1236b5005f1b095ae0d2a88695327f8cfba2c772d8aec

Observation 4ef7273a-8b91-487c-913e-eeea0dfea38d · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.381765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.381765Z digest=sha256:958f0ccdce9498267582d133d9e14ff1d3c30c015e3ffcc5607edd1c379dcc09

Observation 54ea1290-4ad9-4b80-a076-ad1d10fac364 · outbound

This paper cites Neural discrete representation learning.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Neural discrete representation learning

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:04:03.133397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-08T20:04:02.386558Z digest=sha256:78a1404f35f2e9cd106ef706e38b73fc94fa8824ed0de365f68aa9b59cbc86e6

Observation 3c79048f-2f57-4847-90c7-e066fc8b41dc · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.390941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.390941Z digest=sha256:bae601d10f43502f4db7f4a791e78d14fe3365628248835917f99e38612683c7

Observation fe84ed4d-84af-4a46-995e-7f2fdec335e1 · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Emu3: Next-Token Prediction is All You Need

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.395739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.395739Z digest=sha256:441539503cf734350e1f6d9ea12115a5cedb335ad846d9624f6f3df5a3a013bd

Observation afac164b-4635-4703-9a45-4c183a40e4aa · outbound

This paper cites Image quality assessment: from error visibility to structural similarity.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Image quality assessment: from error visibility to structural similarity

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:04:03.120627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-08T20:04:02.399952Z digest=sha256:f59e4542991a35ee288a0d5d3cbef483d68426f316fa79cad53fa52bdf65d0b0

Observation 5dbfe271-8bd9-4e92-94f5-23d6550ba9b5 · outbound

This paper cites Diffusion models as masked autoencoders.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Diffusion models as masked autoencoders

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:04:03.106327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-08T20:04:02.404211Z digest=sha256:c35901046338f19c3404979c6ab0ab219035859dcfaad87bef681aedc643d8f6

Observation 6ccbb9ee-5de6-4445-b8b7-efc5d7d6f211 · outbound

This paper cites Next-gpt: Any-to-any multimodal llm.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Next-gpt: Any-to-any multimodal llm

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:04:03.091696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-08T20:04:02.408117Z digest=sha256:b82bba9a548c76f6700f6986da83ec195baf2d137d09b452090c75661772a398

Observation 21678f6a-67f3-4732-80f8-d69a1c6719e0 · outbound

This paper cites VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.412305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.412305Z digest=sha256:68cca19c8d29e88c7fbc3a3f38a3df22d94058b79e114a218093ce2fba58ed0d

Observation 4ca01e07-7e29-478c-84da-50fb65cac641 · outbound

This paper cites Show-o: One Single Transformer to Unify Multimodal Understanding and Generation.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.416333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.416333Z digest=sha256:f876997e52ce25c42b7786b3b5fc639adfd74e44dd04de62ac95df6155404c5a

Observation 11781f39-f765-4cfe-ba2a-9bcd7bb75604 · outbound

This paper cites Vector-quantized image modeling with improved vqgan.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Vector-quantized image modeling with improved vqgan

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:04:03.075755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-08T20:04:02.420675Z digest=sha256:972ae3a8d24f1a40f4883cf44045569460caf8922ab679ac6746aa3c2dcfcdd7

Observation 1b4ddafe-09af-4ef2-94c8-7fe345d7b8b8 · outbound

This paper cites Spae: Seman- tic pyramid autoencoder for multimodal generation with frozen llms.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Spae: Seman- tic pyramid autoencoder for multimodal generation with frozen llms

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:04:03.061179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-08T20:04:02.424933Z digest=sha256:663a06a32fa048552d246386f5704742b498272faea7118d2c781a1903a5f7d7

Observation 64b8f8ad-79f2-4ea7-9c33-1fd36c74b33b · outbound

This paper cites Lan- guage model beats diffusion–tokenizer is key to visual gen- eration.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Lan- guage model beats diffusion–tokenizer is key to visual gen- eration

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:04:03.046218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-08T20:04:02.429265Z digest=sha256:267d7ad7827bcc099dba7db5953a0e7cca1d86c4d589910ccdc6f766124e3056

Observation bdece9a2-041f-4fa8-ab57-b22e7422889d · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.433718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.433718Z digest=sha256:918f5362cd862718b740090dc28a12c90b45512a69f608649c8a4ef5a3519109

Observation bf3a4357-bba0-42f1-b08f-455a2bc8821b · outbound

This paper cites Hellaswag: Can a machine really finish your sentence? In ACL, 2019.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Hellaswag: Can a machine really finish your sentence? In ACL, 2019

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:04:03.031418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-08T20:04:02.438286Z digest=sha256:7ff4d1e25c91a95e71e128372f10fc3a781611d623b8c169a705ba4e2f19037e

Observation 394930f6-e2d5-40eb-9008-0dfd22aa5589 · outbound

This paper cites Sigmoid loss for language image pre-training.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Sigmoid loss for language image pre-training

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:04:03.016591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-08T20:04:02.442574Z digest=sha256:280862f6875d55692b6bed254a3a68e980ce9838ce918dd2743ef20a7277a649

Observation 30127f51-8631-4c74-b980-39de953a66f7 · outbound

This paper cites The unreasonable effectiveness of deep features as a perceptual metric.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation The unreasonable effectiveness of deep features as a perceptual metric

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.447133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.447133Z digest=sha256:2425170c661562e0bcbdc36f02e054c701e09e50337476dce84bcae0fd424f86

Observation 0d66fc1e-a3a2-4698-87f2-d800bc7c8802 · outbound

This paper cites Image and Video Tokenization with Binary Spherical Quantization.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Image and Video Tokenization with Binary Spherical Quantization

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.451468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.451468Z digest=sha256:8d5d954588f8e715851350ff8623f06884320f4943c655ee0c2313ae1f4e7296

Observation 54aca9c2-4c8d-4dc7-8e30-9885c2e3eafe · outbound

This paper cites Online clustered codebook.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Online clustered codebook

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:04:02.992318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-08T20:04:02.456061Z digest=sha256:89c6a11e00639ab98d08280f10924d23b15d018691733a6b0ba9f054b352dd64

Observation 89097363-e6d8-42b1-b2ac-dea405150346 · outbound

This paper cites Movq: Modulating quantized vectors for high- fidelity image generation.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Movq: Modulating quantized vectors for high- fidelity image generation

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:04:02.977744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-08T20:04:02.460265Z digest=sha256:a5aac2af1dfd46fda96650e0cc9520dbc2c1e6a56267a2715abf1ce2392b96f6

Pith citing papers

Observation 6543b0fb-cbdb-48bc-b1ee-67fe13b2fed6 · inbound

Next Patch Prediction for Autoregressive Visual Generation cites this paper.

Next Patch Prediction for Autoregressive Visual Generation QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation

Reference 120

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:42.274831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:42.274831Z digest=sha256:6b4f40a620aad60143e5910c951878ad5049856ae056695a30fe213258f8c563

Observation 012862ec-3128-4e45-a7c8-4a0b725dbfcc · inbound

TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation cites this paper.

TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-15T23:09:11.045334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:09:11.045334Z digest=sha256:82809d7574efbaf610d86ee020c93ba529a7e7ae2d2012551a2358fb1cc8bf6c

Observation 4e72eeab-ef41-47a9-8833-b49d375efc23 · inbound

JarvisArt: Liberating Human Artistic Creativity via an Intelligent Photo Retouching Agent cites this paper.

JarvisArt: Liberating Human Artistic Creativity via an Intelligent Photo Retouching Agent QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-15T19:09:18.209986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:09:18.209986Z digest=sha256:4c1378277fa069240f0e9acd1c152d9496e9358170eb9c9ff2cb6f75bf8e5753

Observation cf687e1a-b5f5-4563-a753-5427cbecb4a8 · inbound

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation cites this paper.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.928203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.928203Z digest=sha256:28f12806a04b40e6e797c032bb34fb83a45ae7d703b309af777ef8cb18608417

Observation e389c8ad-e413-4311-96c6-dedb6e571120 · inbound

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings cites this paper.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:26.204540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:26.204540Z digest=sha256:22639ad3a02b1f5d2ddf1ad34323d2656b7868c136058b573d458f9f10304971

Observation 78ce552d-ceb8-4c56-a03b-f1a590da903a · inbound

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture cites this paper.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation

Reference 167

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T05:17:18.666961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:94a2e4288bcdc62446ff2a314cc43fd58e7bb1528c566e65340929aa32bf345e

Observation f6deb1b8-5c81-4bfc-b684-300bf9f9134d · inbound

Unified Pix Token And Word Token Generative Language Model cites this paper.

Unified Pix Token And Word Token Generative Language Model QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T21:35:04.591361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T21:30:45.355012Z digest=sha256:1be6254951b0d49c62c7080fe46da5cf453427b7347beba94ae22503cb08aa0c

Observation 231eadc4-01d1-4ae1-ba6e-263234ac0b94 · inbound

WinTok: A Win-Win Hybrid Tokenizer via Decomposing Visual Understanding and Generation with Transferable Tokens cites this paper.

WinTok: A Win-Win Hybrid Tokenizer via Decomposing Visual Understanding and Generation with Transferable Tokens QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation

Reference 107

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:08:15.768573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T12:04:19.761430Z digest=sha256:141152511a4f4e565e26c26e2f726f1505c39c6832d01f0b651bbc204fd9be67

Observation 0d553d7d-d8bb-4acb-bb92-f667a5f324c6 · inbound

Diffusing in the Right Space: A Systematic Study of Latent Diffusability cites this paper.

Diffusing in the Right Space: A Systematic Study of Latent Diffusability QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation

Reference 113

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:36:27.631392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-28T10:44:24.318786Z digest=sha256:8b4fdcb76cc0cb1b832478432263a75021640d2e6e95275b25c24424e8daaee2

Observation b7cfaf30-7abf-4a31-8663-d096ee5faa72 · inbound

NSVQ: Mitigating Codebook Collapse by Stabilizing Encoder Drift in Vector Quantization cites this paper.

NSVQ: Mitigating Codebook Collapse by Stabilizing Encoder Drift in Vector Quantization QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T05:27:39.702114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T13:18:47.472178Z digest=sha256:9a7938870d87105ddf9c49a155f03e2b3a485ab7c96a0b742c40e1f8fef5d92d

Observation 6d93c88c-7d14-4a5d-ae65-d42304fbf072 · inbound

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers cites this paper.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:38:28.819010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:10928cc07bbf0f4bc1f370c548e6bf0697d8a3c54dd8b35cc4eb787ae73871e0

Observation 3d54901a-9046-4a6d-9181-a448ae5835b3 · inbound

Unified Multimodal Autoregressive Modeling with Shared Context-Visual Tokenizer is Key to Unification cites this paper.

Unified Multimodal Autoregressive Modeling with Shared Context-Visual Tokenizer is Key to Unification QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation

Reference 71

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T20:28:55.821216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T01:18:03.846908Z digest=sha256:3755bd82fa01ae38517a984a0fd0fdff5530c758196220226aefe3b23a8424a2

Observation aff2e8d0-fbac-4880-9063-09d70c87d728 · inbound

dRAE: Representation Autoencoder with Hyper-Spherical Codes cites this paper.

dRAE: Representation Autoencoder with Hyper-Spherical Codes QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-01T05:45:53.491983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T05:45:53.491983Z digest=sha256:c533d5d79bac794c50f2c2366de38c4aab0f1c1ad08566e3acba65c5be70a38c

Observation 63171561-62da-4f9d-9d63-c7eec9039483 · inbound

Twins: Learn to Predict Unified Representations with Focal Loss cites this paper.

Twins: Learn to Predict Unified Representations with Focal Loss QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-01T04:29:50.752053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T04:29:50.752053Z digest=sha256:28d0ccb36211d1e2b1e6c75ffec2504c83d8f048fd6e9dc49ab12c1c3271b8e6

Observation dc275fa0-7300-4446-9dc3-cc9c0a1d4d35 · inbound

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes cites this paper.

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation

Reference 157

Resolution
unresolved
no resolver link, observed 2026-08-06T11:55:28.219048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:55:28.219048Z digest=sha256:8747cf72b6da2a132a6fa09b5b7928a0d6ab36612cd1c1a734a63619893ae0aa

Observation 59f237bf-22fe-4d5c-8e3e-c99693e2e1b9 · inbound

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes cites this paper.

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation

Reference 157

Resolution
unresolved
no resolver link, observed 2026-08-08T17:08:56.983795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:08:56.983795Z digest=sha256:6cf6635c5023faf82cea1a7d800c84051292a108dba4fff2ef9f69c99134c034

Observation d10f0e56-a29e-428a-8be7-940bff895fdf · inbound

UniSpace: Unified Visual Representation and Scalable Multimodal Modeling cites this paper.

UniSpace: Unified Visual Representation and Scalable Multimodal Modeling QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-14T04:35:00.075374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:35:00.075374Z digest=sha256:62b86cf02bbbf10d8381cb0168dde0f778651772cdd12eae080f425a224c7c6e