Pith. sign in

Paper Citation Record · LEDGER

TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

As of 21 August 2026, this Paper Citation Record lists 77 of 77 outbound references and 47 inbound Pith citation observations for arXiv:2412.03069.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.03069 v2

Coverage vector

measured 77 of 77 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T22:54:20.710373Z

measured 124 of 124 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 47 of 47 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T04:41:47.730706Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

77 of 77 outbound references displayed

  • verified exact0
  • verified fuzzy20
  • unresolved55
  • parse uncertain1
  • malformed identifier1
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 6d0b9588-1f0b-4fde-98c5-e1d5b1e072db · outbound

This paper cites GPT-4 Technical Report.

TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T22:54:20.432103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:54:20.432103Z digest=sha256:07831190aacfd6d06e07ef5abfc15ca65b0b88677a6b80ac724dc2e9513d90a3

Observation 9a9c9a12-e813-42db-b699-35d41b04f9fe · outbound

This paper cites Qwen Technical Report.

TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Qwen Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T22:54:20.437208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:54:20.437208Z digest=sha256:0d2061a6ceb4d66e700db7086d3fdbf5274c8c6de36fbcf88d26d7cff6ce6654

Observation a600daf4-fa30-4998-98c4-aaf2a5cf9cc6 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T22:54:20.440971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:54:20.440971Z digest=sha256:dcaeef9b7b608a5db559c8dbd4e80a11d25a42352e6e8848f594eb75ef6843dd

Observation 1e42f739-a626-4adb-9fdb-6eac2aac593f · outbound

This paper cites Improving image generation with better captions.

TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Improving image generation with better captions

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:54:21.533353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T22:54:20.444874Z digest=sha256:fdf4a23f46a2f0d17797c1576a7dbb85e1ee5bc4e4f52e8582647010da10bafb

Observation 0126eda4-bacf-4cd4-8c8e-490363c9c0a4 · outbound

This paper cites Coyo-700m: Image-text pair dataset.

TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Coyo-700m: Image-text pair dataset

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T22:54:20.448365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:54:20.448365Z digest=sha256:c2eabb7682d99641870c2c92afca662c3c6c02ff9285cad8d82e92d597b34825

Observation 83b9e60b-9a7b-4007-8f90-4783bc6ac15f · outbound

This paper cites Muse: Text-To-Image Generation via Masked Generative Transformers.

TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T22:54:20.452018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:54:20.452018Z digest=sha256:630bfbd9fe1dcb0cf16e1783ff45166bddabffddca6c214be25e39d4549b0b76

Observation bde1a2d9-d9a1-4920-9081-92afa09a5756 · outbound

This paper cites PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis.

TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T22:54:20.456159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:54:20.456159Z digest=sha256:e53e532faf95e7012a7f01825ac0ee7a127408d93dff3650eede02e61cbe2152

Observation ee66adad-dd0e-4eae-b10f-14d17ca91814 · outbound

This paper cites Vitamin: Designing scalable vision models in the vision-language era.

TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Vitamin: Designing scalable vision models in the vision-language era

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:54:21.515254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T22:54:20.459983Z digest=sha256:b3646cbf38799dd8f4f8f21c9e55ea0b8e8b146358cac84b0918ee92e52859e6

Observation 3bcc1fd3-13ea-4a9a-ab99-7c06c4a9641e · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T22:54:20.463765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:54:20.463765Z digest=sha256:9b5503751cec467962a7283d81dc5d4d66293327f77bca95b738a2ba2ebec177

Observation 62f6e974-6380-4f05-8464-83a3484c7266 · outbound

This paper cites Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.

TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T22:54:20.467232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:54:20.467232Z digest=sha256:541efd2eaaf0962860fee3f2b2a4ae7638ca2da8d63378398c569f1bbe1ac3e2

Observation d0bcf96c-b243-4657-8491-d244bdf98ed7 · outbound

This paper cites Scaling vision transformers to 22 billion pa- rameters.

TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Scaling vision transformers to 22 billion pa- rameters

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T22:54:20.470257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:54:20.470257Z digest=sha256:e5c4e3bb324c7335e82a204d22f571b122ea319a75398c6a354c952c68f75ec5

Observation 277ced2b-40c5-4f62-8949-5edbf2e6179c · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Imagenet: A large-scale hierarchical image database

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T22:54:20.473283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:54:20.473283Z digest=sha256:ab53e407709ab72808771ff655c411ed59858bfbb11b81a1c6765466a0a2854a

Observation 3b41a819-0890-419e-aa81-45f85d443dd6 · outbound

This paper cites Taming transformers for high-resolution image synthesis.

TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Taming transformers for high-resolution image synthesis

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:54:21.483238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T22:54:20.476220Z digest=sha256:49ed39f4c66de8c89036725a1f0c644d214cc73eb00c79f2d57ac919ee682cc7

Observation 8442a6c1-1f43-4240-a106-4becb8e2f815 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T22:54:20.479149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:54:20.479149Z digest=sha256:1643931781f8ecfd687dd0fcec7aff9a4c6199aa4df58da1b2fe264096f40c28

Observation 6b935a5e-e2e8-49b5-a594-cceb16895586 · outbound

This paper cites Geneval: An object-focused framework for evaluating text- to-image alignment.

TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Geneval: An object-focused framework for evaluating text- to-image alignment

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:54:21.471133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T22:54:20.482418Z digest=sha256:5ec7e1f3141ab5a5672bc85a2069f71516d72b3ece63aba3f29b6fa267854785

Observation 34ba906f-63fd-43aa-9dc1-7255b69946b8 · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing.

TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T22:54:20.485745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:54:20.485745Z digest=sha256:682628d1e04596dbe69473afbf2cc3b4fc8acc4169353566215908791536b5b3

Observation 77e81f5a-c599-44c3-8940-1b40d1538d1a · outbound

This paper cites Classifier-Free Diffusion Guidance.

TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Classifier-Free Diffusion Guidance

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T22:54:20.489286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:54:20.489286Z digest=sha256:0af358aa16ae1f2a5bb35f367a1a720218db21ac49118c472fdfdb9d287086b2

Observation 1eb8ae94-2fc6-4c62-8163-8cd34638aa07 · outbound

This paper cites ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment.

TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T22:54:20.493411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:54:20.493411Z digest=sha256:f7f98b25932e732485b76639b36d2ae17c3d25f253e74bd2d3e710965fce5f8a

Observation 275b0fea-09db-4020-94dc-f24fb0517144 · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:54:21.452078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T22:54:20.497296Z digest=sha256:b7d3d892c016c5fef2ba89befc2ecdcf669d83b939f472697ef45b1e191dca2c

Observation 9465b5a9-afb4-4bd5-9d1d-47c20d3b6441 · outbound

This paper cites A diagram is worth a dozen images.

TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation A diagram is worth a dozen images

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T22:54:20.500771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:54:20.500771Z digest=sha256:d150013a9febea1bcf4a593368e29a89439a9a155cbf8d7c3bfb02767352656e

Observation cc895db4-5e7e-4816-9e78-2b335767cbc0 · outbound

This paper cites Autoregressive image generation using residual quantization.

TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Autoregressive image generation using residual quantization

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:54:21.419022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T22:54:20.507767Z digest=sha256:00eb6fee5ee9394a57525752ef01f3970f668aceb772ce58ee4546a41b5e3c30

Observation 1b2f7917-e8a3-4552-8f38-da51b150f588 · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T22:54:20.511365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:54:20.511365Z digest=sha256:c5bcb6175aac318c9fc846494291fe3c8cd1b87433c38fd23c43f9f4d4071687

Observation 6c17fdfb-fa90-4736-8ba1-43dc7d9771e7 · outbound

This paper cites Seed-bench: Bench- marking multimodal large language models.

TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Seed-bench: Bench- marking multimodal large language models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:54:21.405613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T22:54:20.515091Z digest=sha256:80d6084b5ffb2981f3610c547306534599fc8a6eb581bf088eb554138193e66e

Observation baa69641-420a-4977-aaee-7f0f29323e22 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T22:54:20.518661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:54:20.518661Z digest=sha256:bd182ae2e6ce3c441bdde66cea5b3942956f6904f661f775f253d8d3e5886524

Observation cc9ee8dd-5b24-4b43-a77a-ddfbe1152e28 · outbound

This paper cites ImageFolder: Autoregressive Image Generation with Folded Tokens.

TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation ImageFolder: Autoregressive Image Generation with Folded Tokens

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T22:54:20.522317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:54:20.522317Z digest=sha256:55f4ef0f4c424824887e424057b1202fc89f42c096be1a5f6bc1a8e082fa6c6b

Observation 882d1ad6-b025-4264-988f-d5b5ab787d8d · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Evaluating Object Hallucination in Large Vision-Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T22:54:20.526195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:54:20.526195Z digest=sha256:4d8755598f40ed6e04d1c96a30a6b71a35c6529b0719eaa8818296d66999fdb8

Observation 8190837a-a921-4522-8ec4-960382bb187b · outbound

This paper cites Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models.

TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T22:54:20.530112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:54:20.530112Z digest=sha256:0fde1770576e7b443c2006d577f4d8b6e179b349e1d2df6075b2a60e2234c30b

Observation cd2555aa-4a88-491b-aa82-b381cddce663 · outbound

This paper cites Lumina-mGPT: Illuminate Flexible Photorealistic Text-to-Image Generation with Multimodal Generative Pretraining.

TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Lumina-mGPT: Illuminate Flexible Photorealistic Text-to-Image Generation with Multimodal Generative Pretraining

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T22:54:20.534492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:54:20.534492Z digest=sha256:68ffa21b489e3665c6abd40cd29b52c688a0687d773478aaee2bbd8eb878a098

Observation bd19d9d4-709d-4af7-885b-579623b7efa9 · outbound

This paper cites Improved baselines with visual instruction tuning.

TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Improved baselines with visual instruction tuning

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:54:21.383765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T22:54:20.538662Z digest=sha256:95a6a76c4afb847c4517f33996b36a61487b89b4979d985b3074b7a4936c12c3

Observation 81e0e3a7-ba44-455b-9d72-1a14cfebe7db · outbound

This paper cites Visual instruction tuning.

TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Visual instruction tuning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T22:54:20.542673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:54:20.542673Z digest=sha256:8abb370b81d761d1eb790e14ddd527bae057beec0c2f9ffab410c3b345f66db1

Observation 0ecdae2f-e6b6-43b7-b059-12bb71766442 · outbound

This paper cites World Model on Million-Length Video And Language With Blockwise RingAttention.

TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation World Model on Million-Length Video And Language With Blockwise RingAttention

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T22:54:20.546454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:54:20.546454Z digest=sha256:9b3d1c6651e2229f1e2d5f409e425f6949cc3a71a5fdd78483655569fde341e9

Observation 91137bbb-6da4-4879-b7ed-4efac947e0ca · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? In European Conference on Computer Vision, pages 216–233.

TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Mmbench: Is your multi-modal model an all-around player? In European Conference on Computer Vision, pages 216–233

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:54:21.362830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T22:54:20.550584Z digest=sha256:8c32d318f69e53470c32c155a0d11cd06d5af2e537d506732e0716fa374e8300

Observation 25bbee17-e5d8-4dfa-a61d-269935bd978d · outbound

This paper cites Open-MAGVIT2: An Open-Source Project Toward Democratizing Auto-regressive Visual Generation.

TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Open-MAGVIT2: An Open-Source Project Toward Democratizing Auto-regressive Visual Generation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T22:54:20.554558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:54:20.554558Z digest=sha256:aceba1e2ff10b25881cd0b863caae1ae6dd2421967c74fce803fb4dbd2c73ca5

Observation 24522319-7885-4295-8579-54bddb232fc1 · outbound

This paper cites STAR: Scale-wise Text-conditioned AutoRegressive image generation.

TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation STAR: Scale-wise Text-conditioned AutoRegressive image generation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T22:54:20.558121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:54:20.558121Z digest=sha256:970295b6e4d8154cd3df20e51299f1195c974e527bad9584389a139d91a60529

Observation d8332b65-6797-4a6c-b077-13e575320895 · outbound

This paper cites BEiT v2: Masked Image Modeling with Vector-Quantized Visual Tokenizers.

TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation BEiT v2: Masked Image Modeling with Vector-Quantized Visual Tokenizers

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T22:54:20.561377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:54:20.561377Z digest=sha256:0d047810cece2df7d15ef0e09fe7f25862a7b3beea76c1f4f01a9849aeb2ad4c

Observation 2a49e759-944f-4cc3-be32-39fcc5dc92f6 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T22:54:20.564718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:54:20.564718Z digest=sha256:c1528ee1b38161e4d0a75cca882130b363f130d2d919a889ec0862d717721988

Observation 236c09c7-6794-46f9-b308-3c4730825e3f · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Learning transferable visual models from natural language supervi- sion

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:54:21.349471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T22:54:20.568157Z digest=sha256:b8b5ac456ee0f484130908f2775e0fb220571fc1baec5938b6550912514181f9

Observation d8f784aa-ffff-4faa-a70f-414a7e0a1104 · outbound

This paper cites Zero-shot text-to-image generation.

TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Zero-shot text-to-image generation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T22:54:20.571310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:54:20.571310Z digest=sha256:df912ab02c4a99d733133ba3eae3c7e0656799726172ef233fa2e558844c77f9

Observation 8d85a53c-c0cc-4ca7-a145-53fe4e229c54 · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T22:54:20.575726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:54:20.575726Z digest=sha256:54868c4ac64ae70dd8efabe569feb275e7b3770bb2fa8671ebe55c89c2298cd1

Observation e145cf4c-c565-448b-9a35-40ca5035c484 · outbound

This paper cites Gener- ating diverse high-fidelity images with vq-vae-2.

TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Gener- ating diverse high-fidelity images with vq-vae-2

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T22:54:20.579618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:54:20.579618Z digest=sha256:f9a24c5637399db5eb654df4998bf711e29be3e31abe28751f075d44d1204dca

Observation 38b9f5a3-b13a-49ec-97d3-3cc4ef2844cd · outbound

This paper cites High-resolution image syn- thesis with latent diffusion models, 2021.

TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation High-resolution image syn- thesis with latent diffusion models, 2021

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:54:21.324026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T22:54:20.583099Z digest=sha256:e096b689a543e1105ed5fb6aa1d824a39a1191fb8691f6ca6332dddc3b46b5f0

Observation adf34022-b3d7-42cc-8bec-9feb71bf7ddf · outbound

This paper cites Laion-5b: An open large-scale dataset for training next generation image-text models.

TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Laion-5b: An open large-scale dataset for training next generation image-text models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T22:54:20.586572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:54:20.586572Z digest=sha256:a85ac161970ef1b1dc7bece26cfb8d822042b584534f2ff7c4426c77c20a4783

Observation a54e9c2c-08d4-49d7-a53f-49f48bad442b · outbound

This paper cites Towards vqa models that can read.

TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Towards vqa models that can read

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:54:21.304060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T22:54:20.589992Z digest=sha256:a2fd04520c3455ea42ebe6ab285a191e38f9a11a8574719fe694a5f2263607a5

Observation 32dc96f1-7c70-44c4-ad5d-7dedb75769a3 · outbound

This paper cites Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation.

TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T22:54:20.593468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:54:20.593468Z digest=sha256:54002c107d1d506a69b6a4b9370e4cd92808ce7706750e671fd8c27d2b260e94

Observation 94e41039-f9a3-4a65-aece-a7f03e0d5363 · outbound

This paper cites EVA-CLIP: Improved Training Techniques for CLIP at Scale.

TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T22:54:20.597289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:54:20.597289Z digest=sha256:08c8a1b7d63b0774906693086971adfbd8722b55bc2813e18aadcd3f2efeece8

Observation 39475a2a-77f4-4587-8cf4-16dbd376f497 · outbound

This paper cites Emu: Generative Pretraining in Multimodality.

TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Emu: Generative Pretraining in Multimodality

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T22:54:20.600724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:54:20.600724Z digest=sha256:983a112f38984376469c213fa93b9805c7d3eb9fe85e71358ff3c8708b9e77fa

Observation c578ad44-d825-4931-a306-606d9cd9bbdc · outbound

This paper cites HART: Efficient Visual Generation with Hybrid Autoregressive Transformer.

TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation HART: Efficient Visual Generation with Hybrid Autoregressive Transformer

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T22:54:20.604450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:54:20.604450Z digest=sha256:5a1e24fb2ee7f74ea37f5ade7b72d48c4501291e1890d9b0ed19e04086128c1b

Observation 9661d7ca-622d-44f1-af91-f94d215ba70b · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T22:54:20.608473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:54:20.608473Z digest=sha256:c43904da08d9f94eb71c9477a754344abfd5c586e78362e80004455013a1c116

Observation d99523a9-a43e-4a55-9857-3de3ba680703 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Gemini: A Family of Highly Capable Multimodal Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T22:54:20.612561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:54:20.612561Z digest=sha256:189615ea808fa47cb23c08f903b9072ae2a95c467249ecd3e04ef5adff79ac42

Observation 4162a1dc-3979-49d4-ae3a-06d2d8f3a122 · outbound

This paper cites Qwen2.5: A party of foundation models, 2024.

TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Qwen2.5: A party of foundation models, 2024

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:54:21.291175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T22:54:20.616198Z digest=sha256:05e6dd1cd56a8e3e530bdac3432c586a9cf79a760e98e9dff9439c8ea605a7ea

Observation 830a81bd-5396-4b1b-a97e-418a0757aa38 · outbound

This paper cites Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction.

TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T22:54:20.619798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:54:20.619798Z digest=sha256:bd9769275dd80005e3370f97bd63540067021dc1c22c29070c5c00bc716700c8

Observation 5e749ca7-9d6e-4057-9833-238eb6aa9bc6 · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.

TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T22:54:20.623671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:54:20.623671Z digest=sha256:6f0200da0f3a03c5e70714ec5acd17996a9f9a06bd7020649aece492f73d87fb

Observation efa95757-7033-4ac6-bdf8-8aad11a1e1bd · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T22:54:20.627445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:54:20.627445Z digest=sha256:2aa0f39b9f8b1c1c4805e76d8958608047c185d8eec28e547cc27ecd55ea514e

Observation bbc79bd2-d88d-4710-b43e-82da69f3eec0 · outbound

This paper cites Neural discrete representation learning.

TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Neural discrete representation learning

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T22:54:20.631321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:54:20.631321Z digest=sha256:34a4f92401d44b64ee7441c6bde57080c802ea1e2785ea4b23082740e9b7e3fc

Observation 3e529a2b-a17c-4374-91d6-c2067f818c68 · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Emu3: Next-Token Prediction is All You Need

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T22:54:20.634598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:54:20.634598Z digest=sha256:b2668ea31c3a892bdd10bd245d367292d0dab4470138267853b25a66dda5cb37

Observation 35d89243-fe77-40d6-9d27-edea4a0642c6 · outbound

This paper cites Small-scale proxies for large-scale Transformer training instabilities.

TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Small-scale proxies for large-scale Transformer training instabilities

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T22:54:20.637724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:54:20.637724Z digest=sha256:ab511137e423518426e67e81f782dddece5491b737f312ef1cea4ad1bc3c4b9d

Observation e178fc9a-8e52-494c-a1f5-4f06585b27a2 · outbound

This paper cites Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation.

TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T22:54:20.641452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:54:20.641452Z digest=sha256:517b185607c5c6a32265dc1c63481de0d96c75c94ed544458f07061ed926b95b

Observation afb78ad8-e84f-47e9-ba30-fc20ec80c3e9 · outbound

This paper cites NExT-GPT: Any-to-Any Multimodal LLM.

TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation NExT-GPT: Any-to-Any Multimodal LLM

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T22:54:20.644671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:54:20.644671Z digest=sha256:e6d6cd30189c0f9a4cb71348b5219a03330068e9661b27b2e587cc0108ceaa8a

Observation 87a6c221-2386-4a65-871c-cc21054f86ae · outbound

This paper cites GPT-4o: Visual perception performance of multimodal large language models in piglet activity understanding.

TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation GPT-4o: Visual perception performance of multimodal large language models in piglet activity understanding

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T22:54:20.647753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:54:20.647753Z digest=sha256:b198a237a28268a9c0470db9b7de0cd4fad88f22ac61fdebc226a9e7f81862da

Observation 2f2b19e7-e8e9-4aa4-85e0-30d1f732fc58 · outbound

This paper cites VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation.

TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-11T22:54:20.651175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:54:20.651175Z digest=sha256:9bf8e00a2a9f3d5d4c7c3a0793979a4f2f1b88bd0aa31bc7671bf3c0db4d9ffe

Observation 2f44a8ff-e9ce-412b-8206-ffe46e8cf1d5 · outbound

This paper cites Realworldqa, 2024.

TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Realworldqa, 2024

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:54:21.271690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T22:54:20.654934Z digest=sha256:09f8d112013b9db63e48aef70e50fde0648d321372b25950f75254af16fc56f1

Observation fcbb2340-86c3-4e17-8b36-e5a415ef2cec · outbound

This paper cites Show-o: One Single Transformer to Unify Multimodal Understanding and Generation.

TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-11T22:54:20.658340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:54:20.658340Z digest=sha256:fcea42d3d16666d74d055ec1379b25bbc1d153dd608eeb50e681bc8ea2412716

Observation d6e2e3f1-cc77-4f06-86c4-fc44a43933be · outbound

This paper cites Imagere- ward: Learning and evaluating human preferences for text- to-image generation.

TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Imagere- ward: Learning and evaluating human preferences for text- to-image generation

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-11T22:54:20.661987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:54:20.661987Z digest=sha256:cb44ed1952c548b5d6800f2de5592dfe11efbed849fa48db4e643d75f0103dff

Observation f1be0404-742e-4ae5-a5f0-6ac69a505cde · outbound

This paper cites Vector-quantized Image Modeling with Improved VQGAN.

TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Vector-quantized Image Modeling with Improved VQGAN

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-11T22:54:20.665388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:54:20.665388Z digest=sha256:3839e0be4727d9c9ce278b884ae039a47b1a8c500cbeb45916aa49dba18c3ffe

Observation 6ae76a07-7f94-4cd5-b142-23eecc37f21b · outbound

This paper cites Scaling Autoregressive Models for Content-Rich Text-to-Image Generation.

TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-11T22:54:20.669264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:54:20.669264Z digest=sha256:8f2bc1dc1e573e895d899f1b73eea547fe45a418e0a19c674a22a57b9c108b61

Observation fbb3738f-1681-489b-ba3c-9f5bb0bf6e5d · outbound

This paper cites Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation.

TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-11T22:54:20.672941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:54:20.672941Z digest=sha256:c5babe8acbddbe0d5317a0711829f9baf4a64d2de79f70890d31880c187fe6f4

Observation a7433235-cd95-4bab-90c6-6caff9664116 · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-11T22:54:20.677080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:54:20.677080Z digest=sha256:64dc5d746415a69df1bbde1ece81acaa0e447187014163dcca2ab62de32f0162

Observation 97924604-63c1-4e25-a780-deb3a4ece033 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for ex- pert agi.

TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for ex- pert agi

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:54:21.252796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T22:54:20.680901Z digest=sha256:fa55534ce0f87d7d300df0311600ff63dbc492e8b756a3abf0ce3aae7011b0b4

Observation 790683d6-8e4c-4ea1-903b-2ed8dc57ebe6 · outbound

This paper cites Sigmoid loss for language image pre-training.

TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Sigmoid loss for language image pre-training

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:54:21.240254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T22:54:20.684367Z digest=sha256:58a9c95cd2326e7d57fa62d2f84e4da3ffe0170a365e2b948bfe90743cf61a6d

Observation af8f51cb-c556-4eed-a38d-a92ff6ef9f5b · outbound

This paper cites Regularized vector quantization for tokenized im- age synthesis.

TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Regularized vector quantization for tokenized im- age synthesis

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:54:21.228177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T22:54:20.687701Z digest=sha256:bd551fca83146f89742beed0ec075b41553bd58975518aa8b1797ba4c1ca59e2

Observation 619eabbd-f148-493d-8c4e-a0bb68dfbf66 · outbound

This paper cites Lmms- eval: Reality check on the evaluation of large multimodal models, 2024.

TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Lmms- eval: Reality check on the evaluation of large multimodal models, 2024

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:54:21.216866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T22:54:20.691146Z digest=sha256:08e022ac28de2d65acc3f4003508c1fa7f3b741826e5559b09261952a00213ea

Observation 6bc54eb1-c27a-4c9a-aca0-85f77db61c34 · outbound

This paper cites Tinyllama: An open-source small language model, 2024.

TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Tinyllama: An open-source small language model, 2024

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:54:21.205978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T22:54:20.694544Z digest=sha256:c097e2c3563bfa32efc3951b608300603d75dc314536c3875a9950c8895f6b22

Observation 2fa11b19-356d-4ae3-bce8-de512028d510 · outbound

This paper cites Movq: Modulating quantized vectors for high- fidelity image generation.

TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Movq: Modulating quantized vectors for high- fidelity image generation

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:54:21.194261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T22:54:20.698112Z digest=sha256:4be1d3d57205db7109a15b8a2c9bd626fd239e9024f89042d9df041ef3074a2f

Observation fc46919f-a677-458f-a2bd-1bc7a87618d5 · outbound

This paper cites Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model.

TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-11T22:54:20.702601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:54:20.702601Z digest=sha256:b2770a1389a8b9a6c020eca0debd311aab109ff876f95259a95d18700e5b06a1

Observation be76c2cd-affe-4c68-9975-795ac727ba99 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-11T22:54:20.706909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:54:20.706909Z digest=sha256:d9bf30334d7e23160a86a6844be18c3572ffb277846c6f78fa4d214a7e219626

Observation 52a578d8-715e-4592-bc2e-8762c087d474 · outbound

This paper cites Scaling the Codebook Size of VQGAN to 100,000 with a Utilization Rate of 99%.

TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Scaling the Codebook Size of VQGAN to 100,000 with a Utilization Rate of 99%

Reference 76

Resolution
malformed identifier
no resolver link, observed 2026-08-11T22:54:20.710373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:54:20.710373Z digest=sha256:dfd1ec4501e90cb0766cdc358157ce29684a2f0d43d8d7ee5133e14b8e9b1a27

Observation 83bfaea3-6f60-4619-b6bd-e5e90a504cde · outbound

This paper cites an unresolved cited work.

TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Unresolved cited work

Reference 251

Resolution
parse uncertain
raw_fallback, observed 2026-08-11T22:54:21.432452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T22:54:20.504263Z digest=sha256:23c0342c930f7ee67c1e9a4f0340e8015b2bfbed7c00a569e33102ebd193d0cd

Pith citing papers

Observation 9eae1cb8-a1ef-4948-a502-f8e92d36566b · inbound

MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding cites this paper.

MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T12:37:38.500293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:37:38.500293Z digest=sha256:72f4756e48c86d2a3ba4e4fe0971ab48fe4d9748892e13859f8bef023aa2fc38

Observation 88f6eeff-8513-4a51-849a-7144f2448515 · inbound

Scalable Image Tokenization with Index Backpropagation Quantization cites this paper.

Scalable Image Tokenization with Index Backpropagation Quantization TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:46.411334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:46.411334Z digest=sha256:d9861b1fe9f4799d355e4f147af12cea4db80b1ab24ce581907d2be72349f8d2

Observation f0e22818-1e16-4dc4-ae1e-b17e0b3e6e2b · inbound

Liquid: Language Models are Scalable and Unified Multi-modal Generators cites this paper.

Liquid: Language Models are Scalable and Unified Multi-modal Generators TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T21:35:40.413579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:35:40.413579Z digest=sha256:7d03d443d6eabffa312d64bcc5d9cd5bf7972c3f509b69cf4975c5fa51c5dd94

Observation 5df510ed-3dba-488d-a51c-b6cb1d70606b · inbound

SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization cites this paper.

SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T17:59:36.547616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:59:36.547616Z digest=sha256:44217fc5b829f05e4dc3b871cf8269f5cc0642d643ca99f36d3504e6ae4c38d9

Observation 43d0d774-8cf4-4c06-babb-9e4c11199368 · inbound

Next Patch Prediction for Autoregressive Visual Generation cites this paper.

Next Patch Prediction for Autoregressive Visual Generation TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:42.093550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:42.093550Z digest=sha256:eb85449f224fa3ea826dcf07b5caa029a507a92025e5d2a1cc9200b5ed70ede9

Observation 5b9e76d7-d684-4499-8d69-a1284bc0d394 · inbound

Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling cites this paper.

Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:14:53.158653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-11T08:14:52.890145Z digest=sha256:4b73e9df648e435b8bec75a62d40c0cff50e983f5bd064dac9c081369a9b7e1f

Observation 996164ff-57dc-495c-91de-c36211141636 · inbound

On Fairness of Unified Multimodal Large Language Model for Image Generation cites this paper.

On Fairness of Unified Multimodal Large Language Model for Image Generation TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-09T04:52:40.597797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:52:40.597797Z digest=sha256:2fa6dacae12fe695eb7c7c6445cf73e0e4608c913e606926b84b59be6142ebc6

Observation bfbc628b-1402-4466-848a-af856d58595a · inbound

Masked Autoencoders Are Effective Tokenizers for Diffusion Models cites this paper.

Masked Autoencoders Are Effective Tokenizers for Diffusion Models TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-09T04:47:30.399979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:47:30.399979Z digest=sha256:0846a208c9fbbc563b2ae8617019a622dfc73f7393c5db967d07efbe169470a3

Observation d12143ba-0a32-42a9-a00a-e6f1a22c4ae1 · inbound

Multimodal Medical Code Tokenizer cites this paper.

Multimodal Medical Code Tokenizer TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-09T00:42:57.117020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T00:42:57.117020Z digest=sha256:63ab3e657ae492fed59dbd31376d09040da750ad93ee635ebc3959de9a3b8832

Observation 88e622de-3e0f-4eca-bf68-b219530fdf91 · inbound

UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths cites this paper.

UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T15:24:46.387548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:24:46.387548Z digest=sha256:0c3e5f188e89676650dda69f7bfb56c48eecb8ba9ede34de4d239274a5c295df

Observation cca3d66c-37d2-482a-8e41-0ba9f8c5a2c1 · inbound

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models cites this paper.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.809314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.809314Z digest=sha256:09396450f80cab51d8c45c9a7af3e2f950bbe22621574f5a5e0f85a5fd678258

Observation f55f4173-db02-4d93-a5e4-96030d315cea · inbound

DualToken: Towards Unifying Visual Understanding and Generation with Dual Visual Vocabularies cites this paper.

DualToken: Towards Unifying Visual Understanding and Generation with Dual Visual Vocabularies TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-22T23:52:16.721256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-22T23:51:43.934329Z digest=sha256:dd98710f5c3014b6c4e568bc0a10652eb72530b5b91081f53ceb56a0d514ba62

Observation e6a9ae03-2e7b-4916-846d-3c9c7d390273 · inbound

T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT cites this paper.

T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-16T04:41:47.730706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:41:47.730706Z digest=sha256:09fd971f69c605e05e1e17a8d72aa48f1072c95572f15b7d3be3cf13a600e790

Observation 7807693a-b4d5-4c74-8b19-ebf9a6b6560e · inbound

Ming-Lite-Uni: Advancements in Unified Architecture for Natural Multimodal Interaction cites this paper.

Ming-Lite-Uni: Advancements in Unified Architecture for Natural Multimodal Interaction TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T00:54:39.419665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:54:39.419665Z digest=sha256:bd8392327d6e266e37589786c83f3d35c9386ec0864911793bd3f70fdd10ce0f

Observation e3a2bb0e-85f8-4263-95f9-52aa70fd6cad · inbound

Position: Foundation Models Need Digital Twin Representations cites this paper.

Position: Foundation Models Need Digital Twin Representations TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-16T04:36:49.979479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:36:49.979479Z digest=sha256:7933bebd191def39b09dbdc0f079c460f1e773b46d05e47486bcfa638d0422e8

Observation b2fffa2c-5aa1-4434-9c32-e0510cea9d2e · inbound

TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation cites this paper.

TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-15T23:09:10.873424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:09:10.873424Z digest=sha256:de9f123c63a46c26a4c1f3b3d1f8c37e1faf7a25456a0c827f078186d8ec47ff

Observation a9e53c24-575a-4fc3-b371-025ed7011778 · inbound

Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation cites this paper.

Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-17T07:24:04.603862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-17T07:24:04.460276Z digest=sha256:2db32b492cc87ecf243fcdb58ba3e84d94e43cd19aa937d8e620273c7826aa4c

Observation 2356c3ea-3ece-449c-b702-cc2e7d541be6 · inbound

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning cites this paper.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.561032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.561032Z digest=sha256:5e0783b7052b8910ec8f5a33a6ab969c7f7f27710bdd8265151ddb37e85402e1

Observation 6869a934-ada0-4a3d-a060-0bf823273652 · inbound

BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset cites this paper.

BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:34:27.042984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-11T23:34:26.878354Z digest=sha256:99c211c2098141f22fb983048926f715b3ea06c3fa46c297906cd394fc9149fb

Observation ec67038c-8d8a-4cdd-b29e-5fc2732a4d94 · inbound

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO cites this paper.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.261282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.261282Z digest=sha256:08641f6b2c8e2995bf8608180b0b7eb5691f882602032fbfcbdd4880c5bd2585

Observation 078059fc-f144-40af-b551-fc62dd3a4209 · inbound

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation cites this paper.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:54.732431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:54.732431Z digest=sha256:9600ec3bc4a223195514496b5cdd22c345ff650e33d5243e2833422b9a5784db

Observation 98fed8a3-a44b-49fa-8069-5a61c92b004b · inbound

Emerging Properties in Unified Multimodal Pretraining cites this paper.

Emerging Properties in Unified Multimodal Pretraining TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-10T16:23:42.211623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T16:23:41.854132Z digest=sha256:7ca0e4cca18bf60eb2bde8218ce25797ca04c5fbe978b0e2c9df29dca00583cb

Observation 49711da3-9b0c-44c5-8b6b-ba35915ed162 · inbound

TokBench: Evaluating Your Visual Tokenizer before Visual Generation cites this paper.

TokBench: Evaluating Your Visual Tokenizer before Visual Generation TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:42.504964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:42.504964Z digest=sha256:7f4252ccd11210504f73532e9c1f566d0a526791d47a4a3c749aee1b828da2fd

Observation d9455675-b83a-4d95-beea-188e4e763b43 · inbound

FUDOKI: Discrete Flow-based Unified Understanding and Generation via Kinetic-Optimal Velocities cites this paper.

FUDOKI: Discrete Flow-based Unified Understanding and Generation via Kinetic-Optimal Velocities TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T14:04:59.258830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:04:59.258830Z digest=sha256:09b67dd5496e361340103e32ed9e226ba959b58d488bb2f15e565314212d771f

Observation 8ed1468c-7cf9-4a0c-b1c7-8dcfafffb433 · inbound

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning cites this paper.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:24.689189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:24.689189Z digest=sha256:03bbf86bc30e57e13b7e007e0c1852d68239fa0e57cded7aa5148634d649f5df

Observation 29dcb26a-57e5-44e0-8b57-e43c8271223a · inbound

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation cites this paper.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:18.590800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:18.590800Z digest=sha256:c8bfd671c43ef0d3ca2b807e15f269a9ca2ee869e6ca4f39e301805e8528fb22

Observation deba6181-a615-42f0-841b-d0ff73d2e289 · inbound

HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation cites this paper.

HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:19.588611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:19.588611Z digest=sha256:10b2ab1f46f556a760078001e1567a057a471e9021093c9d15fd1b32f8bd2f90

Observation f69c249c-b6b3-444c-90c8-e4e86784bdca · inbound

Ming-Omni: A Unified Multimodal Model for Perception and Generation cites this paper.

Ming-Omni: A Unified Multimodal Model for Perception and Generation TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:09.513337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:58:09.513337Z digest=sha256:f5346dc2481aea9925d6ed0cae085a12d5c4369d861682d0eaeb00b2313a4167

Observation fb7f0dab-7b70-4b83-b7d9-fbfcbae00a30 · inbound

Show-o2: Improved Native Unified Multimodal Models cites this paper.

Show-o2: Improved Native Unified Multimodal Models TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

Reference 89

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:51:15.542475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-12T18:51:15.428692Z digest=sha256:65c8f16b766e54fd617169dfe8694e85bd44b77ee7fad3d628205f58a87596a5

Observation ce07186b-2fa2-4449-9253-24ff0f49ae5d · inbound

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation cites this paper.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.746510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.746510Z digest=sha256:751c33b0f2ae8f6751e583353c2298dfee262305d124edd4eb8fc27baf91ef78

Observation 78609e91-2f10-4de1-8f9d-266779a24f78 · inbound

Instella-T2I: Pushing the Limits of 1D Discrete Latent Space Image Generation cites this paper.

Instella-T2I: Pushing the Limits of 1D Discrete Latent Space Image Generation TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T22:43:03.846195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:43:03.846195Z digest=sha256:b60cc40ea0a962de740923c7f0180cc92ebba3b853cbdcb3166aa483f4f826fc

Observation c1aa9569-6165-4191-ae36-9f863a158d8a · inbound

Rethinking Discrete Tokens: Treating Them as Conditions for Continuous Autoregressive Image Synthesis cites this paper.

Rethinking Discrete Tokens: Treating Them as Conditions for Continuous Autoregressive Image Synthesis TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T20:49:41.857110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:49:41.857110Z digest=sha256:cf0b185f68a17beea2b7af9ccf495d939292ae4d175d39796cbb435e20cb5f6e

Observation e77a85cc-3c45-4774-b6cd-e6648a3feeb8 · inbound

DC-AR: Efficient Masked Autoregressive Image Generation with Deep Compression Hybrid Tokenizer cites this paper.

DC-AR: Efficient Masked Autoregressive Image Generation with Deep Compression Hybrid Tokenizer TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T19:41:23.743296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:41:23.743296Z digest=sha256:728533597f3bf7592dcd9b6d44035a78790481c348a1c07fa2dc3e3fe12dd2b9

Observation 07056f24-8547-4f77-acc2-e499b83f6a3f · inbound

A Unified Low-level Foundation Model for Enhancing Pathology Image Quality cites this paper.

A Unified Low-level Foundation Model for Enhancing Pathology Image Quality TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T13:00:48.183504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:00:48.183504Z digest=sha256:76dfd6d0cbb1c1aaeda30f55b14d5855269aae156105c7714d1cd456a1f677db

Observation b6a80fd0-984a-4d3a-b633-fccfc4c216a0 · inbound

Interleaving Reasoning for Better Text-to-Image Generation cites this paper.

Interleaving Reasoning for Better Text-to-Image Generation TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:44.942757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:44.942757Z digest=sha256:a60cd2a2334508682e038fb502132231d531650bb0568cf3fb104b62175891d9

Observation d8d924e2-91f0-486d-85fd-6eaf58b666e0 · inbound

Reconstruction Alignment Improves Unified Multimodal Models cites this paper.

Reconstruction Alignment Improves Unified Multimodal Models TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-04T22:36:08.122658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T22:36:08.122658Z digest=sha256:7f148136b78fe9c27e3ae859204008d302ee856c97dbbb88c41ffa8314b47527

Observation b88afb4d-7579-4fad-ae74-0661592f29e0 · inbound

A Unified and Controllable Framework for Layered Image Generation with Visual Effects cites this paper.

A Unified and Controllable Framework for Layered Image Generation with Visual Effects TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-16T11:57:50.252224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-16T11:54:46.989748Z digest=sha256:18e12ccabf6ea81985ff072deecabb3aeb727dfd58a943f6333655c862811491

Observation 609a9b36-99f1-4287-a95c-8be0c11ae5cb · inbound

FeynmanBench: Benchmarking Multimodal LLMs on Diagrammatic Physics Reasoning cites this paper.

FeynmanBench: Benchmarking Multimodal LLMs on Diagrammatic Physics Reasoning TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T16:52:59.470222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-13T16:51:48.705876Z digest=sha256:61b53aba80c9c264b99be1402feaf685aae9403b1dfff7d19fd5b6ec61345a10

Observation f7bec7b5-3764-4502-9b7e-e1511547bb45 · inbound

FeynmanBench: Benchmarking Multimodal LLMs on Diagrammatic Physics Reasoning cites this paper.

FeynmanBench: Benchmarking Multimodal LLMs on Diagrammatic Physics Reasoning TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-13T12:10:53.720348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T12:10:53.720348Z digest=sha256:15995e8bdd7d0cf67521a2ed808d5f4c6fbc1e0f027a7d68bcc8adbf5abb44a5

Observation 211c994b-f3f4-45f2-a9ff-72e6f011aac3 · inbound

Think in Strokes, Not Pixels: Process-Driven Image Generation via Interleaved Reasoning cites this paper.

Think in Strokes, Not Pixels: Process-Driven Image Generation via Interleaved Reasoning TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:10:53.714152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T19:16:58.323955Z digest=sha256:362206db64d08213e7295ab1b7014c981005409258334fc163c3b77abd76e066

Observation 85b178fd-4ad0-47f0-974e-080507de5924 · inbound

Free Lunch for Unified Multimodal Models: Enhancing Generation via Reflective Rectification with Inherent Understanding cites this paper.

Free Lunch for Unified Multimodal Models: Enhancing Generation via Reflective Rectification with Inherent Understanding TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T14:15:28.979944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T14:15:02.723774Z digest=sha256:93b1474fdf5fea0dea92a13aab4152de0cd5592dd2c1de6b76a75194643e636c

Observation fec4d6af-7265-4a0e-a833-957939712b01 · inbound

Meta-CoT: Enhancing Granularity and Generalization in Image Editing cites this paper.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:41:19.367534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:64195565073db5a29e4f28220aa0d5bfc0d8312342914ac0251311a7cf9e8326

Observation 226722f2-0983-4fef-883d-c2374909187d · inbound

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling cites this paper.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

Reference 63

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T10:16:29.041921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:b3975c4df2b157f0eb0e99caa42765bb369867ef65e6f24bac7afea6c0bbcd93

Observation b7e1ded3-b160-45dd-bb73-6d9cfdeb56cc · inbound

What Matters for Diffusion-Friendly Latent Manifold? Prior-Aligned Autoencoders for Latent Diffusion cites this paper.

What Matters for Diffusion-Friendly Latent Manifold? Prior-Aligned Autoencoders for Latent Diffusion TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

Reference 62

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:05:57.477353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-11T01:57:24.033068Z digest=sha256:f75f7269abbfac2b0e87c88543c0a2ad21d00eb855e83e42515dade0d6b6358f

Observation ab1bc623-546d-41bc-9581-ddde19cc7ab6 · inbound

Vision Foundation Models as Generalist Tokenizers for Image Generation cites this paper.

Vision Foundation Models as Generalist Tokenizers for Image Generation TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:03:13.455436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T11:01:24.738195Z digest=sha256:4a192e99e945b175ee7e255c7c4dff43a93c809801f84481785f4a92938fe5cb

Observation 179e42ab-e52d-411d-96ca-81da6fff4ca9 · inbound

Histogram-constrained Image Generation cites this paper.

Histogram-constrained Image Generation TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

Reference 48

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T10:05:41.186922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-01T05:51:19.756153Z digest=sha256:979b1476bc1cc9c66d98be9ad1a882eb9a7ef4cd46da9f8a3c5599d4a5c7c6b0

Observation c07c4d84-88bb-4bb0-bf4c-91baa3e94539 · inbound

OSVE: One Step Video Editing with One Step Diffusion Models cites this paper.

OSVE: One Step Video Editing with One Step Diffusion Models TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-01T11:29:16.883675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:29:16.883675Z digest=sha256:49ac9c608c19a8cdf6e4f04e38a95facfaa1b9a81cf7c702bc4e1de1c9754966