Pith. sign in

Paper Citation Record · LEDGER

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation

As of 7 August 2026, this Paper Citation Record lists 94 of 94 outbound references and 0 inbound Pith citation observations for arXiv:2506.20214.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.20214 v2

Coverage vector

measured 94 of 94 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:01:49.976800Z

measured 94 of 94 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

94 of 94 outbound references displayed

  • verified exact1
  • verified fuzzy9
  • unresolved84
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 872bb111-9d23-41b6-9b8e-0ae57d180100 · outbound

This paper cites GPT-4 Technical Report.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.550995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.550995Z digest=sha256:95091ba96b341b40ffe948f7159081e46d108022a78addbef81d585504727ebf

Observation c519187e-b57e-478a-bbb5-2104d25374b8 · outbound

This paper cites Foundation models defining a new era in vision: a survey and outlook.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Foundation models defining a new era in vision: a survey and outlook.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.556467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.556467Z digest=sha256:e3600947595a56e122af93ca3d1a4f69032c1bcf61e5d41db0f7464efd7953c2

Observation ff2c862b-d8d5-4b29-89a1-d4071d76d42e · outbound

This paper cites Qwen Technical Report.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Qwen Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.561153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.561153Z digest=sha256:51c2d98f7e6d1af1d1fec12a7a931de37844782328e65e0e561c8b65c0aea89c

Observation 9e2bc2f0-20ff-4014-95f4-5d937639e659 · outbound

This paper cites Qwen2.5-VL Technical Report.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Qwen2.5-VL Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.566349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.566349Z digest=sha256:4b0d2704f1108ce5298dfc57979115da3007520d01a21372bf3fea7c47388673

Observation c6c07830-669c-4a0c-a60d-4e6f2a51b4ff · outbound

This paper cites Factorized Visual Tokenization and Generation.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Factorized Visual Tokenization and Generation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.571001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.571001Z digest=sha256:ab30ab578726af1702f00cf655273ec13d88d3278cde11fc62556fa1c1df0e46

Observation f8fb298e-2b8e-4ef4-b4a7-557b3cc17120 · outbound

This paper cites BEiT: BERT Pre-Training of Image Transformers.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation BEiT: BERT Pre-Training of Image Transformers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.575932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.575932Z digest=sha256:cd6e35a60ef6598ef881e082c1382e9ae65f18cc8dbd1f75be7a3dbd3dd9d668

Observation bd2f5d7a-2895-443e-b3ae-5a59da144700 · outbound

This paper cites Improving image generation with better captions.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Improving image generation with better captions

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.581184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.581184Z digest=sha256:7fa67dd46a8a0205424ca9cd5c8ac85eb4c4636d7e06548aa8b91acf0bd01bc2

Observation 1b4190ce-a9a4-46cb-a7da-e4fbc838671b · outbound

This paper cites Efficient-vqgan: Towards high-resolution image generation with efficient vision transformers.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Efficient-vqgan: Towards high-resolution image generation with efficient vision transformers

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.586310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.586310Z digest=sha256:45d955b47dac74a4b5664faeadac274a0dd5f7458d9e7659f6bc9eaae32303e2

Observation 816e7878-d28f-448d-a0f6-60b9fb2112c2 · outbound

This paper cites PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.590864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.590864Z digest=sha256:43967d7511917f0d131e39ee81ad1c667da9957c030bd431c52eced07270ff63

Observation 9506d091-090d-4eb3-acaf-49d4a46c20cc · outbound

This paper cites Sharegpt4v: Improving large multi-modal models with better captions.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Sharegpt4v: Improving large multi-modal models with better captions

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.596284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.596284Z digest=sha256:68b37c8467e53f82f61640421bf0ea40fdc3c284f72f9f98e144a7cdbed3ea57

Observation 64a43b10-e34a-4164-b74c-3f6c65919106 · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.600626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.600626Z digest=sha256:1e5b5a6d795f9c2d46d183c1072e4c98cc11a5ad61e06b5ec7fbf991ba3428fc

Observation 58253b53-f465-41f8-972c-0209037c50fb · outbound

This paper cites Mai: A multi-turn aggregation- iteration model for composed image retrieval.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Mai: A multi-turn aggregation- iteration model for composed image retrieval

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.605012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.605012Z digest=sha256:8157aae1e394d46714518ec490347a985465abb6dc73c62743e2ba8c32663db3

Observation 0e6275c0-886e-40ec-9c56-8ffdf8686cb0 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.609193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.609193Z digest=sha256:4759ed4801aa2095866cd13a1398542a2152abfddcdb5fcbfcfcfb9623be6870

Observation a7fe78e5-91cf-4dcf-82af-54c450071935 · outbound

This paper cites Semhitok: A unified image tokenizer via semantic-guided hierarchical codebook for multimodal understanding and generation.arXiv preprint arXiv:2503.06764, 2025.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Semhitok: A unified image tokenizer via semantic-guided hierarchical codebook for multimodal understanding and generation.arXiv preprint arXiv:2503.06764, 2025

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.613792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.613792Z digest=sha256:beabb2aa3a1adbea00a62874fa0cf7d2ce375227e2fb5e889aa9c12607116c90

Observation 5b1cc51a-8109-4340-8a18-1e8590487aa4 · outbound

This paper cites ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.618223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.618223Z digest=sha256:c2de173dd5cf306169d06a480072a01f91a7ae38c16f4815d62f0707eae3d6ba

Observation df0a9eeb-3faf-47ed-84cb-1e5802130594 · outbound

This paper cites DreamLLM: Synergistic Multimodal Comprehension and Creation.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation DreamLLM: Synergistic Multimodal Comprehension and Creation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.622707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.622707Z digest=sha256:4e3876c0efd1897c8721c1041eb870074140926ed0989368687b4d326c85cdd5

Observation 8ed18f1b-c6ee-44bf-8092-f2139cc88e44 · outbound

This paper cites Taming transformers for high-resolution image synthesis.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Taming transformers for high-resolution image synthesis

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.627390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.627390Z digest=sha256:1b2cbd34c842c5ac05293a5c3812990a189c4068416115103d21c16fb434e1f9

Observation ec0c0661-7393-4f8f-8e30-722e711d7eea · outbound

This paper cites Eva-02: A visual representation for neon genesis.Image and Vision Computing, 149:105171, 2024.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Eva-02: A visual representation for neon genesis.Image and Vision Computing, 149:105171, 2024

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.631713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.631713Z digest=sha256:5e9f5584cc05e797002c4ae4474112fc3bfc99e34e3fa41e560f96b62edd1160

Observation f3ff8017-8113-478e-a766-99933325d007 · outbound

This paper cites Eva: Exploring the limits of masked visual representation learning at scale.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Eva: Exploring the limits of masked visual representation learning at scale

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.636345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.636345Z digest=sha256:8f07a34fdfa3601d8d7bd916fe94106fc344523e91414b9158c41e5e9df5a908

Observation a9370ca3-f9b0-4ecd-9790-ca3683e7cd07 · outbound

This paper cites SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.640623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.640623Z digest=sha256:dd2c885d7c007dfedf6594e1f16e11828b476980d09b14776632de5d918f75dc

Observation d21b8d60-4825-4754-93c6-e0b072fb4956 · outbound

This paper cites Geneval: An object-focused framework for evaluating text-to-image alignment.Advances in Neural Information Processing Systems, 36:52132–52152, 2023.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Geneval: An object-focused framework for evaluating text-to-image alignment.Advances in Neural Information Processing Systems, 36:52132–52152, 2023

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.645064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.645064Z digest=sha256:933b14f3eeb4b5dbf82aa134287b542c7a5e7a01b1567b547240d1dcbda3eacd

Observation c5dfdd9f-140d-4262-92ef-2e1fb81f2bbf · outbound

This paper cites A survey on self-supervised learning: Algorithms, applications, and future trends.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation A survey on self-supervised learning: Algorithms, applications, and future trends.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.649556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.649556Z digest=sha256:ef705c8398b8eb8e37a898cebf6ad897609d5390436b216d70de383a5dfa0d7c

Observation 07d11dca-9be3-4a51-9a81-ff3f8b66c356 · outbound

This paper cites ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.654183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.654183Z digest=sha256:5ca51b9c688bcb45bd85078ada1a735823878d3ba3b9970c6a6216ffa9ccc8ad

Observation 4d2afe1e-87aa-4b55-9add-f11e96ef60c8 · outbound

This paper cites ILLUME+: Illuminating Unified MLLM with Dual Visual Tokenization and Diffusion Refinement.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation ILLUME+: Illuminating Unified MLLM with Dual Visual Tokenization and Diffusion Refinement

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.658476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.658476Z digest=sha256:e05833625040859c8ef2c610a84a8cc03e35ce15051458b21e86c54184a25ebf

Observation 808afccd-3930-4e07-9246-7950dbbabb36 · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.662775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.662775Z digest=sha256:c234c12efaca1790dbad958b8a77c4a5484deece45ee1794f07d1b75641f7730

Observation 185a68ec-e17a-4264-a362-0c073d654f6d · outbound

This paper cites UniToken: Harmonizing Multimodal Understanding and Generation through Unified Visual Encoding.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation UniToken: Harmonizing Multimodal Understanding and Generation through Unified Visual Encoding

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.667271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.667271Z digest=sha256:05e6bdf2e87c807ba432686ca9dc503d587bc6bf7e19196f3b288ac0cd338524

Observation 5e474d8d-49d9-4e70-baca-5c8574236730 · outbound

This paper cites A diagram is worth a dozen images.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation A diagram is worth a dozen images

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.671702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.671702Z digest=sha256:1771c8a2ba719ff62b3254c3bf9cd48bf599d2fbb536e057192b2237a05dff1e

Observation d1a0a8bb-4238-4767-ab93-68179d7ffadf · outbound

This paper cites an unresolved cited work.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.676112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.676112Z digest=sha256:fa3d8ba6f721102552704bae9f1c6795fda1c315e1e8582664117a8248af432b

Observation e2dad82a-0490-4401-a0fa-dd809acb4476 · outbound

This paper cites Autoregressive image generation using residual quantization.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Autoregressive image generation using residual quantization

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.680540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.680540Z digest=sha256:88efacb0f796563b1b1b665272950aa59e58f4abac4994630f12227c1bc95a36

Observation 9a5a63a1-7d84-4fa7-9c57-fe42d513137f · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation LLaVA-OneVision: Easy Visual Task Transfer

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.684769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.684769Z digest=sha256:c60bbcf2baef4aae981caf2a8252a6e85e45c66d67bbeee647600c091bc4eebc

Observation e6328eb0-7a09-4e34-b5f5-ade67c75f138 · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.688963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.688963Z digest=sha256:6a1940029d9c3b9b5debb56e2bc3cce14476ccf6a8bea284ad3cb28cfc5c29e8

Observation dfb0530e-9d56-4664-9f36-9be5475d79c5 · outbound

This paper cites A survey of multimodel large language models.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation A survey of multimodel large language models

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:51.849440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:01:49.693312Z digest=sha256:649b1e745659dea0964b1f87859eb6d3c6d02c19498227e8e11e2f632658bc97

Observation 2208cd6b-3b7a-431b-924a-1ad2e4eedd7d · outbound

This paper cites Toklip: Marry visual tokens to clip for multimodal comprehension and generation, 2025.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Toklip: Marry visual tokens to clip for multimodal comprehension and generation, 2025

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:51.831628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:01:49.697319Z digest=sha256:2efde61d94b8485ace3487996865ba163362cd1c09b39cfe44fad45d09e45de5

Observation cb6bdd6f-6d7e-4939-8f90-cacfc8886da7 · outbound

This paper cites World Model on Million-Length Video And Language With Blockwise RingAttention.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation World Model on Million-Length Video And Language With Blockwise RingAttention

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.702024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.702024Z digest=sha256:b75e6d52b5d03b1380714e889815b6ea2137a73eed47428221539d876333dd2c

Observation 69d9c7a3-63e9-4f07-b34c-8b154f8d7278 · outbound

This paper cites Improved baselines with visual instruction tuning.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Improved baselines with visual instruction tuning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.706281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.706281Z digest=sha256:e4cbe5c4c625307ec17dbb7e9fd7556aa4bb561a55154e0f2833e1fb6db6439d

Observation 801b98e7-a0c7-4180-875f-475dbc187d5f · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.710584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.710584Z digest=sha256:cef2427e20c616ed2e8a43222cd01600f25630c380dd4d8631ef23a916bff333

Observation b7310155-5a00-4fd4-bf02-b6a5ee3a3858 · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? InEuropean conference on computer vision, pages 216–233.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Mmbench: Is your multi-modal model an all-around player? InEuropean conference on computer vision, pages 216–233

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.715276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.715276Z digest=sha256:6144800ce1b549ebd8b324a4874235cf9487de72cc7f2ebae5f13e7206c9325e

Observation c07dfd61-d9a0-43b6-aea3-adad6cb73cf1 · outbound

This paper cites Decoupled Weight Decay Regularization.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Decoupled Weight Decay Regularization

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.719739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.719739Z digest=sha256:770121bd32b6032ac66e03ad8fb1ca1e8f0641937fb8142f524a2f0a285db206

Observation b9e15661-e791-481c-8847-73d935275cfd · outbound

This paper cites Ovis: Structural Embedding Alignment for Multimodal Large Language Model.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.723984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.723984Z digest=sha256:40103dd2f8e121147132d10884417d511f8525ce84380425ca1ef526f0e4010b

Observation 0297c9ca-c531-4fde-bd37-81888fd2e27f · outbound

This paper cites Unitok: A unified tokenizer for visual generation and understanding.arXiv preprint arXiv:2502.20321, 2025.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Unitok: A unified tokenizer for visual generation and understanding.arXiv preprint arXiv:2502.20321, 2025

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.728398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.728398Z digest=sha256:42dc18243ee59c99af65f2a1123afd204b120b3094634821bc9a97103df4e096

Observation b82b6d9c-43e2-4f05-9534-8aad015c5f2a · outbound

This paper cites JanusFlow: Harmonizing Autoregression and Rectified Flow for Unified Multimodal Understanding and Generation.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation JanusFlow: Harmonizing Autoregression and Rectified Flow for Unified Multimodal Understanding and Generation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.732504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.732504Z digest=sha256:f4f1340c0954e816768474da455ba7e957a31bd7424f68b057d1b612c8f220f3

Observation 9393d3d7-75c3-4e01-895d-a6536e095847 · outbound

This paper cites BEiT v2: Masked Image Modeling with Vector-Quantized Visual Tokenizers.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation BEiT v2: Masked Image Modeling with Vector-Quantized Visual Tokenizers

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.737348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.737348Z digest=sha256:649bc015ae99e951f0f581524d8e280aee06e8119d5a491cf4a64275a4853c1b

Observation 0f870014-1798-4722-a736-f277b0887bea · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.742089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.742089Z digest=sha256:f78c9453d3c06cf61cd38c318ebf21d9b78da61277560e3253e1bd68e1e58423

Observation ce07186b-2fa2-4449-9253-24ff0f49ae5d · outbound

This paper cites TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.746510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.746510Z digest=sha256:ace59a9c0236c869ae71bcdd6e589becb617b1535542e38e022779e2b64abb9f

Observation 25f9e83f-3624-4516-9e8a-ef6009465aa1 · outbound

This paper cites Beyond Next-Token: Next-X Prediction for Autoregressive Visual Generation.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Beyond Next-Token: Next-X Prediction for Autoregressive Visual Generation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.751169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.751169Z digest=sha256:3d8f6f615a3c020c49bfc8e12ed1bed044a56c272b705e4233a74c2c7769b37d

Observation bcf6da21-bf3c-4997-9986-9c8c1ee82db2 · outbound

This paper cites High- resolution image synthesis with latent diffusion models.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation High- resolution image synthesis with latent diffusion models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.755625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.755625Z digest=sha256:6651fcb814acdc21dbaeef857ab080b8006ef54059978caeb35012dc31f574f0

Observation c99e1186-de25-478d-9674-808f355f0b44 · outbound

This paper cites Scalable Image Tokenization with Index Backpropagation Quantization.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Scalable Image Tokenization with Index Backpropagation Quantization

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.760298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.760298Z digest=sha256:4378f57fd6971c1f3eff1b19fa1bc6149313163bf5a122c7e8e66d3bc3270a2a

Observation 741dca58-6de6-4114-b0c6-f8429c8d91b6 · outbound

This paper cites Towards vqa models that can read.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Towards vqa models that can read

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.765263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.765263Z digest=sha256:5ed2c46b1fbe82d2e10f8366b609674bec4bb5476625c9c3a6f664650dad36f5

Observation c0101e35-3818-4596-b9e2-af2b176910a6 · outbound

This paper cites DualToken: Towards Unifying Visual Understanding and Generation with Dual Visual Vocabularies.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation DualToken: Towards Unifying Visual Understanding and Generation with Dual Visual Vocabularies

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.769637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.769637Z digest=sha256:255071a0f2a995d9c84fbb5d1ed1262a338fd646ac8f3431dfdc70b0a12a7802

Observation 9bea6b54-9227-4f47-b7d0-d89077778d14 · outbound

This paper cites Journeydb: A benchmark for generative image understanding.Advances in neural information processing systems, 36:49659–49678, 2023.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Journeydb: A benchmark for generative image understanding.Advances in neural information processing systems, 36:49659–49678, 2023

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.774524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.774524Z digest=sha256:82db323fe78ccdaeca887ca7f8b38207207093b2b71e886e25d0bf9f0926c81c

Observation 9f449a2c-e8a0-495d-93cf-5922ea91ff05 · outbound

This paper cites Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.779091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.779091Z digest=sha256:99cce6573e86ced418b459f80349d7ced8c99cc7306f6b269a93162746c993ac

Observation 6425bd17-4f4f-4304-9a6a-3c9b8059245d · outbound

This paper cites Generative multimodal models are in-context learners.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Generative multimodal models are in-context learners

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.784005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.784005Z digest=sha256:51b5207d5a9b62583ab0fa9b72ad3edbe1ec7ff52971493b175d3889b768f24c

Observation 5ce7e1ae-a651-46ff-b5e7-cbb259228f94 · outbound

This paper cites Emu: Generative Pretraining in Multimodality.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Emu: Generative Pretraining in Multimodality

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.788445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.788445Z digest=sha256:2b654e6b181ddf25ab14f29a7870bd661f8d2cf583d787d1932743ce2037e22c

Observation ef4f571a-8c4f-4036-bd58-ec9d840575eb · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.793100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.793100Z digest=sha256:7bd93832df48b4f2c99a544ee9c56bf34bfb0326de6809652df23049d48b228e

Observation 9780d705-0bc5-4c10-9ff9-0cdb2158c955 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Gemini: A Family of Highly Capable Multimodal Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.797577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.797577Z digest=sha256:4672688ad7f97fe0b136201b44662542599ec586c23fe135c0ece327091d8339

Observation 61084d5c-aaa5-43f4-af9d-aa1f029122d5 · outbound

This paper cites Visual autoregressive modeling: Scalable image generation via next-scale prediction.Advances in neural information processing systems, 37:84839–84865, 2024.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Visual autoregressive modeling: Scalable image generation via next-scale prediction.Advances in neural information processing systems, 37:84839–84865, 2024

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.801688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.801688Z digest=sha256:755fdc86639cbc83bcedd1ec6dd507cd27292bbbfbf549b50a39094890f26311

Observation 7a2b6382-157a-4e49-9676-8fe6d40a94d6 · outbound

This paper cites MetaMorph: Multimodal Understanding and Generation via Instruction Tuning.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation MetaMorph: Multimodal Understanding and Generation via Instruction Tuning

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.805911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.805911Z digest=sha256:6f10cde3167f48045650000e5ff7a56635378dfa8ec45f236251ff5c63550a7c

Observation 1e630c62-7010-4cd4-90b4-3c0407fe3f0f · outbound

This paper cites Neural discrete representation learning.Advances in neural information processing systems, 30, 2017.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Neural discrete representation learning.Advances in neural information processing systems, 30, 2017

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.810741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.810741Z digest=sha256:279ec9ca13ce91065c0957dc5261030b3b9c3211350701a67a02e951402b3cd0

Observation 5a97a6a8-26bc-4f56-9a80-f7f0ead67f4b · outbound

This paper cites Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.815176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.815176Z digest=sha256:12341bbe7e1efb8316c3758f46e2d6c7c9fef7476dfcb0ec6ad0e0670d15b2c4

Observation 8fe5a1de-7e37-4c94-9eb1-8fc0804bb377 · outbound

This paper cites ILLUME: Illuminating Your LLMs to See, Draw, and Self-Enhance.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation ILLUME: Illuminating Your LLMs to See, Draw, and Self-Enhance

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.819756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.819756Z digest=sha256:2b3c2df029dcc02cac9e790eaa29577fac53dc1460a5bf0b1f8c9490ae3a5768

Observation 5e13fc24-4806-4e23-bae5-cebaa0c79874 · outbound

This paper cites Omnitok- enizer: A joint image-video tokenizer for visual generation.Advances in Neural Information Processing Systems, 37:28281–28295, 2024.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Omnitok- enizer: A joint image-video tokenizer for visual generation.Advances in Neural Information Processing Systems, 37:28281–28295, 2024

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:51.690847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:01:49.824085Z digest=sha256:9a39b95c98002c5beafab25bb320501c133a7777718b885f180e7fb088972e07

Observation 5a6f2efd-9e75-46b7-9c18-974d95f567e1 · outbound

This paper cites Image under- standing makes for a good tokenizer for image generation.Advances in Neural Information Processing Systems, 37:31015–31035, 2024.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Image under- standing makes for a good tokenizer for image generation.Advances in Neural Information Processing Systems, 37:31015–31035, 2024

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:51.670739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:01:49.828530Z digest=sha256:6864b6e2ac7f0f6f902474e3a93c9e20631beabc9ba48b0ac4b308ef28ccfb70

Observation 7a8093d1-edb0-4dca-b5f1-fc8c28582e59 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.832665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.832665Z digest=sha256:f6b6029ede34cdc212da91e9d70a2b47a10e0f8b7b9f54dc721e807f0b4adb92

Observation 638346e8-1ae3-4000-b244-7d8e29d0f2e1 · outbound

This paper cites Image as a foreign lan- guage: Beit pretraining for vision and vision-language tasks.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Image as a foreign lan- guage: Beit pretraining for vision and vision-language tasks

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.837077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.837077Z digest=sha256:daa30568e05c4d3019ae44f9e48d77c55a8c938aadf50786cc387d24a3d638cb

Observation fb6c6446-7d35-42f3-9ac2-a12b09e0ca80 · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Emu3: Next-Token Prediction is All You Need

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.841242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.841242Z digest=sha256:6a132c2079c451faeb3e290334a2c47934982dc10614bf1bb5055177f977a429

Observation 5d49eb60-9db7-43d2-894d-a55ed865ddb7 · outbound

This paper cites Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.845919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.845919Z digest=sha256:001ba58784ef13dd0efe309ffa63870309a89c19c790955245f882ec83c35922

Observation fa027142-f6a4-4ae2-bfcd-7fbd7e199ab2 · outbound

This paper cites Liquid: Language Models are Scalable and Unified Multi-modal Generators.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Liquid: Language Models are Scalable and Unified Multi-modal Generators

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.850509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.850509Z digest=sha256:a602f56e19bd61c604b6b4e622c8ed63d4532a42e41c2b9916647376f44006dc

Observation 4640a3a0-953a-4ead-be75-c1d4759418dc · outbound

This paper cites Next-gpt: Any-to-any multimodal llm.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Next-gpt: Any-to-any multimodal llm

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:51.636941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:01:49.854970Z digest=sha256:f8e78c4ee9d9fc50dcfd1b5c5eb667aeaeb86f604b2ff94adfecfdd6590cd91e

Observation 58a09def-c607-40c0-b40f-9ea28861d1b1 · outbound

This paper cites Harmonizing Visual Representations for Unified Multimodal Understanding and Generation.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Harmonizing Visual Representations for Unified Multimodal Understanding and Generation

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.859294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.859294Z digest=sha256:296288d7df932c88017a5af6dff8502ba0283e53f1e3b6ad58af242bf9bfe674

Observation 94a77910-e224-46a6-a252-85d2bdc4603f · outbound

This paper cites VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.863781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.863781Z digest=sha256:e58c8b6073e6c16671ef8696f2fa4c2a36420c472089f240b69e7b6e8962d90c

Observation f5143d83-63be-4c9c-be19-75acb5c5b9d2 · outbound

This paper cites Grok-1.5 vision preview, 6 2024.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Grok-1.5 vision preview, 6 2024

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:51.616407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:01:49.868739Z digest=sha256:c38c41b63e794e7b941d7ced4e61c872f8b8ec205553904199275b44b3327091

Observation 0c1d8e90-0ad1-47b6-9d5e-1fdb31e97b4f · outbound

This paper cites OmniGen: Unified Image Generation.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation OmniGen: Unified Image Generation

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.873348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.873348Z digest=sha256:8141590ed05b94dd233bb4953b2b92169706c507932cefdae996c14ab3c941a2

Observation f267fa33-2a43-494f-843e-1cc236f425db · outbound

This paper cites Show-o: One Single Transformer to Unify Multimodal Understanding and Generation.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.878313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.878313Z digest=sha256:a152a8d711f15e932606c33f19842ffe5dfc7f337a4b7e65b0caeaaf975d5a12

Observation eca16e4f-e042-4a5e-9470-7a5a5991c70b · outbound

This paper cites MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.883147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.883147Z digest=sha256:8a8bc5fec2f77a094ba289e475b54b16df6261e3906d6a6b1c56093504b8ef24

Observation 1e4c5d69-034a-48fb-8da2-858ccd1c4813 · outbound

This paper cites Efficientsam: Leveraged masked image pretraining for efficient segment anything.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Efficientsam: Leveraged masked image pretraining for efficient segment anything

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:51.595479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:01:49.887716Z digest=sha256:d11a695c5995832cc2e38ea6a7e9a8830391f4764f8667e7dd24c8f53eaeec0c

Observation 1095fca2-3ae7-4175-a3c6-136135bc4670 · outbound

This paper cites Qwen2.5 Technical Report.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Qwen2.5 Technical Report

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.891862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.891862Z digest=sha256:4ae3533a38f590b614d8a4aebc5b91d81e09c4fb8f9ef4261648d39ef6c24577

Observation ec5bb5e2-5f5d-428a-98da-e6ec69ca4aac · outbound

This paper cites Vector-quantized Image Modeling with Improved VQGAN.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Vector-quantized Image Modeling with Improved VQGAN

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.896781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.896781Z digest=sha256:da2b6ee43bc0144c013634dec7c1ba3fcc958332a6106e605f90b475f4c02959

Observation 29635d4e-8db5-4ee0-ab66-e5d08ade24e7 · outbound

This paper cites Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.901168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.901168Z digest=sha256:a73ae69c797c96d046c7ce6cb9e55f2add98a77bf529245b1d43fc9f204cfa80

Observation a7f74818-5d00-4ea5-92df-f2b5ab82163e · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.905672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.905672Z digest=sha256:c4b4d20bb58df75a977ef26a0f19757b2166ee75dec8ec4d9b7517a5dbf9173e

Observation cf029f21-64fd-48ee-a43b-14918d731c36 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.910134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.910134Z digest=sha256:d74e2c3be7d3029c5563243afb401c52bbf250e849a1a89a42650f7d7a5b6a4f

Observation cf06dac7-6d75-4178-848f-b79dcaa58976 · outbound

This paper cites Sigmoid loss for language image pre-training.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Sigmoid loss for language image pre-training

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.914438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.914438Z digest=sha256:b5b6517fd7bfe646778c84ddc94b46fba0a6408e02060b3574173d2f41ddcde9

Observation 7ffabdac-a96f-46a5-84fe-6fb89690a743 · outbound

This paper cites Token dynamics: Towards efficient and dynamic video token representation for video large language models.arXiv preprint arXiv:2503.16980, 2025.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Token dynamics: Towards efficient and dynamic video token representation for video large language models.arXiv preprint arXiv:2503.16980, 2025

Reference 82

Resolution
verified exact
raw_fallback, observed 2026-08-06T23:01:50.399667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:01:49.919010Z digest=sha256:d9d3cff4b7f2bdc151399b864f596e29baa53ca8851e0d26b0d4645e213a77f1

Observation 5b8708d8-348c-4e4b-9285-0ba2dd27d76a · outbound

This paper cites Unified multimodal understanding and generation models: Advances, challenges, and opportunities.arXiv preprint arXiv:2505.02567, 2025.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Unified multimodal understanding and generation models: Advances, challenges, and opportunities.arXiv preprint arXiv:2505.02567, 2025

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.923951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.923951Z digest=sha256:a05992c424824a1aa38aa69812f013196e3ab04f8afde5937f4b78e1263a657f

Observation cf687e1a-b5f5-4563-a753-5427cbecb4a8 · outbound

This paper cites QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.928203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.928203Z digest=sha256:16a9d951d8d0897fcafbf72619beebf1ca6e72387b6b973cedbc46a7c2cc4862

Observation f4809218-b996-4377-b2ec-af22cdef42a5 · outbound

This paper cites Online clustered codebook.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Online clustered codebook

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.932726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.932726Z digest=sha256:ef3bc3e532258466fd5012cae20e1e152d53881fafdc5b37a7afcc6adb82a955

Observation 84e2bf29-c556-4028-856a-f85857f84d14 · outbound

This paper cites Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.937018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.937018Z digest=sha256:73c3818b14e44f03055b611f22b6777287f0bd8fbcd15006148375e3f4c91444

Observation 6459ee8c-5f34-43d2-9b3e-5fba6bd2d812 · outbound

This paper cites Scaling the Codebook Size of VQGAN to 100,000 with a Utilization Rate of 99%.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Scaling the Codebook Size of VQGAN to 100,000 with a Utilization Rate of 99%

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.941622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.941622Z digest=sha256:ae249a92e51a2fcdd4ab49f02ff8d8cf3b6186eb2b6736e8a4e68aa0b6f2a91d

Observation caa43789-6d45-46a6-914e-9f32bf63ae11 · outbound

This paper cites Addressing representation collapse in vector quantized models with one linear layer.arXiv preprint arXiv:2411.02038, 2024.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Addressing representation collapse in vector quantized models with one linear layer.arXiv preprint arXiv:2411.02038, 2024

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.946448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.946448Z digest=sha256:4ade33fec630843eee6590a29ed02d3e5221bd3bea826f3e4749e3833e8734f6

Observation 365b6c50-2d1c-47e0-a6eb-2cab4993cf8b · outbound

This paper cites OmniMamba: Efficient and Unified Multimodal Understanding and Generation via State Space Models.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation OmniMamba: Efficient and Unified Multimodal Understanding and Generation via State Space Models

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.951366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.951366Z digest=sha256:974a87f985eb31c85ef09080a1f3fb4c41399928df9d28ba5831c54f16f44f06

Observation bd75435d-6cf2-420b-88fc-211b99186a4d · outbound

This paper cites Let E={e 1, ...,eN } ⊂Rd denote visual embeddings sampled from a distribution p(e), and let C={c 1, ...,cK} ⊂Rd be a codebook with K discrete centroids.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Let E={e 1, ...,eN } ⊂Rd denote visual embeddings sampled from a distribution p(e), and let C={c 1, ...,cK} ⊂Rd be a codebook with K discrete centroids

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:51.534554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:01:49.957273Z digest=sha256:26495cc61e5a25b36979f208720eed3aa3632daf9a62dfe314f8c9ef7f1a77b8

Observation 5289fcf7-54c8-430d-9dd6-f00fb34f66e5 · outbound

This paper cites an unresolved cited work.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Unresolved cited work

Reference 91

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:01:51.513926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:01:49.961746Z digest=sha256:dcf068e715d97e68e39bd3b9bc71b23fd35e17121e2b83f787e8d1a1bf412c15

Observation 6c9382c1-2830-4073-8eae-1c713deb4564 · outbound

This paper cites Let y be a semantic target label (e.g., object class, scene type), and let v=Q(e) be the discrete token assigned to embeddinge.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Let y be a semantic target label (e.g., object class, scene type), and let v=Q(e) be the discrete token assigned to embeddinge

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:51.489877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:01:49.966452Z digest=sha256:ce467df2caf031100ffa56d2edc33ef64f740add8b5ba295d7ce86d7cd71a958

Observation a003bd6f-da76-4860-be9c-742ba2295ec1 · outbound

This paper cites an unresolved cited work.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Unresolved cited work

Reference 93

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:01:51.469848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:01:49.971454Z digest=sha256:7e8e86ef5a5a8e5121c571c6c77e408959bb4ac7db536bb95c63b1cd739b6eb4

Observation af3ee04e-8a70-4148-bbd3-d27b89482896 · outbound

This paper cites an unresolved cited work.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Unresolved cited work

Reference 94

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:01:51.448653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:01:49.976800Z digest=sha256:383505e32f4cc732bf429a6b33ac5a4c2ebcbffe6bdcb9691f7c3b5b4eebac7b

Pith citing papers

No inbound Pith citation observations are available.