Pith. sign in

Paper Citation Record · LEDGER

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation

As of 7 August 2026, this Paper Citation Record lists 94 of 94 outbound references and 0 inbound Pith citation observations for arXiv:2506.20214.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.20214 v2

Coverage vector

measured 94 of 94 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:01:49.976800Z

measured 94 of 94 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

94 of 94 outbound references displayed

  • verified exact1
  • verified fuzzy9
  • unresolved84
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 872bb111-9d23-41b6-9b8e-0ae57d180100 · outbound

This paper cites GPT-4 Technical Report.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.550995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.550995Z digest=sha256:9755fd2b8e81730e6c8cc5f910956fde88f9a30666c22832457c56f5aab2f933

Observation c519187e-b57e-478a-bbb5-2104d25374b8 · outbound

This paper cites Foundation models defining a new era in vision: a survey and outlook.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Foundation models defining a new era in vision: a survey and outlook.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.556467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.556467Z digest=sha256:333fb879f38806e19a2aeeb335959ac9ab16a9a72f446bd60fb5f8f5dfaf105e

Observation ff2c862b-d8d5-4b29-89a1-d4071d76d42e · outbound

This paper cites Qwen Technical Report.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Qwen Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.561153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.561153Z digest=sha256:d54f2832ae2a5abd1e320b3b843e96f7c5b8170fa83dff88076be607ddb1274c

Observation 9e2bc2f0-20ff-4014-95f4-5d937639e659 · outbound

This paper cites Qwen2.5-VL Technical Report.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Qwen2.5-VL Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.566349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.566349Z digest=sha256:0b017da6f6abef57c8852dc39756e48663a0af33529d9f21bad90a12c5b39be9

Observation c6c07830-669c-4a0c-a60d-4e6f2a51b4ff · outbound

This paper cites Factorized Visual Tokenization and Generation.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Factorized Visual Tokenization and Generation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.571001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.571001Z digest=sha256:7741adbcb98af78c6e41b127d64f6761e35ddb6bfdee613e98aa48d13b403462

Observation f8fb298e-2b8e-4ef4-b4a7-557b3cc17120 · outbound

This paper cites BEiT: BERT Pre-Training of Image Transformers.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation BEiT: BERT Pre-Training of Image Transformers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.575932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.575932Z digest=sha256:b8fdd0978472124dad5e2b71284811774a98952fe33e6b4a02eacbc83098b201

Observation bd2f5d7a-2895-443e-b3ae-5a59da144700 · outbound

This paper cites Improving image generation with better captions.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Improving image generation with better captions

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.581184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.581184Z digest=sha256:1e8478550c558a653c83d01edc84b365d4ec1bb935e329c11bb4fe60c587871e

Observation 1b4190ce-a9a4-46cb-a7da-e4fbc838671b · outbound

This paper cites Efficient-vqgan: Towards high-resolution image generation with efficient vision transformers.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Efficient-vqgan: Towards high-resolution image generation with efficient vision transformers

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.586310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.586310Z digest=sha256:0ab73863598c8de077ece2bc1b3930931253bde914875cbd641305e0cadd6846

Observation 816e7878-d28f-448d-a0f6-60b9fb2112c2 · outbound

This paper cites PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.590864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.590864Z digest=sha256:f31572351327a4d038f8f94ebea7548a71371935a4f59551953e1b257057c0af

Observation 9506d091-090d-4eb3-acaf-49d4a46c20cc · outbound

This paper cites Sharegpt4v: Improving large multi-modal models with better captions.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Sharegpt4v: Improving large multi-modal models with better captions

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.596284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.596284Z digest=sha256:09504e167abac53abdb45fb7666a49a1484cbd2da70f243cf83afbde94dc6a1c

Observation 64a43b10-e34a-4164-b74c-3f6c65919106 · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.600626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.600626Z digest=sha256:8ae17c63460070e4fca8d21fc12d3c81e1cd78f5ff3ff38b199a34961a68d5b9

Observation 58253b53-f465-41f8-972c-0209037c50fb · outbound

This paper cites Mai: A multi-turn aggregation- iteration model for composed image retrieval.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Mai: A multi-turn aggregation- iteration model for composed image retrieval

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.605012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.605012Z digest=sha256:6eca46200eafeae68cf84eb0135779da31625110d5bf40c34f30fa1b8284638b

Observation 0e6275c0-886e-40ec-9c56-8ffdf8686cb0 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.609193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.609193Z digest=sha256:903702b08b33a2c6b1ef575f6599721ff676d91d0ee48b7033f8790ac2ed7f42

Observation a7fe78e5-91cf-4dcf-82af-54c450071935 · outbound

This paper cites Semhitok: A unified image tokenizer via semantic-guided hierarchical codebook for multimodal understanding and generation.arXiv preprint arXiv:2503.06764, 2025.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Semhitok: A unified image tokenizer via semantic-guided hierarchical codebook for multimodal understanding and generation.arXiv preprint arXiv:2503.06764, 2025

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.613792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.613792Z digest=sha256:5285acf5e0e36bd78ceece274a2687d9e0dae5f28654782b0036fed686790523

Observation 5b1cc51a-8109-4340-8a18-1e8590487aa4 · outbound

This paper cites ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.618223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.618223Z digest=sha256:f151ba20228ccd9666f2139bc569cf4ed011f509cc41ab10433e5f08aca91328

Observation df0a9eeb-3faf-47ed-84cb-1e5802130594 · outbound

This paper cites DreamLLM: Synergistic Multimodal Comprehension and Creation.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation DreamLLM: Synergistic Multimodal Comprehension and Creation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.622707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.622707Z digest=sha256:744c78fd66fc69f4a9b2e437827c14d700ab16e28fc14eeb6ef5d87ad6ef56ad

Observation 8ed18f1b-c6ee-44bf-8092-f2139cc88e44 · outbound

This paper cites Taming transformers for high-resolution image synthesis.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Taming transformers for high-resolution image synthesis

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.627390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.627390Z digest=sha256:cc14137f2b0932c5e5df23ae5ee13fb0c3821b55e1ef908b309488a4f09306aa

Observation ec0c0661-7393-4f8f-8e30-722e711d7eea · outbound

This paper cites Eva-02: A visual representation for neon genesis.Image and Vision Computing, 149:105171, 2024.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Eva-02: A visual representation for neon genesis.Image and Vision Computing, 149:105171, 2024

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.631713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.631713Z digest=sha256:c79340637c49e37c57a08889cd0b09827613313fff19f7eb0dc8d94e422e297a

Observation f3ff8017-8113-478e-a766-99933325d007 · outbound

This paper cites Eva: Exploring the limits of masked visual representation learning at scale.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Eva: Exploring the limits of masked visual representation learning at scale

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.636345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.636345Z digest=sha256:567f600ba6d6f7bce3f844686a874a8d08e361c42c6a10805330a3749f62be49

Observation a9370ca3-f9b0-4ecd-9790-ca3683e7cd07 · outbound

This paper cites SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.640623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.640623Z digest=sha256:4ba8c246e0d2b34ecdf0806624b28e60b14551058ca98a7b8dcae7e310b9f34e

Observation d21b8d60-4825-4754-93c6-e0b072fb4956 · outbound

This paper cites Geneval: An object-focused framework for evaluating text-to-image alignment.Advances in Neural Information Processing Systems, 36:52132–52152, 2023.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Geneval: An object-focused framework for evaluating text-to-image alignment.Advances in Neural Information Processing Systems, 36:52132–52152, 2023

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.645064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.645064Z digest=sha256:ab0008b7518afcce6c6de50a1f1dc2f9b6c6bb4d67eb4aa6e9345210ab0e3bd6

Observation c5dfdd9f-140d-4262-92ef-2e1fb81f2bbf · outbound

This paper cites A survey on self-supervised learning: Algorithms, applications, and future trends.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation A survey on self-supervised learning: Algorithms, applications, and future trends.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.649556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.649556Z digest=sha256:4adfaeb195c8e2b5800020179c83f1151b2eaf67cf2b72984661349700b05134

Observation 07d11dca-9be3-4a51-9a81-ff3f8b66c356 · outbound

This paper cites ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.654183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.654183Z digest=sha256:59528a98b2dbb24c7e54101dd7d9cae9069a65e23b9a8627bd24c1aa630a6b87

Observation 4d2afe1e-87aa-4b55-9add-f11e96ef60c8 · outbound

This paper cites ILLUME+: Illuminating Unified MLLM with Dual Visual Tokenization and Diffusion Refinement.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation ILLUME+: Illuminating Unified MLLM with Dual Visual Tokenization and Diffusion Refinement

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.658476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.658476Z digest=sha256:b4fb3046ed48aefb8f1ca15ced6acb43dc90cd084faf955101263dea6260ece9

Observation 808afccd-3930-4e07-9246-7950dbbabb36 · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.662775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.662775Z digest=sha256:9443d40e45e58f36a1b3f88929164bf72278cb7111c180c152352261cbbeb443

Observation 185a68ec-e17a-4264-a362-0c073d654f6d · outbound

This paper cites UniToken: Harmonizing Multimodal Understanding and Generation through Unified Visual Encoding.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation UniToken: Harmonizing Multimodal Understanding and Generation through Unified Visual Encoding

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.667271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.667271Z digest=sha256:fc587cf7c064c5a6da297f15724dda471dedb166af9ded170fe114e297cb7ee5

Observation 5e474d8d-49d9-4e70-baca-5c8574236730 · outbound

This paper cites A diagram is worth a dozen images.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation A diagram is worth a dozen images

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.671702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.671702Z digest=sha256:a67514e95d11aab2d3c98b635ad4b72f2e2ff53649e71b0f29f515bd17a277fc

Observation d1a0a8bb-4238-4767-ab93-68179d7ffadf · outbound

This paper cites an unresolved cited work.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.676112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.676112Z digest=sha256:47804e00404aac097c5a3258838d31dbdf1f3270eeb5d8c4e21481bdf0b86054

Observation e2dad82a-0490-4401-a0fa-dd809acb4476 · outbound

This paper cites Autoregressive image generation using residual quantization.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Autoregressive image generation using residual quantization

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.680540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.680540Z digest=sha256:a63e21875a4d4affdcb5ef9963ed79eac6c6c3e431ccf269494b7c8b1a808976

Observation 9a5a63a1-7d84-4fa7-9c57-fe42d513137f · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation LLaVA-OneVision: Easy Visual Task Transfer

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.684769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.684769Z digest=sha256:03ee7ab92e2aeed692ad8123995e42e4532b46688842c1dcd4ce4a3f03a46826

Observation e6328eb0-7a09-4e34-b5f5-ade67c75f138 · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.688963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.688963Z digest=sha256:b40c1f9343c5fd56bf214df484e3be11efd751e620d8a98e65a7365e4cc8a848

Observation dfb0530e-9d56-4664-9f36-9be5475d79c5 · outbound

This paper cites A survey of multimodel large language models.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation A survey of multimodel large language models

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:51.849440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:01:49.693312Z digest=sha256:1f7f6c40f450f89b80ed0b95c8263c7ec704c46e67ba682cbd6ccd0e0ba974dd

Observation 2208cd6b-3b7a-431b-924a-1ad2e4eedd7d · outbound

This paper cites Toklip: Marry visual tokens to clip for multimodal comprehension and generation, 2025.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Toklip: Marry visual tokens to clip for multimodal comprehension and generation, 2025

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:51.831628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:01:49.697319Z digest=sha256:b6c290fb8dbfb7fcd3dee92e21baae41e077362520a2ba939df338819a869f9b

Observation cb6bdd6f-6d7e-4939-8f90-cacfc8886da7 · outbound

This paper cites World Model on Million-Length Video And Language With Blockwise RingAttention.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation World Model on Million-Length Video And Language With Blockwise RingAttention

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.702024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.702024Z digest=sha256:be4a133b61ae3ad4dfb7e6887cff925d62ffd510f53118e579a5e98cb07087dc

Observation 69d9c7a3-63e9-4f07-b34c-8b154f8d7278 · outbound

This paper cites Improved baselines with visual instruction tuning.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Improved baselines with visual instruction tuning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.706281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.706281Z digest=sha256:7b226ca3fb78d81d84e4c0004920349fd4c3d62ca6d10ceb4c42718a86e97cb7

Observation 801b98e7-a0c7-4180-875f-475dbc187d5f · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.710584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.710584Z digest=sha256:ed22ed82a87016ea4d0a7944069dadded44ab5eb761776f1358190c3e6c412e6

Observation b7310155-5a00-4fd4-bf02-b6a5ee3a3858 · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? InEuropean conference on computer vision, pages 216–233.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Mmbench: Is your multi-modal model an all-around player? InEuropean conference on computer vision, pages 216–233

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.715276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.715276Z digest=sha256:1b07e882cdbfef95cc67962a057f1b64fc99e82c4c9f8075cbb1c6edde0cc81a

Observation c07dfd61-d9a0-43b6-aea3-adad6cb73cf1 · outbound

This paper cites Decoupled Weight Decay Regularization.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Decoupled Weight Decay Regularization

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.719739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.719739Z digest=sha256:1c2ea143088225f108cc8bca6a41bc139a533d97cec6b2a3a5bd882fdcd0ac48

Observation b9e15661-e791-481c-8847-73d935275cfd · outbound

This paper cites Ovis: Structural Embedding Alignment for Multimodal Large Language Model.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.723984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.723984Z digest=sha256:854b00a665b1fd5b7fc41acaf5215f66bc7c1d1c39ebcf26d88fbfdc2ea9a128

Observation 0297c9ca-c531-4fde-bd37-81888fd2e27f · outbound

This paper cites Unitok: A unified tokenizer for visual generation and understanding.arXiv preprint arXiv:2502.20321, 2025.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Unitok: A unified tokenizer for visual generation and understanding.arXiv preprint arXiv:2502.20321, 2025

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.728398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.728398Z digest=sha256:0e143b37746487e4a18d0979630de3ab03e5573741cbd3dadb1dd73c59be1b3d

Observation b82b6d9c-43e2-4f05-9534-8aad015c5f2a · outbound

This paper cites JanusFlow: Harmonizing Autoregression and Rectified Flow for Unified Multimodal Understanding and Generation.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation JanusFlow: Harmonizing Autoregression and Rectified Flow for Unified Multimodal Understanding and Generation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.732504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.732504Z digest=sha256:309ddedb1a617a849b73a714176ba08f36d963fc9d79da3353b4b19eafc82270

Observation 9393d3d7-75c3-4e01-895d-a6536e095847 · outbound

This paper cites BEiT v2: Masked Image Modeling with Vector-Quantized Visual Tokenizers.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation BEiT v2: Masked Image Modeling with Vector-Quantized Visual Tokenizers

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.737348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.737348Z digest=sha256:05c448f55598af2978ba5fdc4cc86c45621ea9da9d9cfde9a9e780518bb60f32

Observation 0f870014-1798-4722-a736-f277b0887bea · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.742089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.742089Z digest=sha256:3bc4570aadf1a33d0c6f2e2e0362c15e7860a8819fb73f5fa6e9b7c43b659e4a

Observation ce07186b-2fa2-4449-9253-24ff0f49ae5d · outbound

This paper cites TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.746510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.746510Z digest=sha256:700ad810adb59a36b3f5049c84bb3e5e0d8dc21fdfebe1012c8eb622950d85eb

Observation 25f9e83f-3624-4516-9e8a-ef6009465aa1 · outbound

This paper cites Beyond Next-Token: Next-X Prediction for Autoregressive Visual Generation.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Beyond Next-Token: Next-X Prediction for Autoregressive Visual Generation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.751169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.751169Z digest=sha256:2758cd1f9b63fdcefd57a4d9ce9969dc1cc093bf1f3be168fd443943337de608

Observation bcf6da21-bf3c-4997-9986-9c8c1ee82db2 · outbound

This paper cites High- resolution image synthesis with latent diffusion models.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation High- resolution image synthesis with latent diffusion models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.755625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.755625Z digest=sha256:4abed2dff6dfc8e05ad867d714ae976b5edc1dd779ef33323b279293484e54bc

Observation c99e1186-de25-478d-9674-808f355f0b44 · outbound

This paper cites Scalable Image Tokenization with Index Backpropagation Quantization.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Scalable Image Tokenization with Index Backpropagation Quantization

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.760298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.760298Z digest=sha256:d60f96a2ace80aef575ff09c64806de97ad52db05c0cb3b9d64a062fcc510aa9

Observation 741dca58-6de6-4114-b0c6-f8429c8d91b6 · outbound

This paper cites Towards vqa models that can read.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Towards vqa models that can read

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.765263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.765263Z digest=sha256:4217f3db1ce8e20a00a321e71db184fab1c6499070063e9708044091ac67b702

Observation c0101e35-3818-4596-b9e2-af2b176910a6 · outbound

This paper cites DualToken: Towards Unifying Visual Understanding and Generation with Dual Visual Vocabularies.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation DualToken: Towards Unifying Visual Understanding and Generation with Dual Visual Vocabularies

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.769637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.769637Z digest=sha256:6c18a25e5dfea268aa8601ad4c9e6f69f749c39c05d84b04de8654b805632bf7

Observation 9bea6b54-9227-4f47-b7d0-d89077778d14 · outbound

This paper cites Journeydb: A benchmark for generative image understanding.Advances in neural information processing systems, 36:49659–49678, 2023.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Journeydb: A benchmark for generative image understanding.Advances in neural information processing systems, 36:49659–49678, 2023

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.774524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.774524Z digest=sha256:d0e176b818de838d5af0a2c27e9804eefd8cc3c9a17be6f2cd1b33f9e933c2f0

Observation 9f449a2c-e8a0-495d-93cf-5922ea91ff05 · outbound

This paper cites Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.779091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.779091Z digest=sha256:2fbf4aa7d6de473a402c5da5489f0b62cfcde5f81bc4df28e504af3e0f508b8b

Observation 6425bd17-4f4f-4304-9a6a-3c9b8059245d · outbound

This paper cites Generative multimodal models are in-context learners.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Generative multimodal models are in-context learners

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.784005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.784005Z digest=sha256:68079725d18a823e51a900fc155430b5795927066d8eb2f4e9127433a9318823

Observation 5ce7e1ae-a651-46ff-b5e7-cbb259228f94 · outbound

This paper cites Emu: Generative Pretraining in Multimodality.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Emu: Generative Pretraining in Multimodality

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.788445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.788445Z digest=sha256:77ead37eb376b6a6aaa43e1292f80cc7e166ed9b03bf412b7e8402f2b83bad38

Observation ef4f571a-8c4f-4036-bd58-ec9d840575eb · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.793100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.793100Z digest=sha256:42f73484bf65e51019b5285d01cab40c7f6112b790f75987ea1fb56e74ee9f6d

Observation 9780d705-0bc5-4c10-9ff9-0cdb2158c955 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Gemini: A Family of Highly Capable Multimodal Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.797577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.797577Z digest=sha256:3ce95c73657001b5f61d35ed0eb49b7dc3361ab1fefbed368cee6be9b46a2e42

Observation 61084d5c-aaa5-43f4-af9d-aa1f029122d5 · outbound

This paper cites Visual autoregressive modeling: Scalable image generation via next-scale prediction.Advances in neural information processing systems, 37:84839–84865, 2024.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Visual autoregressive modeling: Scalable image generation via next-scale prediction.Advances in neural information processing systems, 37:84839–84865, 2024

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.801688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.801688Z digest=sha256:6ce504bdedce9a050f866493e17dd1ee6c1dc7511167792a1b966045d61011ed

Observation 7a2b6382-157a-4e49-9676-8fe6d40a94d6 · outbound

This paper cites MetaMorph: Multimodal Understanding and Generation via Instruction Tuning.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation MetaMorph: Multimodal Understanding and Generation via Instruction Tuning

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.805911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.805911Z digest=sha256:bb72d2592437a8c6b05c52bfe789c8707effcf0682e3f6fd2abc3082f5817598

Observation 1e630c62-7010-4cd4-90b4-3c0407fe3f0f · outbound

This paper cites Neural discrete representation learning.Advances in neural information processing systems, 30, 2017.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Neural discrete representation learning.Advances in neural information processing systems, 30, 2017

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.810741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.810741Z digest=sha256:9bbddaf5df5f6ea8c710d2952a16b9f9d359ad84b95163faecfe49a84cc4c570

Observation 5a97a6a8-26bc-4f56-9a80-f7f0ead67f4b · outbound

This paper cites Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.815176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.815176Z digest=sha256:a1497330c96772581bfe446c1d8937cb438df44f7c29e952e9248d777af1115d

Observation 8fe5a1de-7e37-4c94-9eb1-8fc0804bb377 · outbound

This paper cites ILLUME: Illuminating Your LLMs to See, Draw, and Self-Enhance.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation ILLUME: Illuminating Your LLMs to See, Draw, and Self-Enhance

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.819756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.819756Z digest=sha256:d2ed94026f4e64e7fb53d598fc2661e4cc1afb3333ed9bf17bf12ba799d93669

Observation 5e13fc24-4806-4e23-bae5-cebaa0c79874 · outbound

This paper cites Omnitok- enizer: A joint image-video tokenizer for visual generation.Advances in Neural Information Processing Systems, 37:28281–28295, 2024.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Omnitok- enizer: A joint image-video tokenizer for visual generation.Advances in Neural Information Processing Systems, 37:28281–28295, 2024

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:51.690847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:01:49.824085Z digest=sha256:e72b2611f0c3808b94c4192731e36343925ab918fbddd7278a37a2730574c8b7

Observation 5a6f2efd-9e75-46b7-9c18-974d95f567e1 · outbound

This paper cites Image under- standing makes for a good tokenizer for image generation.Advances in Neural Information Processing Systems, 37:31015–31035, 2024.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Image under- standing makes for a good tokenizer for image generation.Advances in Neural Information Processing Systems, 37:31015–31035, 2024

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:51.670739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:01:49.828530Z digest=sha256:2a4095f8d2e2bb17cc1da6a6efab49aa51b7929fbe426592208c1ec56881670a

Observation 7a8093d1-edb0-4dca-b5f1-fc8c28582e59 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.832665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.832665Z digest=sha256:0f1b175a5e8554dce0add4c4cb0a0a9276d6ec7f47b6b2cf2b5ab0d5f57476a7

Observation 638346e8-1ae3-4000-b244-7d8e29d0f2e1 · outbound

This paper cites Image as a foreign lan- guage: Beit pretraining for vision and vision-language tasks.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Image as a foreign lan- guage: Beit pretraining for vision and vision-language tasks

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.837077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.837077Z digest=sha256:35ae71621f1ff28a45c46daa986adcc9c8a1519da45b56bab5a4989d246d305c

Observation fb6c6446-7d35-42f3-9ac2-a12b09e0ca80 · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Emu3: Next-Token Prediction is All You Need

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.841242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.841242Z digest=sha256:8407dffd8bb1473705c38d353ad303f72d5940980a2f45f2ae58c6d94e87ee26

Observation 5d49eb60-9db7-43d2-894d-a55ed865ddb7 · outbound

This paper cites Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.845919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.845919Z digest=sha256:7f41d54b53f91e2733fa6f8be0c13ed6511b6212d89c95ec01b5b1d2b74e52c7

Observation fa027142-f6a4-4ae2-bfcd-7fbd7e199ab2 · outbound

This paper cites Liquid: Language Models are Scalable and Unified Multi-modal Generators.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Liquid: Language Models are Scalable and Unified Multi-modal Generators

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.850509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.850509Z digest=sha256:cb2a1f1d588b9753f2f7b7e20256f7f5ed51ae3cf8399c4c4e0cc2e4e1f060cd

Observation 4640a3a0-953a-4ead-be75-c1d4759418dc · outbound

This paper cites Next-gpt: Any-to-any multimodal llm.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Next-gpt: Any-to-any multimodal llm

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:51.636941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:01:49.854970Z digest=sha256:0e105e01f9f581467dba353dd9b44bf35a8b13b18bd7cbd6eb4b3b28e43bfbfe

Observation 58a09def-c607-40c0-b40f-9ea28861d1b1 · outbound

This paper cites Harmonizing Visual Representations for Unified Multimodal Understanding and Generation.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Harmonizing Visual Representations for Unified Multimodal Understanding and Generation

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.859294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.859294Z digest=sha256:edf48b52bc047e300e6d31e2b9380bec181f605a4c290059949bb219319f1536

Observation 94a77910-e224-46a6-a252-85d2bdc4603f · outbound

This paper cites VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.863781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.863781Z digest=sha256:661e65ad29bc85d32050d08ad777b491162ece4f170f28215345b98efc70bcbc

Observation f5143d83-63be-4c9c-be19-75acb5c5b9d2 · outbound

This paper cites Grok-1.5 vision preview, 6 2024.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Grok-1.5 vision preview, 6 2024

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:51.616407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:01:49.868739Z digest=sha256:d5d8832e8344c116bf037b05bd848eab08c39ab6b02458f5a1c4f8ab7448186e

Observation 0c1d8e90-0ad1-47b6-9d5e-1fdb31e97b4f · outbound

This paper cites OmniGen: Unified Image Generation.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation OmniGen: Unified Image Generation

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.873348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.873348Z digest=sha256:9ae88c23cecd76223715b698e60fba3a4a30d73e60fc52bc7b6f1f4280870818

Observation f267fa33-2a43-494f-843e-1cc236f425db · outbound

This paper cites Show-o: One Single Transformer to Unify Multimodal Understanding and Generation.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.878313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.878313Z digest=sha256:de7af0dc17a993667343f67a8660ffe39a313f2f4ca10b7b2eb880c2643a933f

Observation eca16e4f-e042-4a5e-9470-7a5a5991c70b · outbound

This paper cites MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.883147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.883147Z digest=sha256:28932e7e370d1e91f5f80a4329aecd7affe533c99be369e5f9a0bea4167a7caa

Observation 1e4c5d69-034a-48fb-8da2-858ccd1c4813 · outbound

This paper cites Efficientsam: Leveraged masked image pretraining for efficient segment anything.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Efficientsam: Leveraged masked image pretraining for efficient segment anything

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:51.595479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:01:49.887716Z digest=sha256:f34a9d97854110708bef61dcf42ab76011b25356323ed5658a208e903e48d501

Observation 1095fca2-3ae7-4175-a3c6-136135bc4670 · outbound

This paper cites Qwen2.5 Technical Report.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Qwen2.5 Technical Report

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.891862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.891862Z digest=sha256:2c3f753cbd20eb35701cab991bc41412b85f8c8c35c4bebdfbd0c872f050251e

Observation ec5bb5e2-5f5d-428a-98da-e6ec69ca4aac · outbound

This paper cites Vector-quantized Image Modeling with Improved VQGAN.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Vector-quantized Image Modeling with Improved VQGAN

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.896781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.896781Z digest=sha256:aa6b9398042feebfce870dd2e8da89451556605423a9a471df2e9848bcda184a

Observation 29635d4e-8db5-4ee0-ab66-e5d08ade24e7 · outbound

This paper cites Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.901168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.901168Z digest=sha256:29b8ff96bd6add8f756d98d3584f8780b00882ce47b19f3ce1b63e76c8bd94f5

Observation a7f74818-5d00-4ea5-92df-f2b5ab82163e · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.905672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.905672Z digest=sha256:a416dcf2c27d63e3915cc2836668e3642f1d83f51ed9f7e9f516994fa053032d

Observation cf029f21-64fd-48ee-a43b-14918d731c36 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.910134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.910134Z digest=sha256:147e740687dffe6e35309bfde87d4c09a947f6059798da6fdfca7fb477fd9aa0

Observation cf06dac7-6d75-4178-848f-b79dcaa58976 · outbound

This paper cites Sigmoid loss for language image pre-training.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Sigmoid loss for language image pre-training

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.914438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.914438Z digest=sha256:df30d03bd2c0a9f6756aa7be9b2e4d1a0d23c1f9c041529a9da7a4eca38b505b

Observation 7ffabdac-a96f-46a5-84fe-6fb89690a743 · outbound

This paper cites Token dynamics: Towards efficient and dynamic video token representation for video large language models.arXiv preprint arXiv:2503.16980, 2025.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Token dynamics: Towards efficient and dynamic video token representation for video large language models.arXiv preprint arXiv:2503.16980, 2025

Reference 82

Resolution
verified exact
raw_fallback, observed 2026-08-06T23:01:50.399667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:01:49.919010Z digest=sha256:82b4e9401c94b6533b7aeae1fbf0349197a03e593245d76c748d152b1a2d1bb4

Observation 5b8708d8-348c-4e4b-9285-0ba2dd27d76a · outbound

This paper cites Unified multimodal understanding and generation models: Advances, challenges, and opportunities.arXiv preprint arXiv:2505.02567, 2025.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Unified multimodal understanding and generation models: Advances, challenges, and opportunities.arXiv preprint arXiv:2505.02567, 2025

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.923951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.923951Z digest=sha256:a757526287d1ca9be662d543088a2efba5377a40dc0412e315e883bf6522afaf

Observation cf687e1a-b5f5-4563-a753-5427cbecb4a8 · outbound

This paper cites QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.928203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.928203Z digest=sha256:54bade39ab286ad62e276dffd7aed549f05ad3a8c4718222a9ddeded2e2a3764

Observation f4809218-b996-4377-b2ec-af22cdef42a5 · outbound

This paper cites Online clustered codebook.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Online clustered codebook

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.932726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.932726Z digest=sha256:05b9a7c5c6020ec391fa209db5b278b04b7bca8b43a589bdea71e21db4cca255

Observation 84e2bf29-c556-4028-856a-f85857f84d14 · outbound

This paper cites Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.937018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.937018Z digest=sha256:6b58c9a42a4ba6d07c8f0aad1a97c67418cb366388530549d7817ff3b97c5b25

Observation 6459ee8c-5f34-43d2-9b3e-5fba6bd2d812 · outbound

This paper cites Scaling the Codebook Size of VQGAN to 100,000 with a Utilization Rate of 99%.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Scaling the Codebook Size of VQGAN to 100,000 with a Utilization Rate of 99%

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.941622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.941622Z digest=sha256:c9a8791d3ff21bf776a86ffe5b42f6dd01254d85ad5c241ca85db6b98b5ccbc1

Observation caa43789-6d45-46a6-914e-9f32bf63ae11 · outbound

This paper cites Addressing representation collapse in vector quantized models with one linear layer.arXiv preprint arXiv:2411.02038, 2024.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Addressing representation collapse in vector quantized models with one linear layer.arXiv preprint arXiv:2411.02038, 2024

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.946448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.946448Z digest=sha256:5efe78c311fb6ab9959e464f84c9c2f41db918252e45af538fcf42d59d0adead

Observation 365b6c50-2d1c-47e0-a6eb-2cab4993cf8b · outbound

This paper cites OmniMamba: Efficient and Unified Multimodal Understanding and Generation via State Space Models.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation OmniMamba: Efficient and Unified Multimodal Understanding and Generation via State Space Models

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.951366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.951366Z digest=sha256:bfa980a071306537da3e89f7693dda134049e77abe0ce19aeb0edc6baf8da3dd

Observation bd75435d-6cf2-420b-88fc-211b99186a4d · outbound

This paper cites Let E={e 1, ...,eN } ⊂Rd denote visual embeddings sampled from a distribution p(e), and let C={c 1, ...,cK} ⊂Rd be a codebook with K discrete centroids.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Let E={e 1, ...,eN } ⊂Rd denote visual embeddings sampled from a distribution p(e), and let C={c 1, ...,cK} ⊂Rd be a codebook with K discrete centroids

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:51.534554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:01:49.957273Z digest=sha256:c10e1119aa5230b783f9bd5b6d32fd4d072a2dda0128987a9ab958cc4e9183e2

Observation 5289fcf7-54c8-430d-9dd6-f00fb34f66e5 · outbound

This paper cites an unresolved cited work.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Unresolved cited work

Reference 91

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:01:51.513926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:01:49.961746Z digest=sha256:a91212edc62a9dec6c70f60eb5a276cf03f3be9a6ef269a05c7aa7f0b5c8c688

Observation 6c9382c1-2830-4073-8eae-1c713deb4564 · outbound

This paper cites Let y be a semantic target label (e.g., object class, scene type), and let v=Q(e) be the discrete token assigned to embeddinge.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Let y be a semantic target label (e.g., object class, scene type), and let v=Q(e) be the discrete token assigned to embeddinge

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:51.489877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:01:49.966452Z digest=sha256:b66e2d1311247bf85e15fe4aff4c15337037f7d1086d7d22ad23d2a6d1b61509

Observation a003bd6f-da76-4860-be9c-742ba2295ec1 · outbound

This paper cites an unresolved cited work.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Unresolved cited work

Reference 93

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:01:51.469848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:01:49.971454Z digest=sha256:5e151862f676a8627c14b506e079c067ed51facbec48991063e94e70f86905bf

Observation af3ee04e-8a70-4148-bbd3-d27b89482896 · outbound

This paper cites an unresolved cited work.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Unresolved cited work

Reference 94

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:01:51.448653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:01:49.976800Z digest=sha256:09da316ddec2b58cd27405f1a95e9bc49ad1d056268d0f172cbe61b5bd0aa906

Pith citing papers

No inbound Pith citation observations are available.