Pith. sign in

Paper Citation Record · LEDGER

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation

As of 6 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 17 inbound Pith citation observations for arXiv:2604.24763.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.24763 v2

Coverage vector

measured 58 of 58 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-20T23:41:25.275207Z

measured 75 of 75 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T11:55:27.994286Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T13:37:06.881254Z

Reference resolution

58 of 58 outbound references displayed

  • verified exact54
  • verified fuzzy1
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f10e9c16-da04-44bd-a042-8b7609c692a2 · outbound

This paper cites Ming-flash-omni: A sparse, unified architecture for multimodal perception and generation.arXiv preprint arXiv:2510.24821.

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation Ming-flash-omni: A sparse, unified architecture for multimodal perception and generation.arXiv preprint arXiv:2510.24821

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-20T23:43:51.165136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T23:41:25.275207Z digest=sha256:68853c56eed91c11bc806416769a83457924d6ad2f3a6f7a8813898c35d06733

Observation aa21ed8b-d792-4f37-8981-a01aba58df48 · outbound

This paper cites LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training.

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-20T23:43:51.182896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T23:41:25.275207Z digest=sha256:6451aa11eb9f085c0c5e9afcb4398ac72e47e5134a4c57ecf511c17ac144ec00

Observation bbb8a7cf-805c-42e0-ab18-994621a85522 · outbound

This paper cites VGGRPO: Towards World-Consistent Video Generation with 4D Latent Reward.

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation VGGRPO: Towards World-Consistent Video Generation with 4D Latent Reward

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-14T02:20:26.920520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T23:41:25.275207Z digest=sha256:6a5b634d9d38da67dad935246746bed04b9f121657ee3116c5c0ab0a14d8d353

Observation d4a23b81-b675-474e-9942-5ea3b1ebd1f7 · outbound

This paper cites Qwen Technical Report.

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation Qwen Technical Report

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T23:43:51.123112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T23:41:25.275207Z digest=sha256:3a5c6bdfcb6c1af92cb59c4e02840e9f2e107d3c5a6790c12fe7d705c63ea35f

Observation 003d9e10-1d2d-4967-ac36-9fab5437f7a8 · outbound

This paper cites Qwen3-VL Technical Report.

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation Qwen3-VL Technical Report

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-20T23:43:51.224063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T23:41:25.275207Z digest=sha256:9bc8aeb46b7ff8b768d530e9583118e32f196a8f151ef93cc0058247cbbe3c68

Observation 72626697-3026-439a-bd03-6981120e3144 · outbound

This paper cites Bie, T., Cao, M., Chen, K., Du, L., Gong, M., Gong, Z., Gu, Y ., Hu, J., Huang, Z., Lan, Z., et al.

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation Bie, T., Cao, M., Chen, K., Du, L., Gong, M., Gong, Z., Gu, Y ., Hu, J., Huang, Z., Lan, Z., et al

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T23:43:51.088011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T23:41:25.275207Z digest=sha256:e9ba129cc2b5cb2e9e06da5a02ba8e7ced131a8981c11cf69574a0d88c871c42

Observation a95c9959-6841-4ac8-b32b-bcfbbdf71840 · outbound

This paper cites Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer.

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-20T23:43:51.133147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T23:41:25.275207Z digest=sha256:4883c941ee39f0357457359674637a92e917d9bb369fff04efec24cc205c9033

Observation 34dbc408-464e-4870-a933-076b1d57bfd0 · outbound

This paper cites BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset.

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-20T23:43:51.126327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T23:41:25.275207Z digest=sha256:6a1a8b0a514729188c41a0f5a3e3afee6bc9ff94a97dca10f7653f834c350025

Observation 02e8942e-6abf-4af7-a3b9-2e719e2aa46e · outbound

This paper cites PixelFlow: Pixel-Space Generative Models with Flow.

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation PixelFlow: Pixel-Space Generative Models with Flow

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-20T23:43:51.084701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T23:41:25.275207Z digest=sha256:d0e69bdebb17a0ebc38e29b5b1109ae99d0ac488674a4abbf8f81a825e7cc672

Observation 43c9d681-bb25-43b8-b687-61e13146f13c · outbound

This paper cites Emu3.5: Native Multimodal Models are World Learners.

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation Emu3.5: Native Multimodal Models are World Learners

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-20T23:43:51.071493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T23:41:25.275207Z digest=sha256:21935751b917767bb14d2d1c4c49c59b770c6a1c3bc1caf1e03d057dfd140619

Observation 72fdccde-a861-4c3e-9420-90339c50ded2 · outbound

This paper cites Emerging Properties in Unified Multimodal Pretraining.

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation Emerging Properties in Unified Multimodal Pretraining

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-20T23:43:51.115609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T23:41:25.275207Z digest=sha256:30b05aeb9e3d9d24fa9caae1fa1ed21cf0ae60f7c52e6a6ec5d7a198a2182a88

Observation 721f847d-105d-4625-abbd-92c9c73210f1 · outbound

This paper cites From pixels to words–towards native vision-language primitives at scale.arXiv preprint arXiv:2510.14979, 2025a.

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation From pixels to words–towards native vision-language primitives at scale.arXiv preprint arXiv:2510.14979, 2025a

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-20T23:43:51.063724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T23:41:25.275207Z digest=sha256:f37f0f8eb401c19b8691851f8919e45d46b0198adc49311f9f993dfd1c78466a

Observation 1683cb3c-e9e0-4994-9a8a-409e6b24e763 · outbound

This paper cites Unified Autoregressive Visual Generation and Understanding with Continuous Tokens.

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation Unified Autoregressive Visual Generation and Understanding with Continuous Tokens

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-20T23:43:51.078460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T23:41:25.275207Z digest=sha256:566ea32a4b230621b0a33af8c54d1c648794d49a791188d4050499b16efd724c

Observation a2363116-a38e-4124-9544-8abb6ebdbed1 · outbound

This paper cites X-Omni: Reinforcement Learning Makes Discrete Autoregressive Image Generative Models Great Again.

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation X-Omni: Reinforcement Learning Makes Discrete Autoregressive Image Generative Models Great Again

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-20T23:43:51.185958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T23:41:25.275207Z digest=sha256:082c26c39acc3349c693ed62e9ccf5a0ea7bf285a094f89d42745f281527453f

Observation 47986ce8-2cf0-4917-ad52-e37215bcd7c2 · outbound

This paper cites The Llama 3 Herd of Models.

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation The Llama 3 Herd of Models

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-20T23:43:51.198204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T23:41:25.275207Z digest=sha256:cd4e118a72f5c15f92d1cc4c9915e8b9c099468a5dcb80021d249ed31c4cb0ae

Observation d48b7176-f7d3-438e-aec4-f7d2d4da5b79 · outbound

This paper cites Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations.

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-20T23:43:51.091129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T23:41:25.275207Z digest=sha256:be7cf271c93b13c22ee11136f54bff18b0aa117a277fe59ecc315f2c0f2dbd14

Observation f1b3d296-1dcc-417c-9045-60c6cfc9b2d4 · outbound

This paper cites Uni-x: Mitigating modality conflict with a two-end- separated architecture for unified multimodal models.arXiv preprint arXiv:2509.24365.

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation Uni-x: Mitigating modality conflict with a two-end- separated architecture for unified multimodal models.arXiv preprint arXiv:2509.24365

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-20T23:43:51.059913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T23:41:25.275207Z digest=sha256:60b188a6867cc868e8dc4e397dbc7c03a68445d3669c11fee96bea616b7d29c6

Observation dd250552-955d-4a8b-a4a5-47a1ff54fc22 · outbound

This paper cites Emma: Efficient multimodal understanding, generation, and editing with a unified architecture.

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation Emma: Efficient multimodal understanding, generation, and editing with a unified architecture

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-20T23:43:51.204198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T23:41:25.275207Z digest=sha256:06cf3f8cd9ce2024f9fcf9c60ff95834244c1a5969e1f6289dd86cc66cf1c294

Observation 05581f8c-820e-4922-8863-1bbc53f66c78 · outbound

This paper cites ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment.

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-20T23:43:51.081493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T23:41:25.275207Z digest=sha256:18e59b212d5f4b21f53103a675b68290b9675fc510701dfeaf2adfcf4b73ac44

Observation 86b12f4f-653a-482f-b4f1-a599f959860e · outbound

This paper cites Ming-univision: Joint image understanding and generation with a unified continuous tokenizer.arXiv preprint arXiv:2510.06590, 2025a.

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation Ming-univision: Joint image understanding and generation with a unified continuous tokenizer.arXiv preprint arXiv:2510.06590, 2025a

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-20T23:43:51.221343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T23:41:25.275207Z digest=sha256:043bdbb8cb6720b3e24a7c665cc02f48a28438f0c7538b2986765e689c5a602c

Observation ba1c9eb7-70df-43e6-8237-b274202c8ade · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation LLaVA-OneVision: Easy Visual Task Transfer

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-20T23:43:51.201029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T23:41:25.275207Z digest=sha256:aed3a631c5e28c481d5b86f5a7d849d1c74bef0b337e99ff89596a80f62ff191

Observation b8e1917f-cc1f-4dec-b203-0b3bf18af1c9 · outbound

This paper cites Back to Basics: Let Denoising Generative Models Denoise.

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation Back to Basics: Let Denoising Generative Models Denoise

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-20T23:43:51.230376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T23:41:25.275207Z digest=sha256:fd56e7e1ac297a2264699f2fb7242a11f1adfc548edf4ecc71322f82b93e015b

Observation d3a281b8-a4c6-4ec7-9e22-7b5e6f809ea1 · outbound

This paper cites VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling.

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-20T23:43:51.161913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T23:41:25.275207Z digest=sha256:b167fbf4152f434100c879eb0ba94430267819c8b2cb073a87ab5cc13781c9c4

Observation a5036c6b-1a90-4e68-a26d-226ec35f09c4 · outbound

This paper cites UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation.

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-20T23:43:51.215328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T23:41:25.275207Z digest=sha256:cdf91575fd45bb1ba5a140031285c6b6741b640b81607c15bb4f6dd89b1259b8

Observation 6fa77aed-f707-4f87-9e84-f8485345517d · outbound

This paper cites Tuna: Taming unified visual representations for native unified multimodal models.

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation Tuna: Taming unified visual representations for native unified multimodal models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-20T23:43:51.207144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T23:41:25.275207Z digest=sha256:84356c456deef1ae283bb1b7e3d42557a7c48048555c96bbd59ed9f088fc60cc

Observation e885048b-1347-4ffe-96a9-dcf1154235a3 · outbound

This paper cites Decoupled Weight Decay Regularization.

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation Decoupled Weight Decay Regularization

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-20T23:43:51.209927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T23:41:25.275207Z digest=sha256:0356eeb6d20cd62b466d3d7d01716846762515bcbbd081ea48a0f48c43032b3f

Observation f82ec47e-2cdb-4108-8ad9-40949b9471f4 · outbound

This paper cites Unitok: A unified tokenizer for visual generation and understanding.arXiv preprint arXiv:2502.20321, 2025a.

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation Unitok: A unified tokenizer for visual generation and understanding.arXiv preprint arXiv:2502.20321, 2025a

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-20T23:43:51.094637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T23:41:25.275207Z digest=sha256:df1099748b44a806a1bac8dfdd2957696781a6995e794d1078c7d98db9ca7992

Observation c41ea23d-ac1f-4cae-bca3-3a3b1838ce8d · outbound

This paper cites Does understanding inform generation in unified multimodal models? from analysis to path forward.arXiv preprint arXiv:2511.20561.

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation Does understanding inform generation in unified multimodal models? from analysis to path forward.arXiv preprint arXiv:2511.20561

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-20T23:43:51.098439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T23:41:25.275207Z digest=sha256:5a6856209f5ed4b3bc4732e4ac2020391e4f270d6fa8bf3a4292b0911c18aba8

Observation cee7d3ff-0c0d-4d00-bf20-3d8ed4726065 · outbound

This paper cites Roni Paiss, Ariel Ephrat, Omer Tov, Shiran Zada, Inbar Mosseri, Michal Irani, and Tali Dekel.

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation Roni Paiss, Ariel Ephrat, Omer Tov, Shiran Zada, Inbar Mosseri, Michal Irani, and Tali Dekel

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T23:43:51.767832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T23:41:25.275207Z digest=sha256:fe1e358ed7f0f11d022a116112a49c6f9e965cb3d0802f4944501fd3535e76fc

Observation d46546bf-b744-497b-9b7a-abbb50206e52 · outbound

This paper cites Histream: Efficient high-resolution video generation via redundancy-eliminated streaming.arXiv preprint arXiv:2512.21338.

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation Histream: Efficient high-resolution video generation via redundancy-eliminated streaming.arXiv preprint arXiv:2512.21338

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-20T23:43:51.158464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T23:41:25.275207Z digest=sha256:7ad53d8c0b73fa6ae06a71a6e36841bb5039ea46d89b90e970aeb5e0d7233f39

Observation 542a1cf9-2fe7-4fa3-a9b7-0c64c5f89c55 · outbound

This paper cites Mammothmoda2: A unified ar-diffusion framework for multimodal understanding and generation.

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation Mammothmoda2: A unified ar-diffusion framework for multimodal understanding and generation

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-20T23:43:51.115596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T23:41:25.275207Z digest=sha256:082eae9943f009cffc7eb84952447e45488eea95717ef1ec87c5b63a8580765b

Observation 7b303691-2cef-4775-90ec-5adcbc8b568d · outbound

This paper cites Muddit: Liberating Generation Beyond Text-to-Image with a Unified Discrete Diffusion Model.

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation Muddit: Liberating Generation Beyond Text-to-Image with a Unified Discrete Diffusion Model

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-20T23:43:51.168069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T23:41:25.275207Z digest=sha256:00cb0806d1e36bc76e2a9b640718fd69d5babaa22aa71db01ebae0dc5f11e200

Observation f466b11f-5c7c-44ea-9f02-f15502eb6161 · outbound

This paper cites Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation.

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-20T23:43:51.180253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T23:41:25.275207Z digest=sha256:c72808cdf10ffa9ddf775cff883ebe16fbd7838b657355a2b55f707b3339c149

Observation d467844e-d54c-44a3-bdc1-36a6dbb2340a · outbound

This paper cites Unilip: Adapting clip for unified multimodal understanding, generation and editing.

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation Unilip: Adapting clip for unified multimodal understanding, generation and editing

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-20T23:43:51.108996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T23:41:25.275207Z digest=sha256:397892c9820fc5ae3c0fa4bf974267d86fa754ce749242865a6986a54a822013

Observation a446c5a9-1e2f-4cf8-b6dc-531f51947ee5 · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-20T23:43:51.177504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T23:41:25.275207Z digest=sha256:f2f30e9bfc3efca0c1c9feb7f80182af08f035203526125c5acd17b3c5da93f0

Observation 0454ea17-2581-470b-a4cf-6ed4919cf638 · outbound

This paper cites LongCat-Image Technical Report.

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation LongCat-Image Technical Report

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-20T23:43:51.192137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T23:41:25.275207Z digest=sha256:b4b2ef3bf18ac8034563957dad26f5016a50f43e4cbfe3855c58e6d0f1d2edfc

Observation 7768f308-46d6-429a-84bb-f3be31651339 · outbound

This paper cites Internvl-u: Democratizing unified multimodal models for understanding, reasoning, generation and editing.

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation Internvl-u: Democratizing unified multimodal models for understanding, reasoning, generation and editing

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-20T23:43:51.052384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T23:41:25.275207Z digest=sha256:c4e5ab6ab425bbd2f1d4e0707421c4b9730f4ac3cf40bc5afdbbda19f093edf2

Observation 2f5c6bb5-4b09-495a-979c-3596d0ab9730 · outbound

This paper cites Beyond language modeling: An exploration of multimodal pretraining.

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation Beyond language modeling: An exploration of multimodal pretraining

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-20T23:43:51.218228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T23:41:25.275207Z digest=sha256:efa55c4fdb2459dc169673feffcdd567549ff76556860ab251ad568504db0f8c

Observation 9f938dcc-30c8-454c-a350-3808e2b5e8b5 · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-05-20T23:43:51.119286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T23:41:25.275207Z digest=sha256:83a89ce7d6aac905bfeb1518ad6e2910af9ad77250eeb0f213871a6b0dc346b4

Observation 4fdd0d68-d172-4b47-aca6-75aba365f915 · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation Wan: Open and Advanced Large-Scale Video Generative Models

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-20T23:43:51.136076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T23:41:25.275207Z digest=sha256:403c5c482305dfd040b3124b09b412aa69a169f909b6c38a3b88754acc0e8738

Observation ed7f43f2-3360-479d-bc81-2260ec90bc9f · outbound

This paper cites Ovis-U1 Technical Report.

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation Ovis-U1 Technical Report

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-20T23:43:51.147802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T23:41:25.275207Z digest=sha256:418856c0dea3e76e8a86d0f916aa745853c30850c114e722e6f8571f91140c90

Observation 334ecb48-2067-4346-91be-f1a6d78d489f · outbound

This paper cites UniVideo: Unified Understanding, Generation, and Editing for Videos.

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation UniVideo: Unified Understanding, Generation, and Editing for Videos

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-07-07T03:17:13.018030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T23:41:25.275207Z digest=sha256:10e11d58b1693ec597581667e5086f267f97f05cffe12fd73c85aba8937d56e3

Observation 89c604ab-8f2a-4aab-8fee-696b595a2204 · outbound

This paper cites Qwen-Image Technical Report.

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation Qwen-Image Technical Report

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-05-20T23:43:51.101640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T23:41:25.275207Z digest=sha256:a18b2834976737038eb6778044b2b458d7c3e56bad1c12f997a9fab3aaed708d

Observation 921dab5b-a36c-4778-adfb-20bfbf7fbcb0 · outbound

This paper cites OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation.

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-20T23:43:51.195508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T23:41:25.275207Z digest=sha256:8d18e03a35d019a0eb3b91da78bf5730bfd022cc4f2fa0fa49f0a0409c194b9f

Observation 5c35c0d9-e5c1-416d-b475-ff16dc69a0a5 · outbound

This paper cites Reconstruction Alignment Improves Unified Multimodal Models.

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation Reconstruction Alignment Improves Unified Multimodal Models

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-06-26T02:15:36.738364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T23:41:25.275207Z digest=sha256:5eea59da03c58cfbd581cf923b74f10f444879a28f50fcd35bec497bfe34ca63

Observation 2d4ec96d-b77a-4d76-ac64-d13e0329d4fb · outbound

This paper cites Show-o2: Improved Native Unified Multimodal Models.

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation Show-o2: Improved Native Unified Multimodal Models

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-05-20T23:43:51.113006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T23:41:25.275207Z digest=sha256:8e78d0db0e7672e08d14ecb9077b6d37d07a3372571779ac67a2309133403c02

Observation a0811705-ddff-44b3-bd81-7a3f14c18916 · outbound

This paper cites VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language Models.

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language Models

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-20T23:43:51.138864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T23:41:25.275207Z digest=sha256:bb6f9b9b02b6d6deffb7b657c4dd241ebd66c6f7baa48ed2b6436fd72859b3ad

Observation 9e3482fc-0701-46c0-8218-8693526b6cbe · outbound

This paper cites MMaDA: Multimodal Large Diffusion Language Models.

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation MMaDA: Multimodal Large Diffusion Language Models

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-05-20T23:43:51.129679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T23:41:25.275207Z digest=sha256:0b5cff8218c5771362bf0a0d7707e67ccf5eabf48655aca4ce6c2ddae5dcbb26

Observation 1dd6f632-2917-4b90-9e78-3438c90654cb · outbound

This paper cites ImgEdit: A Unified Image Editing Dataset and Benchmark.

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation ImgEdit: A Unified Image Editing Dataset and Benchmark

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-05-20T23:43:51.154165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T23:41:25.275207Z digest=sha256:fa11091d5e7136797fe877cd8ad58ed5f44bed3a62f8a5409a37325a84d4eab6

Observation e54215e6-0e04-43d0-937d-becee4f02a46 · outbound

This paper cites Llada-o: An effective and length-adaptive omni diffusion model.arXiv preprint arXiv:2603.01068.

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation Llada-o: An effective and length-adaptive omni diffusion model.arXiv preprint arXiv:2603.01068

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-20T23:43:51.189154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T23:41:25.275207Z digest=sha256:79406aaaa58e328399e5e3fe4933fb7a194fb60e393a0280712d188336b4c6de

Observation 0d4639ad-a10a-4ed7-bc50-b533bc33d46a · outbound

This paper cites Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think.

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-05-20T23:43:51.150877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T23:41:25.275207Z digest=sha256:51619f10e83c775b00a947c25452f46920d2c089049febf20d7a8475691909b2

Observation 1a465710-521a-472a-8533-a8a4f9e8c82f · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-05-20T23:43:51.111865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T23:41:25.275207Z digest=sha256:c9a52d943fd7be3c4bdd55d4951583a3a1c2d445601bb47f5dc0b5deaa2c7633

Observation cf1e5e24-eef8-46fd-9f45-268c9e4766cd · outbound

This paper cites PixelDiT: Pixel Diffusion Transformers for Image Generation.

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation PixelDiT: Pixel Diffusion Transformers for Image Generation

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-05-20T23:43:51.212501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T23:41:25.275207Z digest=sha256:ef882ff70436ea6444b29b60b5bf605f97def13bff4a6b54f31c4d57facf9f9f

Observation dce48e95-9d8f-475e-95de-a077b01672f5 · outbound

This paper cites Uniflow: A unified pixel flow tokenizer for visual understanding and generation.arXiv preprint arXiv:2510.10575.

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation Uniflow: A unified pixel flow tokenizer for visual understanding and generation.arXiv preprint arXiv:2510.10575

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-20T23:43:51.144460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T23:41:25.275207Z digest=sha256:32cad31e36c1bb3949cb131816e23a9eabd01f1f561a5e6ded3a35fab07d3c11

Observation 6e715181-b3df-4444-bf5c-054d92fbe1bd · outbound

This paper cites Penguin-vl: Exploring the efficiency limits of vlm with llm-based vision encoders.arXiv preprint arXiv:2603.06569, 2026a.

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation Penguin-vl: Exploring the efficiency limits of vlm with llm-based vision encoders.arXiv preprint arXiv:2603.06569, 2026a

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-20T23:43:51.122848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T23:41:25.275207Z digest=sha256:6207c4b3bd1bf5fef6c767577c0e51a1d93f41e4e7aeafbbb7d4e766a91a6870

Observation a4cd0c68-afba-4977-9d8b-f5f60c3df444 · outbound

This paper cites Diffusion Transformers with Representation Autoencoders.

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation Diffusion Transformers with Representation Autoencoders

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-05-20T23:43:51.141634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T23:41:25.275207Z digest=sha256:adcf641df6e0e150ab45d26cd002411491dea62c71d74239149ff63eb367d729

Observation 8c462b17-eef0-4708-bb87-b3f03c431587 · outbound

This paper cites Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model.

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-05-20T23:43:51.227568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T23:41:25.275207Z digest=sha256:82ab6598fa7267dc117d9d13b3db84153151ce12fde0e9b07841301a1cafa3b1

Observation c276a854-1ceb-4b4a-9443-5f50ba344815 · outbound

This paper cites Scaling zero-shot reference-to-video generation.arXiv preprint arXiv:2512.06905.

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation Scaling zero-shot reference-to-video generation.arXiv preprint arXiv:2512.06905

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-20T23:43:51.174771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T23:41:25.275207Z digest=sha256:45a5d02be400aa28223ab4cfd32a35b809c5b0ffcd7f007b4660454ed9216af2

Pith citing papers

Observation 5cd4c18f-214b-4a12-8456-b454291dc81c · inbound

STARFlow2: Bridging Language Models and Normalizing Flows for Unified Multimodal Generation cites this paper.

STARFlow2: Bridging Language Models and Normalizing Flows for Unified Multimodal Generation Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-11T02:45:57.023589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T02:44:59.644215Z digest=sha256:6fda0218e634ea055200bb5fc15057804f1376d2903fb3682af6e9431231ae6c

Observation 79cb73b9-d6e8-4043-be91-48f40f65f8aa · inbound

Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm cites this paper.

Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-05-13T06:07:22.907523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T06:02:39.114660Z digest=sha256:879feeeab971407360be5bc13e1875e98c2149959982831decb0ebd6aa78a6c8

Observation f919d687-9f33-4bbc-91b7-f97d15e76730 · inbound

Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm cites this paper.

Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:05:46.658675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T22:18:55.933311Z digest=sha256:465c7291b52ca484e81e4506220bb3693b0a02531d42d4a453127a250c029e37

Observation 7e87a0de-826d-44e1-b9fd-4df531930344 · inbound

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture cites this paper.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation

Reference 85

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:17:18.656071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:273b55cfe03ac9e9f11e40fea7c65533879be3865f65561f35faa300978d42e6

Observation eecd669e-9280-4a33-9e24-cb7ad660f112 · inbound

Lance: Unified Multimodal Modeling by Multi-Task Synergy cites this paper.

Lance: Unified Multimodal Modeling by Multi-Task Synergy Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation

Reference 78

Resolution
verified exact
local_arxiv, observed 2026-05-20T11:48:14.832936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T11:46:52.658984Z digest=sha256:54e6828157cb38173f6165e17dfb7a40176aaf501213aa13c4e454cafa1954a4

Observation 1841f1f9-3efe-45ae-ad77-4ed7a70d1b87 · inbound

Lance: Unified Multimodal Modeling by Multi-Task Synergy cites this paper.

Lance: Unified Multimodal Modeling by Multi-Task Synergy Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation

Reference 79

Resolution
verified exact
local_arxiv, observed 2026-05-21T07:59:50.672370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:56cb8cd760bd597477c09b7223d9bb042ca7009581f6a95986aec0308fc31fc6

Observation ecb394ca-831f-4321-9ff1-79a7ca9d4c73 · inbound

Toward Native Multimodal Modeling: A Roadmap cites this paper.

Toward Native Multimodal Modeling: A Roadmap Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-06-29T23:04:01.807470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T22:58:38.610609Z digest=sha256:652fcbd412c0a20d49322304e668a9d58e3be58b351a407f971664cff225904d

Observation aaae50c0-09d6-4294-91d0-c03f773b6f88 · inbound

How and What to Imagine? Visual Thinking in Unified Multimodal Models for Cross-View Spatial Reasoning cites this paper.

How and What to Imagine? Visual Thinking in Unified Multimodal Models for Cross-View Spatial Reasoning Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-06-29T18:23:50.680893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-29T18:19:40.323661Z digest=sha256:09c451fbda0834577f4b32847a72d0466104015e72ed5da7ff5375322321fdb3

Observation 057076fe-1fe7-45c2-b07d-830f1064b729 · inbound

Representation Forcing for Bottleneck-Free Unified Multimodal Models cites this paper.

Representation Forcing for Bottleneck-Free Unified Multimodal Models Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-07-01T19:16:00.974604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T22:54:10.460872Z digest=sha256:c3650ab9417fda2fc579d47e38cd9646498e24f50163adacd8797b054c631c80

Observation 3fd2ec9a-f921-45db-a224-2f68be5cd6e7 · inbound

Representation Forcing for Bottleneck-Free Unified Multimodal Models cites this paper.

Representation Forcing for Bottleneck-Free Unified Multimodal Models Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-12T15:31:57.426559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T15:31:57.426559Z digest=sha256:c949addd07df26d8f17676c088b3389a38c84a1f034cccff2c122714cf9a1794

Observation ce8add9b-090a-4cf1-9c68-236101331b6c · inbound

LLM can Read Spectrogram: Encoder-free Speech-Language Modeling cites this paper.

LLM can Read Spectrogram: Encoder-free Speech-Language Modeling Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-07-03T03:47:35.738546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T14:40:40.673091Z digest=sha256:4ac86388803f24ea1b1433f69c8861bf8f7cb726e3414e9936fcd222ad9592e6

Observation b55fc75b-67a4-49dd-a29b-6a47fc00ccde · inbound

LLM can Read Spectrogram: Encoder-free Speech-Language Modeling cites this paper.

LLM can Read Spectrogram: Encoder-free Speech-Language Modeling Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-07-03T23:59:06.141726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-03T23:55:11.698728Z digest=sha256:f3f13248e951ec254bbebc38be57d83accab499ec9d53468589ce90b4e77e6ec

Observation 8c24bcf8-cb8a-4126-9c27-970d8c9c3642 · inbound

IV-CoT: Implicit Visual Chain-of-Thought for Structure-Aware Text-to-Image Generation cites this paper.

IV-CoT: Implicit Visual Chain-of-Thought for Structure-Aware Text-to-Image Generation Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation

Reference 25

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T16:39:58.204439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-26T00:19:49.071495Z digest=sha256:fd784ea63dca1d1d01af1ca6d587fecd2b6cdae1eba3b1ede60487e9d1a31b93

Observation 08c59d78-89a3-47f3-89a3-027c525b7e77 · inbound

GEAR: Guided End-to-End AutoRegression for Image Synthesis cites this paper.

GEAR: Guided End-to-End AutoRegression for Image Synthesis Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-07-01T10:35:42.603436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-01T05:19:21.647714Z digest=sha256:3821dae4ab910ff5f538091d2735889702b047fecf80024c2ef6ab09442f500b

Observation 10f0d88f-3a7f-4a92-9a22-8170a221360b · inbound

APIVOT: Adaptive Planning with Interleaved Vision-Language Thoughts cites this paper.

APIVOT: Adaptive Planning with Interleaved Vision-Language Thoughts Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-07-10T13:37:06.882491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-10T13:33:17.574108Z digest=sha256:2eeb3a97130cd8eddf7a723dea24012ec55094103da4b9df9381f60944c85fc7

Observation e72af9cb-8b76-4ba0-9cf0-57fbbfa3ebad · inbound

STBridge: Shared-Target Alignment for Bridging Understanding and Generation in UMMs cites this paper.

STBridge: Shared-Target Alignment for Bridging Understanding and Generation in UMMs Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T18:56:57.561814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:56:57.561814Z digest=sha256:8b4ec44cd653666eadb2a3ea3896274358a3fa5d10543903bc644699b4630580

Observation 376e13c2-d6df-4e2b-b34a-86dc7b3a7e44 · inbound

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes cites this paper.

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-06T11:55:27.994286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:55:27.994286Z digest=sha256:4fa6692392ee37e8f6864b4cff2cb22a9a1963c6867162846c475ef7d4974f4a