Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-20T23:41:25.275207Z
Paper Citation Record · LEDGER
As of 6 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 17 inbound Pith citation observations for arXiv:2604.24763.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-20T23:41:25.275207Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T11:55:27.994286Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-10T13:37:06.881254Z
58 of 58 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation f10e9c16-da04-44bd-a042-8b7609c692a2 · outbound
Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation Ming-flash-omni: A sparse, unified architecture for multimodal perception and generation.arXiv preprint arXiv:2510.24821
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation aa21ed8b-d792-4f37-8981-a01aba58df48 · outbound
Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation bbb8a7cf-805c-42e0-ab18-994621a85522 · outbound
Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation VGGRPO: Towards World-Consistent Video Generation with 4D Latent Reward
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d4a23b81-b675-474e-9942-5ea3b1ebd1f7 · outbound
Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation Qwen Technical Report
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 003d9e10-1d2d-4967-ac36-9fab5437f7a8 · outbound
Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation Qwen3-VL Technical Report
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 72626697-3026-439a-bd03-6981120e3144 · outbound
Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation Bie, T., Cao, M., Chen, K., Du, L., Gong, M., Gong, Z., Gu, Y ., Hu, J., Huang, Z., Lan, Z., et al
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a95c9959-6841-4ac8-b32b-bcfbbdf71840 · outbound
Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 34dbc408-464e-4870-a933-076b1d57bfd0 · outbound
Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 02e8942e-6abf-4af7-a3b9-2e719e2aa46e · outbound
Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation PixelFlow: Pixel-Space Generative Models with Flow
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 43c9d681-bb25-43b8-b687-61e13146f13c · outbound
Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation Emu3.5: Native Multimodal Models are World Learners
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 72fdccde-a861-4c3e-9420-90339c50ded2 · outbound
Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation Emerging Properties in Unified Multimodal Pretraining
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 721f847d-105d-4625-abbd-92c9c73210f1 · outbound
Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation From pixels to words–towards native vision-language primitives at scale.arXiv preprint arXiv:2510.14979, 2025a
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1683cb3c-e9e0-4994-9a8a-409e6b24e763 · outbound
Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation Unified Autoregressive Visual Generation and Understanding with Continuous Tokens
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a2363116-a38e-4124-9544-8abb6ebdbed1 · outbound
Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation X-Omni: Reinforcement Learning Makes Discrete Autoregressive Image Generative Models Great Again
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 47986ce8-2cf0-4917-ad52-e37215bcd7c2 · outbound
Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation The Llama 3 Herd of Models
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d48b7176-f7d3-438e-aec4-f7d2d4da5b79 · outbound
Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f1b3d296-1dcc-417c-9045-60c6cfc9b2d4 · outbound
Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation Uni-x: Mitigating modality conflict with a two-end- separated architecture for unified multimodal models.arXiv preprint arXiv:2509.24365
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation dd250552-955d-4a8b-a4a5-47a1ff54fc22 · outbound
Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation Emma: Efficient multimodal understanding, generation, and editing with a unified architecture
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 05581f8c-820e-4922-8863-1bbc53f66c78 · outbound
Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 86b12f4f-653a-482f-b4f1-a599f959860e · outbound
Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation Ming-univision: Joint image understanding and generation with a unified continuous tokenizer.arXiv preprint arXiv:2510.06590, 2025a
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ba1c9eb7-70df-43e6-8237-b274202c8ade · outbound
Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation LLaVA-OneVision: Easy Visual Task Transfer
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b8e1917f-cc1f-4dec-b203-0b3bf18af1c9 · outbound
Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation Back to Basics: Let Denoising Generative Models Denoise
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d3a281b8-a4c6-4ec7-9e22-7b5e6f809ea1 · outbound
Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a5036c6b-1a90-4e68-a26d-226ec35f09c4 · outbound
Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6fa77aed-f707-4f87-9e84-f8485345517d · outbound
Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation Tuna: Taming unified visual representations for native unified multimodal models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e885048b-1347-4ffe-96a9-dcf1154235a3 · outbound
Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation Decoupled Weight Decay Regularization
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f82ec47e-2cdb-4108-8ad9-40949b9471f4 · outbound
Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation Unitok: A unified tokenizer for visual generation and understanding.arXiv preprint arXiv:2502.20321, 2025a
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c41ea23d-ac1f-4cae-bca3-3a3b1838ce8d · outbound
Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation Does understanding inform generation in unified multimodal models? from analysis to path forward.arXiv preprint arXiv:2511.20561
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation cee7d3ff-0c0d-4d00-bf20-3d8ed4726065 · outbound
Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation Roni Paiss, Ariel Ephrat, Omer Tov, Shiran Zada, Inbar Mosseri, Michal Irani, and Tali Dekel
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d46546bf-b744-497b-9b7a-abbb50206e52 · outbound
Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation Histream: Efficient high-resolution video generation via redundancy-eliminated streaming.arXiv preprint arXiv:2512.21338
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 542a1cf9-2fe7-4fa3-a9b7-0c64c5f89c55 · outbound
Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation Mammothmoda2: A unified ar-diffusion framework for multimodal understanding and generation
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 7b303691-2cef-4775-90ec-5adcbc8b568d · outbound
Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation Muddit: Liberating Generation Beyond Text-to-Image with a Unified Discrete Diffusion Model
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f466b11f-5c7c-44ea-9f02-f15502eb6161 · outbound
Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d467844e-d54c-44a3-bdc1-36a6dbb2340a · outbound
Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation Unilip: Adapting clip for unified multimodal understanding, generation and editing
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a446c5a9-1e2f-4cf8-b6dc-531f51947ee5 · outbound
Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation Chameleon: Mixed-Modal Early-Fusion Foundation Models
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0454ea17-2581-470b-a4cf-6ed4919cf638 · outbound
Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation LongCat-Image Technical Report
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 7768f308-46d6-429a-84bb-f3be31651339 · outbound
Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation Internvl-u: Democratizing unified multimodal models for understanding, reasoning, generation and editing
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2f5c6bb5-4b09-495a-979c-3596d0ab9730 · outbound
Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation Beyond language modeling: An exploration of multimodal pretraining
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9f938dcc-30c8-454c-a350-3808e2b5e8b5 · outbound
Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4fdd0d68-d172-4b47-aca6-75aba365f915 · outbound
Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation Wan: Open and Advanced Large-Scale Video Generative Models
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ed7f43f2-3360-479d-bc81-2260ec90bc9f · outbound
Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation Ovis-U1 Technical Report
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 334ecb48-2067-4346-91be-f1a6d78d489f · outbound
Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation UniVideo: Unified Understanding, Generation, and Editing for Videos
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 89c604ab-8f2a-4aab-8fee-696b595a2204 · outbound
Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation Qwen-Image Technical Report
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 921dab5b-a36c-4778-adfb-20bfbf7fbcb0 · outbound
Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5c35c0d9-e5c1-416d-b475-ff16dc69a0a5 · outbound
Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation Reconstruction Alignment Improves Unified Multimodal Models
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2d4ec96d-b77a-4d76-ac64-d13e0329d4fb · outbound
Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation Show-o2: Improved Native Unified Multimodal Models
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a0811705-ddff-44b3-bd81-7a3f14c18916 · outbound
Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language Models
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9e3482fc-0701-46c0-8218-8693526b6cbe · outbound
Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation MMaDA: Multimodal Large Diffusion Language Models
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1dd6f632-2917-4b90-9e78-3438c90654cb · outbound
Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation ImgEdit: A Unified Image Editing Dataset and Benchmark
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e54215e6-0e04-43d0-937d-becee4f02a46 · outbound
Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation Llada-o: An effective and length-adaptive omni diffusion model.arXiv preprint arXiv:2603.01068
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0d4639ad-a10a-4ed7-bc50-b533bc33d46a · outbound
Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1a465710-521a-472a-8533-a8a4f9e8c82f · outbound
Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation cf1e5e24-eef8-46fd-9f45-268c9e4766cd · outbound
Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation PixelDiT: Pixel Diffusion Transformers for Image Generation
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation dce48e95-9d8f-475e-95de-a077b01672f5 · outbound
Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation Uniflow: A unified pixel flow tokenizer for visual understanding and generation.arXiv preprint arXiv:2510.10575
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6e715181-b3df-4444-bf5c-054d92fbe1bd · outbound
Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation Penguin-vl: Exploring the efficiency limits of vlm with llm-based vision encoders.arXiv preprint arXiv:2603.06569, 2026a
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a4cd0c68-afba-4977-9d8b-f5f60c3df444 · outbound
Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation Diffusion Transformers with Representation Autoencoders
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8c462b17-eef0-4708-bb87-b3f03c431587 · outbound
Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c276a854-1ceb-4b4a-9443-5f50ba344815 · outbound
Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation Scaling zero-shot reference-to-video generation.arXiv preprint arXiv:2512.06905
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5cd4c18f-214b-4a12-8456-b454291dc81c · inbound
STARFlow2: Bridging Language Models and Normalizing Flows for Unified Multimodal Generation Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 79cb73b9-d6e8-4043-be91-48f40f65f8aa · inbound
Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f919d687-9f33-4bbc-91b7-f97d15e76730 · inbound
Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 7e87a0de-826d-44e1-b9fd-4df531930344 · inbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation eecd669e-9280-4a33-9e24-cb7ad660f112 · inbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1841f1f9-3efe-45ae-ad77-4ed7a70d1b87 · inbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ecb394ca-831f-4321-9ff1-79a7ca9d4c73 · inbound
Toward Native Multimodal Modeling: A Roadmap Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation aaae50c0-09d6-4294-91d0-c03f773b6f88 · inbound
How and What to Imagine? Visual Thinking in Unified Multimodal Models for Cross-View Spatial Reasoning Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 057076fe-1fe7-45c2-b07d-830f1064b729 · inbound
Representation Forcing for Bottleneck-Free Unified Multimodal Models Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3fd2ec9a-f921-45db-a224-2f68be5cd6e7 · inbound
Representation Forcing for Bottleneck-Free Unified Multimodal Models Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce8add9b-090a-4cf1-9c68-236101331b6c · inbound
LLM can Read Spectrogram: Encoder-free Speech-Language Modeling Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b55fc75b-67a4-49dd-a29b-6a47fc00ccde · inbound
LLM can Read Spectrogram: Encoder-free Speech-Language Modeling Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8c24bcf8-cb8a-4126-9c27-970d8c9c3642 · inbound
IV-CoT: Implicit Visual Chain-of-Thought for Structure-Aware Text-to-Image Generation Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 08c59d78-89a3-47f3-89a3-027c525b7e77 · inbound
GEAR: Guided End-to-End AutoRegression for Image Synthesis Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 10f0d88f-3a7f-4a92-9a22-8170a221360b · inbound
APIVOT: Adaptive Planning with Interleaved Vision-Language Thoughts Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e72af9cb-8b76-4ba0-9cf0-57fbbfa3ebad · inbound
STBridge: Shared-Target Alignment for Bridging Understanding and Generation in UMMs Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 376e13c2-d6df-4e2b-b34a-86dc7b3a7e44 · inbound
Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.