Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-12T15:31:57.426559Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 3 inbound Pith citation observations for arXiv:2605.31604.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-12T15:31:57.426559Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-01T05:45:51.915796Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-03T23:59:06.123206Z
65 of 65 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation ee98f9b5-ddb5-4680-8fd8-3fb7b7d0803a · outbound
Representation Forcing for Bottleneck-Free Unified Multimodal Models Latent forcing: Reordering the diffusion trajectory for pixel-space image generation.arXiv preprint arXiv:2602.11401, 2026
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf14e559-e5a4-4189-b7be-946ff2e5f16e · outbound
Representation Forcing for Bottleneck-Free Unified Multimodal Models Improving image generation with better captions.OpenAI Technical Report,https: // cdn
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf610772-11ae-44d0-bbc5-116cb59b4169 · outbound
Representation Forcing for Bottleneck-Free Unified Multimodal Models FLUX.https://github.com/black-forest-labs/flux, 2024
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79f71b40-2ea1-4a39-8f84-a480ba44c8ef · outbound
Representation Forcing for Bottleneck-Free Unified Multimodal Models Unsupervised learning of visual features by contrasting cluster assignments
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56d802b2-668a-405c-ad12-93d97afc94b1 · outbound
Representation Forcing for Bottleneck-Free Unified Multimodal Models Chameleon: Mixed-Modal Early-Fusion Foundation Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25fa5aff-909c-428d-b74a-cb1e89301b66 · outbound
Representation Forcing for Bottleneck-Free Unified Multimodal Models BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 447b46f2-0f61-4bf4-98bd-b11066588b35 · outbound
Representation Forcing for Bottleneck-Free Unified Multimodal Models PixArt-α: Fast training of diffusion transformer for photorealistic text-to-image synthesis
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d53b2b8-b652-48ba-9a3d-45cb3819a60d · outbound
Representation Forcing for Bottleneck-Free Unified Multimodal Models PixelFlow: Pixel-Space Generative Models with Flow
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0db96706-1a79-4764-93cd-539fa656f5d0 · outbound
Representation Forcing for Bottleneck-Free Unified Multimodal Models Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d335398d-8371-4596-bf76-40b3d80154dd · outbound
Representation Forcing for Bottleneck-Free Unified Multimodal Models Patch n’ Pack: NaViT, a vision transformer for any aspect ratio and resolution
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30a5e4ab-0b1a-46a2-a73a-bc2569e16f6c · outbound
Representation Forcing for Bottleneck-Free Unified Multimodal Models Emerging Properties in Unified Multimodal Pretraining
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4f94780-4672-4075-936e-3be7658574c4 · outbound
Representation Forcing for Bottleneck-Free Unified Multimodal Models Imagenet: A large-scale hierarchical image database
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8b9f36a-4054-4cf4-b61f-51706c7d3227 · outbound
Representation Forcing for Bottleneck-Free Unified Multimodal Models Diffusion models beat gans on image synthesis
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eec1c095-5b55-49ab-8e4e-e111cafe4944 · outbound
Representation Forcing for Bottleneck-Free Unified Multimodal Models SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a21bc18c-463d-4e49-bcd4-b73badce8fc5 · outbound
Representation Forcing for Bottleneck-Free Unified Multimodal Models Taming transformers for high-resolution image synthesis
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de88a2e8-26ff-40b1-9ccf-c70a500b27e8 · outbound
Representation Forcing for Bottleneck-Free Unified Multimodal Models Scaling rectified flow transformers for high-resolution image synthesis
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e14f7bf3-8f66-4a5b-85b2-2ce0bb697474 · outbound
Representation Forcing for Bottleneck-Free Unified Multimodal Models MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb484af2-cfc6-42d6-9414-5fafe928d9c4 · outbound
Representation Forcing for Bottleneck-Free Unified Multimodal Models Smith, Wei-Chiu Ma, and Ranjay Krishna
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 718affd4-fede-46e9-bdcc-0e395725a9a3 · outbound
Representation Forcing for Bottleneck-Free Unified Multimodal Models Seedream 3.0 Technical Report
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac8fcb53-ab2b-476c-bab2-ffa7d099b020 · outbound
Representation Forcing for Bottleneck-Free Unified Multimodal Models SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a601bf0-b870-4cfc-b0e4-8e260a6ddac7 · outbound
Representation Forcing for Bottleneck-Free Unified Multimodal Models GenEval: An object-focused framework for evaluating text-to-image alignment
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a814cb94-4f9f-4d2f-b4ba-5aeac8e4843d · outbound
Representation Forcing for Bottleneck-Free Unified Multimodal Models HallusionBench: An advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a7d2f00-d00b-47bb-8e01-b8369cc96e3a · outbound
Representation Forcing for Bottleneck-Free Unified Multimodal Models Denoising diffusion probabilistic models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4636529-06f0-4b30-82be-b768c435c7ca · outbound
Representation Forcing for Bottleneck-Free Unified Multimodal Models Simpler diffusion (SiD2): 1.5 FID on ImageNet512 with pixel-space diffusion
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c01d60d0-3499-421c-a9b4-401464a2dd2d · outbound
Representation Forcing for Bottleneck-Free Unified Multimodal Models ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9fb6db5a-f2d2-4a8f-89a8-d34761787717 · outbound
Representation Forcing for Bottleneck-Free Unified Multimodal Models A diagram is worth a dozen images
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4811a82a-5758-450f-a856-b6b93d408c76 · outbound
Representation Forcing for Bottleneck-Free Unified Multimodal Models Auto-Encoding Variational Bayes
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67c35343-616c-4e20-9071-2d5de573cb16 · outbound
Representation Forcing for Bottleneck-Free Unified Multimodal Models Back to Basics: Let Denoising Generative Models Denoise
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a6fff01-7a55-423e-ac93-0548897990aa · outbound
Representation Forcing for Bottleneck-Free Unified Multimodal Models UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ba1cf97-7cc6-4897-9298-0e9016bd9db9 · outbound
Representation Forcing for Bottleneck-Free Unified Multimodal Models World Model on Million-Length Video And Language With Blockwise RingAttention
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3fd2ec9a-f921-45db-a224-2f68be5cd6e7 · outbound
Representation Forcing for Bottleneck-Free Unified Multimodal Models Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e76227a-78c1-416b-82bc-04b4366fe459 · outbound
Representation Forcing for Bottleneck-Free Unified Multimodal Models Decoupled weight decay regularization
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eeeff4de-3610-4171-b3a1-0c16591adb5a · outbound
Representation Forcing for Bottleneck-Free Unified Multimodal Models JanusFlow: Harmonizing autoregression and rectified flow for unified multimodal understanding and generation
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91b641b2-8468-4a4b-a317-9af22353db9e · outbound
Representation Forcing for Bottleneck-Free Unified Multimodal Models ChartQA: A benchmark for question answering about charts with visual and logical reasoning
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c9a5c00-a365-4c3e-89fc-85f4d9e504a0 · outbound
Representation Forcing for Bottleneck-Free Unified Multimodal Models DocVQA: A Dataset for VQA on Document Images
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37c5734d-1b23-4b88-aa57-a2b6e51cc4ac · outbound
Representation Forcing for Bottleneck-Free Unified Multimodal Models Unresolved cited work
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ea7e89a-6461-42b2-a2ef-8686c26b6ca3 · outbound
Representation Forcing for Bottleneck-Free Unified Multimodal Models Transfer between Modalities with MetaQueries
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67dcfd7d-4c5d-4302-a93e-d42d9bc1a14d · outbound
Representation Forcing for Bottleneck-Free Unified Multimodal Models SDXL: Improving latent diffusion models for high-resolution image synthesis
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2952c8b-a3dd-46e8-8661-f7d6ecac24fd · outbound
Representation Forcing for Bottleneck-Free Unified Multimodal Models Du, Zehuan Yuan, and Xinglong Wu
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c2ba7e0-ee72-4f80-b79f-6154a7ac1ace · outbound
Representation Forcing for Bottleneck-Free Unified Multimodal Models Hierarchical Text-Conditional Image Generation with CLIP Latents
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a187d1e-824f-4c04-b8ac-fee29126b089 · outbound
Representation Forcing for Bottleneck-Free Unified Multimodal Models High-resolution image synthesis with latent diffusion models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6f43eeb-99ce-4e5a-8ad5-2266efe89ae6 · outbound
Representation Forcing for Bottleneck-Free Unified Multimodal Models Latent diffusion model without variational autoencoder.arxiv: 2510.15301, 2025
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b414f980-526f-4589-83dc-c9fd71210dbb · outbound
Representation Forcing for Bottleneck-Free Unified Multimodal Models DINOv3
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c96917a-63ba-45b4-b2bb-433680bbc4f1 · outbound
Representation Forcing for Bottleneck-Free Unified Multimodal Models Generative multimodal models are in-context learners
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5109de35-9c1f-482e-88e6-7ff8e1a1b3da · outbound
Representation Forcing for Bottleneck-Free Unified Multimodal Models SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6bf52a3e-9c07-4463-ad59-885d32286853 · outbound
Representation Forcing for Bottleneck-Free Unified Multimodal Models Neural discrete representation learning
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5cc0082-fe85-46ef-9135-196aa48d17aa · outbound
Representation Forcing for Bottleneck-Free Unified Multimodal Models ILLUME: Illuminating your LLMs to see, draw, and self-enhance
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3963e05a-db2e-43bd-9c65-33860daaead3 · outbound
Representation Forcing for Bottleneck-Free Unified Multimodal Models PixNerd: Pixel Neural Field Diffusion
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1cadfe5-22f6-4d5e-bd6b-b46cf8b75fdd · outbound
Representation Forcing for Bottleneck-Free Unified Multimodal Models Emu3: Next-Token Prediction is All You Need
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f76f380-4e7a-4895-b126-660d9eb3c36a · outbound
Representation Forcing for Bottleneck-Free Unified Multimodal Models Cubic discrete diffusion: Discrete visual generation on high-dimensional representation tokens
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a48579d-3f14-4131-9451-fec5fa3acfdc · outbound
Representation Forcing for Bottleneck-Free Unified Multimodal Models Wan: Open and Advanced Large-Scale Video Generative Models
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df4b6c76-3515-4b4c-9fa5-74e03582d2c2 · outbound
Representation Forcing for Bottleneck-Free Unified Multimodal Models Qwen-Image Technical Report
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae0397fa-b5f3-465a-a96c-f989e40ee16d · outbound
Representation Forcing for Bottleneck-Free Unified Multimodal Models Janus: Decoupling visual encoding for unified multimodal understanding and generation
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bead864d-e8cf-4f43-97e8-96de1d0b857a · outbound
Representation Forcing for Bottleneck-Free Unified Multimodal Models OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8d0f5fe-c403-4420-bd67-7b0e4a4102be · outbound
Representation Forcing for Bottleneck-Free Unified Multimodal Models RealWorldQA.https://huggingface.co/datasets/xai-org/RealworldQA, 2024
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1302f11a-3c2e-4776-9e08-48ca82a1c58b · outbound
Representation Forcing for Bottleneck-Free Unified Multimodal Models Show-o: One single transformer to unify multimodal understanding and generation
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d952adf4-9e74-4cae-865f-abd790f17f38 · outbound
Representation Forcing for Bottleneck-Free Unified Multimodal Models Show-o2: Improved native unified multimodal models
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6da44f9a-1181-4959-9684-c3f56eb86c31 · outbound
Representation Forcing for Bottleneck-Free Unified Multimodal Models Qwen3 Technical Report
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24966641-45da-4f6d-8c10-5b702c88dc4f · outbound
Representation Forcing for Bottleneck-Free Unified Multimodal Models Context Unrolling in Omni Models
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7509791-14da-459b-a708-57f8acfdcbd4 · outbound
Representation Forcing for Bottleneck-Free Unified Multimodal Models Representation alignment for generation: Training diffusion transformers is easier than you think
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2bdf6708-a946-430a-ab16-e1b28be01d33 · outbound
Representation Forcing for Bottleneck-Free Unified Multimodal Models MMMU: A massive multi-discipline multimodal understanding and reasoning benchmark for expert AGI
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ea2d78c-ba30-4fbc-9249-11a3c2607b33 · outbound
Representation Forcing for Bottleneck-Free Unified Multimodal Models Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46e97e8e-bb93-4246-86e7-a13d4615fb6e · outbound
Representation Forcing for Bottleneck-Free Unified Multimodal Models Sigmoid loss for language image pre-training
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0336a190-c260-48d3-ad0b-83799df53e93 · outbound
Representation Forcing for Bottleneck-Free Unified Multimodal Models Diffusion Transformers with Representation Autoencoders
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6409544-364c-48c5-bb68-87be1b19612b · outbound
Representation Forcing for Bottleneck-Free Unified Multimodal Models Transfusion: Predict the next token and diffuse images with one multi-modal model
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7656482-5415-4733-9a24-0c270ba24f6f · inbound
LLM can Read Spectrogram: Encoder-free Speech-Language Modeling Representation Forcing for Bottleneck-Free Unified Multimodal Models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 921868c2-f87e-478e-92ee-4b377756fa90 · inbound
LLM can Read Spectrogram: Encoder-free Speech-Language Modeling Representation Forcing for Bottleneck-Free Unified Multimodal Models
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b199c5fd-609f-490d-93d7-a54484d3769d · inbound
dRAE: Representation Autoencoder with Hyper-Spherical Codes Representation Forcing for Bottleneck-Free Unified Multimodal Models
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.