Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T15:24:46.543451Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 4 inbound Pith citation observations for arXiv:2502.06474.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T15:24:46.543451Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:53:24.588515Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-06-30T15:34:48.324415Z
45 of 45 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation c63fb4c8-4bfb-409d-8c48-ec1120092cfc · outbound
UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e33dca7-c4f7-4387-b784-8a74b2954b88 · outbound
UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de9dc85a-f968-4cb6-b57d-8584dee79dc9 · outbound
UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths LayerSkip: Enabling Early Exit Inference and Self-Speculative Decoding
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 825f9edd-76e2-419d-83ff-c7084422ff30 · outbound
UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a011c63a-f74a-4271-b87b-80e6fca01e69 · outbound
UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths Auto-Encoding Variational Bayes
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c51bf4c-266a-46f7-ae10-861425d6a9b8 · outbound
UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths VideoPoet: A Large Language Model for Zero-Shot Video Generation
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26df7ed4-86df-468d-a80a-8949774e4c65 · outbound
UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cdd7047f-2df8-4500-8677-8ba2db133a57 · outbound
UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths JanusFlow: Harmonizing Autoregression and Rectified Flow for Unified Multimodal Understanding and Generation
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 095f80cc-32f8-479a-a231-78baef1b8778 · outbound
UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c1aa551-d1e4-4f0c-8e63-5844c4dd034d · outbound
UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bde2103b-7a1e-4b64-a249-5eb81122764f · outbound
UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths Kosmos-2: Grounding Multimodal Large Language Models to the World
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53951b5c-6db3-46ca-b28d-b82e8bb92df1 · outbound
UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88e622de-3e0f-4eca-bf68-b219530fdf91 · outbound
UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc2ca83d-6672-4de8-bb7f-ca53b3d44a5e · outbound
UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths Unresolved cited work
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c4c028e0-6882-4d48-9cac-46a3517af7cf · outbound
UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths Hierarchical Text-Conditional Image Generation with CLIP Latents
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf6657fd-d1cd-4d8e-948c-c77115c146c1 · outbound
UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths LazyDiT: Lazy Learning for the Acceleration of Diffusion Transformers
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4281b462-436e-40f7-899c-6eae8f1ce502 · outbound
UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths LMFusion: Adapting Pretrained Language Models for Multimodal Generation
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b473c12a-d2f8-4217-8a33-c9525174e64f · outbound
UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 110289f0-5ca2-4927-8b8a-e23636be0980 · outbound
UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths Generative Multimodal Models are In-Context Learners
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79059f33-c100-4d1b-b7b4-d571d6ca9bec · outbound
UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths Chameleon: Mixed-Modal Early-Fusion Foundation Models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c148cd9b-4ffb-49af-bc49-4b66fc93056e · outbound
UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1df2559a-1339-48ff-be7d-6d756590c2fb · outbound
UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths Emu3: Next-Token Prediction is All You Need
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation efd15f35-2117-4e19-8fed-ef327207ff6b · outbound
UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths Small-scale proxies for large-scale Transformer training instabilities
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8dd074de-e5b9-4644-b9b5-bd5ae90f6f6a · outbound
UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths Liquid: Language Models are Scalable and Unified Multi-modal Generators
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fbf77a66-4013-4097-8e57-66abf718d4dc · outbound
UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths Show-o: One Single Transformer to Unify Multimodal Understanding and Generation
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b003c0d9-950a-40cd-b724-cc0b7cff4dba · outbound
UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths SEED-Story: Multimodal Long Story Generation with Large Language Model
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1bfa684-4684-409b-bf2e-c390f353b7fb · outbound
UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths X-VILA: Cross-Modality Alignment for Large Language Model
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eee958a3-ce80-4dc1-aa91-877675b86481 · outbound
UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths Scaling Autoregressive Models for Content-Rich Text-to-Image Generation
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6504452a-31dd-4d1c-927c-df0c3bfc2c1d · outbound
UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7bfe05f8-b076-4a0c-9394-a1f47bbda071 · outbound
UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths Learning to Skip for Language Modeling
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e54c091-40f9-4a1f-8be6-6f7de5e88539 · outbound
UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths p-MoD: Building Mixture-of-Depths MLLMs via Progressive Ratio Decay
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 900c9ef1-c8e9-4ba2-b157-c3db2f84643b · outbound
UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a46aff11-8af8-4309-a0b6-a77bf14fa7e3 · outbound
UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths Unresolved cited work
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 0bec555d-1264-4718-b863-d5dba6560ffc · outbound
UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding
Reference 324
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8af2fb02-5c1d-4d85-849f-6d1225605dad · outbound
UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths Gemini: A Family of Highly Capable Multimodal Models
Reference 2004
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05cdec10-e926-4ce6-8263-4117add83473 · outbound
UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths Classifier-Free Diffusion Guidance
Reference 2006
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 482684df-b832-4f66-9fa5-b741493bf46f · outbound
UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths DreamLLM: Synergistic Multimodal Comprehension and Creation
Reference 2009
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36281543-3e94-4460-ba97-80e8278b466f · outbound
UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths Vector-quantized Image Modeling with Improved VQGAN
Reference 2014
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6010ea1-63ee-4612-8273-cfa1f1b8418c · outbound
UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths BERT has a Mouth, and It Must Speak: BERT as a Markov Random Field Language Model
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f680462-8a03-4afc-8fdf-40674abcd6c3 · outbound
UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths Mixtral of Experts
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5af0c397-f79e-465e-be57-6da0bd3d65ed · outbound
UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths A Survey on Mixture of Experts in Large Language Models
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 683d76c9-e5f8-4aa1-a621-f85dc2845bbd · outbound
UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e226a789-8427-4973-bcc9-0d1358c5bb95 · outbound
UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths Muse: Text-To-Image Generation via Masked Generative Transformers
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ab3af2b-2f8d-4b4c-bcae-ff369ea093b4 · outbound
UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths SkipDecode: Autoregressive Skip Decoding with Batching and Caching for Efficient LLM Inference
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d5f93f8-5a0e-40df-8127-edb02d113fb7 · outbound
UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation abbcd6f8-fb9e-4226-bd62-53b52b9d3f5c · inbound
UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58fe0e7f-27ae-411e-a5af-684a1867b2ce · inbound
G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 26e843a3-6b1e-45bc-9c71-5c2a53ef71b9 · inbound
G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 40948551-860f-41bc-9405-333c4030406e · inbound
Tackling Multimodal Learning Challenges with Mixture-of-Expert: A Survey UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.