Pith. sign in

Paper Citation Record · LEDGER

UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths

As of 9 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 4 inbound Pith citation observations for arXiv:2502.06474.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.06474 v1

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T15:24:46.543451Z

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:53:24.588515Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T15:34:48.324415Z

Reference resolution

45 of 45 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved45
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c63fb4c8-4bfb-409d-8c48-ec1120092cfc · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T15:24:46.275978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:24:46.275978Z digest=sha256:acb0d1c2aefe646824ad2bc07a1be39f50e941e61e99be8cf40c8be54a3655b2

Observation 0e33dca7-c4f7-4387-b784-8a74b2954b88 · outbound

This paper cites PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis.

UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T15:24:46.302202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:24:46.302202Z digest=sha256:dc99c50868cf3e353fe33295c239f52287c182c0eecadc159d6acfcef6d19a55

Observation de9dc85a-f968-4cb6-b57d-8584dee79dc9 · outbound

This paper cites LayerSkip: Enabling Early Exit Inference and Self-Speculative Decoding.

UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths LayerSkip: Enabling Early Exit Inference and Self-Speculative Decoding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T15:24:46.318344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:24:46.318344Z digest=sha256:e89d934031ea034fdbbd428149cdba42c87b783a1acbf0ed03168fec9cad1ffe

Observation 825f9edd-76e2-419d-83ff-c7084422ff30 · outbound

This paper cites SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation.

UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T15:24:46.328337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:24:46.328337Z digest=sha256:8174fb74a4a67bf5c3f01ee4e5c3c48a95f908e06fb1141e4e1dcaf074a3f27d

Observation a011c63a-f74a-4271-b87b-80e6fca01e69 · outbound

This paper cites Auto-Encoding Variational Bayes.

UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths Auto-Encoding Variational Bayes

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T15:24:46.344410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:24:46.344410Z digest=sha256:5456cfcf00ca829e35855d8be91505d32f4ee937b13b08cf7e5a270e7c76df9d

Observation 7c51bf4c-266a-46f7-ae10-861425d6a9b8 · outbound

This paper cites VideoPoet: A Large Language Model for Zero-Shot Video Generation.

UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T15:24:46.349114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:24:46.349114Z digest=sha256:6e52e775c18731d4cf4e70c0d5b2680afbcca5fc5007c0fe422835261befa48d

Observation 26df7ed4-86df-468d-a80a-8949774e4c65 · outbound

This paper cites Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads.

UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T15:24:46.353838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:24:46.353838Z digest=sha256:5966112ae0ea69a55fa3a58bc04a3448809eaf6f6ffc2ae39ccf7986874d8c0c

Observation cdd7047f-2df8-4500-8677-8ba2db133a57 · outbound

This paper cites JanusFlow: Harmonizing Autoregression and Rectified Flow for Unified Multimodal Understanding and Generation.

UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths JanusFlow: Harmonizing Autoregression and Rectified Flow for Unified Multimodal Understanding and Generation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T15:24:46.363865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:24:46.363865Z digest=sha256:d26c585e9a0b360ab5eec4e08e273e5fb2bac1051a0bf991622837b7adb6110c

Observation 095f80cc-32f8-479a-a231-78baef1b8778 · outbound

This paper cites MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training.

UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T15:24:46.368796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:24:46.368796Z digest=sha256:3825cba1bb60e28be0b13da784a8a392d96cb3ebf9d2b1fa04678710c2ebbb31

Observation 4c1aa551-d1e4-4f0c-8e63-5844c4dd034d · outbound

This paper cites GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models.

UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T15:24:46.373451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:24:46.373451Z digest=sha256:8f0e9c8b23fb6157c9a575499d6cd427866a331b349717ab1841e07f067e8b99

Observation bde2103b-7a1e-4b64-a249-5eb81122764f · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T15:24:46.377880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:24:46.377880Z digest=sha256:0d9a8e6180276dba3285fc2524b4faef777d88b33989f573e7fe4a0e23bd1465

Observation 53951b5c-6db3-46ca-b28d-b82e8bb92df1 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T15:24:46.382581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:24:46.382581Z digest=sha256:5ac86fdb1cf4f251e7af5fb5f85b6cd65e6857efad68cd20cafec8cc523a4fa2

Observation 88e622de-3e0f-4eca-bf68-b219530fdf91 · outbound

This paper cites TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation.

UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T15:24:46.387548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:24:46.387548Z digest=sha256:427ca5ddfef135268f6492eed53ea397945e5ace55be2078f2924d7b60824568

Observation cc2ca83d-6672-4de8-bb7f-ca53b3d44a5e · outbound

This paper cites an unresolved cited work.

UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-08T15:24:47.354874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:24:46.392401Z digest=sha256:c8c3a08b8bf14385f9172a49b42d78291f3887960c42a143fc5b90259f87f5e7

Observation c4c028e0-6882-4d48-9cac-46a3517af7cf · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T15:24:46.397246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:24:46.397246Z digest=sha256:c7ac459de7741f9e6b2a1a15d7cbfb9e6eb14bcdd3d02880c69418e98804c4ba

Observation cf6657fd-d1cd-4d8e-948c-c77115c146c1 · outbound

This paper cites LazyDiT: Lazy Learning for the Acceleration of Diffusion Transformers.

UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths LazyDiT: Lazy Learning for the Acceleration of Diffusion Transformers

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T15:24:46.402324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:24:46.402324Z digest=sha256:94002dbe897067d57a9e5d7eff2d9e5e28aee61db1ca867a30e299eb44cbd451

Observation 4281b462-436e-40f7-899c-6eae8f1ce502 · outbound

This paper cites LMFusion: Adapting Pretrained Language Models for Multimodal Generation.

UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T15:24:46.407707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:24:46.407707Z digest=sha256:268b95c6b8720430815ad5509087c194219fa70d4af9101a627e644cc9edb8ae

Observation b473c12a-d2f8-4217-8a33-c9525174e64f · outbound

This paper cites Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation.

UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T15:24:46.412515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:24:46.412515Z digest=sha256:dfe9824c9f6588959c05ac6c832e283cb1bbefeafde47a11ce7cd0c881d9721f

Observation 110289f0-5ca2-4927-8b8a-e23636be0980 · outbound

This paper cites Generative Multimodal Models are In-Context Learners.

UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths Generative Multimodal Models are In-Context Learners

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T15:24:46.417390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:24:46.417390Z digest=sha256:c912432a39d34662204b39d81c2633ea251c253fd550e2a342d0c7299e7c49da

Observation 79059f33-c100-4d1b-b7b4-d571d6ca9bec · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T15:24:46.421942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:24:46.421942Z digest=sha256:5d9f95eb8f3d3e3ba73c06b38f93704a7d2cc8b6e27f3a51603ff46780fa21df

Observation c148cd9b-4ffb-49af-bc49-4b66fc93056e · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.

UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T15:24:46.426860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:24:46.426860Z digest=sha256:58fc7f1dd21897db07c71b262b9a6472ad208805ccd5a991e1706161af2b8f4e

Observation 1df2559a-1339-48ff-be7d-6d756590c2fb · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths Emu3: Next-Token Prediction is All You Need

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-08T15:24:46.436695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:24:46.436695Z digest=sha256:b7d0e37ae70f4e30fb58c5279df27a3c2d8c80816ac656c37ced91a98e9692a6

Observation efd15f35-2117-4e19-8fed-ef327207ff6b · outbound

This paper cites Small-scale proxies for large-scale Transformer training instabilities.

UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths Small-scale proxies for large-scale Transformer training instabilities

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T15:24:46.441487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:24:46.441487Z digest=sha256:cc8cd237f67388c350d2ef7a3fb516b7eb688e6a50525388a076b7a5e5bb1fdc

Observation 8dd074de-e5b9-4644-b9b5-bd5ae90f6f6a · outbound

This paper cites Liquid: Language Models are Scalable and Unified Multi-modal Generators.

UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths Liquid: Language Models are Scalable and Unified Multi-modal Generators

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-08T15:24:46.446643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:24:46.446643Z digest=sha256:3dd79ed23a183648a033e3a67aa4039dd6726a2970ca6a9397cdf00bacd5e4a1

Observation fbf77a66-4013-4097-8e57-66abf718d4dc · outbound

This paper cites Show-o: One Single Transformer to Unify Multimodal Understanding and Generation.

UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-08T15:24:46.451544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:24:46.451544Z digest=sha256:50837793d1ea7fac6212387dac3b7ba75e3dd8507702dc7aaf078ecc5225ca77

Observation b003c0d9-950a-40cd-b724-cc0b7cff4dba · outbound

This paper cites SEED-Story: Multimodal Long Story Generation with Large Language Model.

UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths SEED-Story: Multimodal Long Story Generation with Large Language Model

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-08T15:24:46.456245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:24:46.456245Z digest=sha256:28940d566913b57d44ff510892922bc9352a1466f62b59badfd07012bfd81d77

Observation e1bfa684-4684-409b-bf2e-c390f353b7fb · outbound

This paper cites X-VILA: Cross-Modality Alignment for Large Language Model.

UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths X-VILA: Cross-Modality Alignment for Large Language Model

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T15:24:46.461321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:24:46.461321Z digest=sha256:92e114eb09d155a48e5e8b0544179d0a150472da8ef9427051696beb835149b0

Observation eee958a3-ce80-4dc1-aa91-877675b86481 · outbound

This paper cites Scaling Autoregressive Models for Content-Rich Text-to-Image Generation.

UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-08T15:24:46.475345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:24:46.475345Z digest=sha256:957d1ed3979dd7666ea725496e073b18b83c7eb233c617d2dce03c4160e43841

Observation 6504452a-31dd-4d1c-927c-df0c3bfc2c1d · outbound

This paper cites Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation.

UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-08T15:24:46.486047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:24:46.486047Z digest=sha256:7ec792ad6fb5b75910fa484dc00cb95ed91667e9fd04b7f308162c402359eee5

Observation 7bfe05f8-b076-4a0c-9394-a1f47bbda071 · outbound

This paper cites Learning to Skip for Language Modeling.

UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths Learning to Skip for Language Modeling

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-08T15:24:46.499373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:24:46.499373Z digest=sha256:714ffc1d993c289d0069c4955cd41725c6e49f206baeb7f08f40f59ec8626375

Observation 3e54c091-40f9-4a1f-8be6-6f7de5e88539 · outbound

This paper cites p-MoD: Building Mixture-of-Depths MLLMs via Progressive Ratio Decay.

UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths p-MoD: Building Mixture-of-Depths MLLMs via Progressive Ratio Decay

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-08T15:24:46.513648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:24:46.513648Z digest=sha256:987d10eb2613ba1467937d5074d2b53b9bb53f2449f33165c6cd8142efcc510a

Observation 900c9ef1-c8e9-4ba2-b157-c3db2f84643b · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-08T15:24:46.528792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:24:46.528792Z digest=sha256:65078b89caaf0283df3c46ca2b16b7e4a5123e4430429d2e3d96641281d4b0a0

Observation a46aff11-8af8-4309-a0b6-a77bf14fa7e3 · outbound

This paper cites an unresolved cited work.

UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-08T15:24:47.339079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:24:46.543451Z digest=sha256:bf07a8277ae571ea39e96317d5e788ff6ac4cef5d1e0fd51e6da0900bcff133d

Observation 0bec555d-1264-4718-b863-d5dba6560ffc · outbound

This paper cites SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding.

UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding

Reference 324

Resolution
unresolved
no resolver link, observed 2026-08-08T15:24:46.358407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:24:46.358407Z digest=sha256:74482d98d3cbb1e7c65d0d925cf677ecc3a0e7563f9e7dfdb7f8f5db21fcb23d

Observation 8af2fb02-5c1d-4d85-849f-6d1225605dad · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths Gemini: A Family of Highly Capable Multimodal Models

Reference 2004

Resolution
unresolved
no resolver link, observed 2026-08-08T15:24:46.281260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:24:46.281260Z digest=sha256:952a8f6e3ed4212f22272d25d46323a62b2967f82d0b18bf327cb815c636f169

Observation 05cdec10-e926-4ce6-8263-4117add83473 · outbound

This paper cites Classifier-Free Diffusion Guidance.

UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths Classifier-Free Diffusion Guidance

Reference 2006

Resolution
unresolved
no resolver link, observed 2026-08-08T15:24:46.334262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:24:46.334262Z digest=sha256:a4495787bf88521e08710bcfd7a5f628141cb1c1179b8b4a89c574d7a658feea

Observation 482684df-b832-4f66-9fa5-b741493bf46f · outbound

This paper cites DreamLLM: Synergistic Multimodal Comprehension and Creation.

UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths DreamLLM: Synergistic Multimodal Comprehension and Creation

Reference 2009

Resolution
unresolved
no resolver link, observed 2026-08-08T15:24:46.313275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:24:46.313275Z digest=sha256:7ba79b1b0a594e45fe6f4adbff923054ed5ec7506ce1e8e7283b7482a05eda95

Observation 36281543-3e94-4460-ba97-80e8278b466f · outbound

This paper cites Vector-quantized Image Modeling with Improved VQGAN.

UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths Vector-quantized Image Modeling with Improved VQGAN

Reference 2014

Resolution
unresolved
no resolver link, observed 2026-08-08T15:24:46.466171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:24:46.466171Z digest=sha256:1edd3a80cdcea19dc364eb1e3f6e736df7f499a66b0514e40f539b1c93b766e7

Observation e6010ea1-63ee-4612-8273-cfa1f1b8418c · outbound

This paper cites BERT has a Mouth, and It Must Speak: BERT as a Markov Random Field Language Model.

UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths BERT has a Mouth, and It Must Speak: BERT as a Markov Random Field Language Model

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-08T15:24:46.431621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:24:46.431621Z digest=sha256:dbb8b0028f1a248cea8c29a9a5d0450053885ee2b1d803f577e84120bcba100f

Observation 7f680462-8a03-4afc-8fdf-40674abcd6c3 · outbound

This paper cites Mixtral of Experts.

UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths Mixtral of Experts

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-08T15:24:46.339298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:24:46.339298Z digest=sha256:e14f19deddb56c98f91397560419090613f7882591afb4f3899f7eb09780dc47

Observation 5af0c397-f79e-465e-be57-6da0bd3d65ed · outbound

This paper cites A Survey on Mixture of Experts in Large Language Models.

UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths A Survey on Mixture of Experts in Large Language Models

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-08T15:24:46.292018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:24:46.292018Z digest=sha256:359c55c6927165d6eb297843c6847ccc70590e8a165eeec4744f2b984d93f83c

Observation 683d76c9-e5f8-4aa1-a621-f85dc2845bbd · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-08T15:24:46.286665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:24:46.286665Z digest=sha256:0f206d53b21168da5dde2f87c64209286dce6d7a5a4e8727fa788a928b8c7f61

Observation e226a789-8427-4973-bcc9-0d1358c5bb95 · outbound

This paper cites Muse: Text-To-Image Generation via Masked Generative Transformers.

UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-08T15:24:46.297104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:24:46.297104Z digest=sha256:03c1e0506b460ef8bfce74b1232ceba049cce57cadee233bc78adab57d60162a

Observation 9ab3af2b-2f8d-4b4c-bcae-ff369ea093b4 · outbound

This paper cites SkipDecode: Autoregressive Skip Decoding with Batching and Caching for Efficient LLM Inference.

UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths SkipDecode: Autoregressive Skip Decoding with Batching and Caching for Efficient LLM Inference

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-08T15:24:46.308313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:24:46.308313Z digest=sha256:59a8ece7b53fdef8e78386468af97743e4f515254fb231a40c78796c50899103

Observation 7d5f93f8-5a0e-40df-8127-edb02d113fb7 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-08T15:24:46.323219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:24:46.323219Z digest=sha256:75b501e0cf292071b787475c828a0fab819ff8e6e1b229355512cf18b81c38f3

Pith citing papers

Observation abbcd6f8-fb9e-4226-bd62-53b52b9d3f5c · inbound

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning cites this paper.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:24.588515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:24.588515Z digest=sha256:0905b0942c0d92294bdc8cefc5bfcf095750d83a3c98b315d6b4de6f34226177

Observation 58fe0e7f-27ae-411e-a5af-684a1867b2ce · inbound

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models cites this paper.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:42:21.515237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T05:37:36.329835Z digest=sha256:5493b5ade155162013440051ea8c245e2b2679fbcd60ddd8ec0c6a27d0f7ea66

Observation 26e843a3-6b1e-45bc-9c71-5c2a53ef71b9 · inbound

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models cites this paper.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-19T16:57:40.364921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:e08e54f2715a58b7fe9ec823333704fc3bce0443a82412edcbd4e199bd1d3878

Observation 40948551-860f-41bc-9405-333c4030406e · inbound

Tackling Multimodal Learning Challenges with Mixture-of-Expert: A Survey cites this paper.

Tackling Multimodal Learning Challenges with Mixture-of-Expert: A Survey UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-06-30T15:34:48.325959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T15:26:44.904121Z digest=sha256:466a91b6157cb2a3cd84bb02e3cb7123eacd5f463b15fd73b053f3b8b32af2dd