Pith. sign in

Paper Citation Record · LEDGER

Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 55 inbound Pith citation observations for arXiv:2405.05945.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2405.05945 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 55 of 55 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:26:35.752294Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T16:59:58.189803Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation d816aaad-1368-4eb2-81e8-ddabfeea946b · inbound

PoM: Efficient Image and Video Generation with the Polynomial Mixer cites this paper.

PoM: Efficient Image and Video Generation with the Polynomial Mixer Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:47.036130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:47.036130Z digest=sha256:5844d953f2f5660cdf37bc2829f5ada7969341008bd95ae87dea997b3749f8c1

Observation 9e7b1660-f2e0-45cf-9b58-eb7aebe52921 · inbound

High-Resolution Image Synthesis via Next-Token Prediction cites this paper.

High-Resolution Image Synthesis via Next-Token Prediction Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T14:55:41.976670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:55:41.976670Z digest=sha256:d339ef3489ff45be4e52243d8ab4ffc86d15267ce1f0b1bd3c61fbeced8ae67d

Observation 9dd5c5e8-c743-4fee-a509-75eef8c6e838 · inbound

Towards Precise Scaling Laws for Video Diffusion Transformers cites this paper.

Towards Precise Scaling Laws for Video Diffusion Transformers Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T12:58:57.129116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:58:57.129116Z digest=sha256:b1c7018f2be6823f32400416ef6600a8ce9d41171e04737371f6553c36ce1ae6

Observation 9bc9becb-6b5c-43dd-8715-302a39da6d37 · inbound

AC3D: Analyzing and Improving 3D Camera Control in Video Diffusion Transformers cites this paper.

AC3D: Analyzing and Improving 3D Camera Control in Video Diffusion Transformers Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T11:07:39.497285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:07:39.497285Z digest=sha256:0284e487a588e47ab07892b91f6427f2e7f680939f0913ae9310f7a9ca25b377

Observation 9b23476b-6142-4793-9401-469a30043998 · inbound

Scaling Image Tokenizers with Grouped Spherical Quantization cites this paper.

Scaling Image Tokenizers with Grouped Spherical Quantization Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-11T23:20:11.067426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:20:11.067426Z digest=sha256:1c7ad6bf95da9fb81efd5e74404e2b7ae7aa6f6e82d8d2fcfeef3c223b8662b9

Observation 45b05fbf-8760-4267-bc57-45d5b7dfed17 · inbound

Mimir: Improving Video Diffusion Models for Precise Text Understanding cites this paper.

Mimir: Improving Video Diffusion Models for Precise Text Understanding Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:11.670823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:11.670823Z digest=sha256:bb3047c2e64c7151447353a98e47a5a4ddcf80e79873a8417411875c299f7553

Observation abf81c51-b893-4d32-ac0d-024683d3f1ca · inbound

Imagine360: Immersive 360 Video Generation from Perspective Anchor cites this paper.

Imagine360: Immersive 360 Video Generation from Perspective Anchor Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T22:23:15.450275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:23:15.450275Z digest=sha256:de01d97d36c87f3581f548d1690dbee65014ba65c1e02256fab627667a2f2c92

Observation 1a8be5e9-fde3-4620-85d7-e12ba2b9effc · inbound

HunyuanVideo: A Systematic Framework For Large Video Generative Models cites this paper.

HunyuanVideo: A Systematic Framework For Large Video Generative Models Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-23T07:42:43.487031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:6c8e271a3de8124844d4096d4f890919898d3a16bef7f6f585acc581878c7908

Observation 6a0b6586-9b77-499a-b583-d36255a26bf2 · inbound

CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation cites this paper.

CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T22:05:33.668005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:05:33.668005Z digest=sha256:ed45fd5e92ba883b9a05d4dea7fa27a32c7100e2adc127fe1f3e429b8cf64aa9

Observation b0514003-a314-4573-b676-3ebcc310c0e8 · inbound

Mind the Time: Temporally-Controlled Multi-Event Video Generation cites this paper.

Mind the Time: Temporally-Controlled Multi-Event Video Generation Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T20:55:01.928899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:55:01.928899Z digest=sha256:4b7fa9abdefcb0b93072500f33a2a8d99f86e418c8eb929d3ae8d752e6e53d16

Observation 8b34c3ec-711c-49a9-bd46-a26009a8df44 · inbound

MotionStone: Decoupled Motion Intensity Modulation with Diffusion Transformer for Image-to-Video Generation cites this paper.

MotionStone: Decoupled Motion Intensity Modulation with Diffusion Transformer for Image-to-Video Generation Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T20:20:36.493002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:20:36.493002Z digest=sha256:21460634a3508272705f093e2d38549b28064e2ba11ff213cfc9cf50428b0336

Observation f5969d1a-10eb-459c-8dac-467337f136ca · inbound

SnapGen: Taming High-Resolution Text-to-Image Models for Mobile Devices with Efficient Architectures and Training cites this paper.

SnapGen: Taming High-Resolution Text-to-Image Models for Mobile Devices with Efficient Architectures and Training Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:01.614056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:01.614056Z digest=sha256:1907aa6045ab52bf5a5afe5b81891bb80108a3c22e6d2ae54651d36dbd3e3ce5

Observation b5e1b2c8-5007-4e4c-a5a6-da577476013d · inbound

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation cites this paper.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T15:43:00.877098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:43:00.877098Z digest=sha256:73dbc8cff4daca8d9af7a8ffbd734b4a0f2fb943ef511f6ffbf0a84a8941193d

Observation 60c5fffb-27ba-46fc-ba1a-9f59906940ed · inbound

Video Diffusion Transformers are In-Context Learners cites this paper.

Video Diffusion Transformers are In-Context Learners Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:19.798497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:19.798497Z digest=sha256:42491a52bba18b28b4b7234d246c28f3415283fc62ce108596d21752849d24ff

Observation 24708269-1164-4a1c-8351-c328662ed805 · inbound

Causal Diffusion Transformers for Generative Modeling cites this paper.

Causal Diffusion Transformers for Generative Modeling Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T14:20:49.289133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:20:49.289133Z digest=sha256:6643594f838088bfdeef54e0e091a333b0153b67067a03ea52fae42cee443b76

Observation 0d33ef00-2426-439d-9090-5595f082775d · inbound

CLEAR: Conv-Like Linearization Revs Pre-Trained Diffusion Transformers Up cites this paper.

CLEAR: Conv-Like Linearization Revs Pre-Trained Diffusion Transformers Up Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T10:54:30.795049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:54:30.795049Z digest=sha256:712487135c397d9209e52e4d96c636f11907df41be52350ecb32c81ef643434e

Observation c5a220a6-cc78-4066-bc0d-e24c36358e57 · inbound

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement cites this paper.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:35.030224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:35.030224Z digest=sha256:e56b59095578ca464170067a598a1b7e7a14f39b791f46dd2148038423d18852

Observation 70ca07cd-d460-41aa-aafc-a56c2210b7c7 · inbound

Dual Diffusion for Unified Image Generation and Understanding cites this paper.

Dual Diffusion for Unified Image Generation and Understanding Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T23:02:07.643640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:02:07.643640Z digest=sha256:84b195cbdc026d318cd4ef921bb38dd737c1dea3f1b596c6a3f671d04ea98213

Observation 7f2264ac-3a17-4ecb-93cb-f017f9639c8d · inbound

Reconstruction vs. Generation: Taming Optimization Dilemma in Latent Diffusion Models cites this paper.

Reconstruction vs. Generation: Taming Optimization Dilemma in Latent Diffusion Models Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T22:34:14.232420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:34:14.232420Z digest=sha256:7a647af78b9449b32cf79077ab33d9957c66d1cb6a7a524140a48ab8febee947

Observation 01546f54-6fc4-4311-8884-8941a3753d23 · inbound

Ingredients: Blending Custom Photos with Video Diffusion Transformers cites this paper.

Ingredients: Blending Custom Photos with Video Diffusion Transformers Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T22:25:25.600000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:25:25.600000Z digest=sha256:2106303d8fc06a06a8e06104bef11eaed49d4b9f8ea971bfaac1413f30ad64a5

Observation 86cbbaa3-b068-402a-ae93-7afc7168ddb0 · inbound

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models cites this paper.

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T15:31:35.500088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:31:35.500088Z digest=sha256:3762ceb6ac041882468af5e26e712a2766ed4e913337b0250988e6e84f8b73ad

Observation 0e61a512-9c5d-48db-b995-acd0aeb50426 · inbound

DiffSplat: Repurposing Image Diffusion Models for Scalable Gaussian Splat Generation cites this paper.

DiffSplat: Repurposing Image Diffusion Models for Scalable Gaussian Splat Generation Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T10:56:30.925033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:56:30.925033Z digest=sha256:3f4bbbf56389631bf13e6a81d4963fd26ca13513d6dadd318a62178a0c61038b

Observation cf5811c7-1ecb-4db6-856d-72512fd18b65 · inbound

Lumina-Video: Efficient and Flexible Video Generation with Multi-scale Next-DiT cites this paper.

Lumina-Video: Efficient and Flexible Video Generation with Multi-scale Next-DiT Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T14:26:46.667065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:26:46.667065Z digest=sha256:cb20fdcbae832c66aaee4d04ab70de62344e3b91e857fbaa4bff974665218489

Observation f0b70c9f-20b8-4842-9c61-371730ec43de · inbound

Enhance-A-Video: Better Generated Video for Free cites this paper.

Enhance-A-Video: Better Generated Video for Free Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T12:32:46.422166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:32:46.422166Z digest=sha256:ab63d0f0fd7a96f76484f090669fbd0f1bc9a42b05a8374a9958ddec3f1d6969

Observation e6debad6-641c-428b-8ca2-12cf1f884dcf · inbound

RealCam-I2V: Real-World Image-to-Video Generation with Interactive Complex Camera Control cites this paper.

RealCam-I2V: Real-World Image-to-Video Generation with Interactive Complex Camera Control Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T19:40:21.192239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T19:40:21.192239Z digest=sha256:7506f17b039806598f1c9f4e46f171b059c64dde31e993083b1ac433c05880e8

Observation 222b390e-0123-4cce-af76-0225f43efd21 · inbound

TSGS: Improving Gaussian Splatting for Transparent Surface Reconstruction via Normal and De-lighting Priors cites this paper.

TSGS: Improving Gaussian Splatting for Transparent Surface Reconstruction via Normal and De-lighting Priors Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T12:26:35.752294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:26:35.752294Z digest=sha256:57d779bee76088536691554daebebe367d6353a7cf280025c4ecde01929adf34

Observation 54cce58d-49b5-405c-89f4-974dfe26e1e4 · inbound

T2S: High-resolution Time Series Generation with Text-to-Series Diffusion Models cites this paper.

T2S: High-resolution Time Series Generation with Text-to-Series Diffusion Models Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T00:57:28.802001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:57:28.802001Z digest=sha256:bef11e502099e40dfcff842ca740ee071867599e914bd70c095909d7dd9bc8d8

Observation 5d519a93-288e-437c-85e1-b0880a5f1302 · inbound

Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis cites this paper.

Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T21:21:22.351916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:21:22.351916Z digest=sha256:4185878a8b142bfe1922e6ae944f2d3e6584f66186ce6b496519d921eb737c04

Observation 7604101b-74f4-4077-a73b-b47f410605a5 · inbound

DRAGON: A Large-Scale Dataset of Realistic Images Generated by Diffusion Models cites this paper.

DRAGON: A Large-Scale Dataset of Realistic Images Generated by Diffusion Models Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T20:58:13.828764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:58:13.828764Z digest=sha256:6e1fc902d6e098ba1c5acb6c3058e50a795b0fde2469711debce335da5907ebf

Observation c5dc8b99-5a05-40bf-834a-a48566c6b0a6 · inbound

RetinaLogos: Fine-Grained Synthesis of High-Resolution Retinal Images Through Captions cites this paper.

RetinaLogos: Fine-Grained Synthesis of High-Resolution Retinal Images Through Captions Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T20:28:44.370098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:28:44.370098Z digest=sha256:d8724c4825d5504633836e222bccbb2e61865ba5cc8187cc79c53806d9883b75

Observation 369b99fb-2f79-4006-ae23-413dac14c3a7 · inbound

R2I-Bench: Benchmarking Reasoning-Driven Text-to-Image Generation cites this paper.

R2I-Bench: Benchmarking Reasoning-Driven Text-to-Image Generation Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:50.794369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:50.794369Z digest=sha256:2d7834edc05663424a8f1e54d1081bdd1f3ab14bc06c19ce979cf637d6a83117

Observation becaf3ad-2e97-42fc-9820-e757bec27070 · inbound

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation cites this paper.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.030846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.030846Z digest=sha256:5135ee3ddf354d6d395bf325fa8fff0f4cd9c1d2f2b422f152812f40e4fa4218

Observation 4e0b2401-591c-44d4-8424-d67df269a7a7 · inbound

The Hidden Cost of an Image: Quantifying the Energy Consumption of AI Image Generation cites this paper.

The Hidden Cost of an Image: Quantifying the Energy Consumption of AI Image Generation Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T19:17:12.904990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:17:12.904990Z digest=sha256:a5523df672399cadda39b7a62998e9495c583fd85436fb9f63d744b6669aa67e

Observation ab2614db-bee4-43c3-946a-a1eb435a82ed · inbound

Diffusion Transformer-to-Mamba Distillation for High-Resolution Image Generation cites this paper.

Diffusion Transformer-to-Mamba Distillation for High-Resolution Image Generation Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T18:44:38.319824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:44:38.319824Z digest=sha256:da2e21ba6021a9289b611b6f78b29f1f52ab4f87890fbfe5686988acd24d319b

Observation 68f82bf5-177a-43f2-bd05-a159abd0926b · inbound

Consistent Zero-shot 3D Texture Synthesis Using Geometry-aware Diffusion and Temporal Video Models cites this paper.

Consistent Zero-shot 3D Texture Synthesis Using Geometry-aware Diffusion and Temporal Video Models Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T22:41:38.549537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:41:38.549537Z digest=sha256:f819a5fe7848402083e049d3377922e8bfaa391aa2740f8e4554064125e4a8ee

Observation 507a892c-920f-4d4f-8c5d-79a184db2557 · inbound

UniMC: Taming Diffusion Transformer for Unified Keypoint-Guided Multi-Class Image Generation cites this paper.

UniMC: Taming Diffusion Transformer for Unified Keypoint-Guided Multi-Class Image Generation Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T20:26:37.854469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:26:37.854469Z digest=sha256:43ce3c1a473957a457f66270c1938fd4771b15e68573cf914346f7fc769b685e

Observation f36ad844-76c0-432f-88db-c24ea48c0fe3 · inbound

FreeCus: Free Lunch Subject-driven Customization in Diffusion Transformers cites this paper.

FreeCus: Free Lunch Subject-driven Customization in Diffusion Transformers Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T15:42:23.696320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:42:23.696320Z digest=sha256:7c6eddffa1c7aa665125427bd425621266d6ee0c03c7cbda9c635aaf88add006

Observation f39e90c7-6473-421f-8a07-88b379c89963 · inbound

SADA: Stability-guided Adaptive Diffusion Acceleration cites this paper.

SADA: Stability-guided Adaptive Diffusion Acceleration Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 1971

Resolution
unresolved
no resolver link, observed 2026-08-06T15:06:47.862588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:06:47.862588Z digest=sha256:fc995fcce94598f31ffcedf3eff72ebc9ef21dc841c7b31eb032e5cc22abf1bb

Observation 5ea0a60b-59a7-460a-b24c-a29539916253 · inbound

Transition Models: Rethinking the Generative Learning Objective cites this paper.

Transition Models: Rethinking the Generative Learning Objective Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T10:19:54.338843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:19:54.338843Z digest=sha256:0c618d78760c2610114fecefb9653d7d489065ade4c74b521c84535034ea6efa

Observation e22bf4ad-90c0-485e-9b86-15ba0e62355b · inbound

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation cites this paper.

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:04.960072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:04.960072Z digest=sha256:91854793c7b3d4146009f53f049612ddab47147a11df71ff62492be7c018f900

Observation 9615c93f-3338-4e9e-9b36-a96ec39ad54c · inbound

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer cites this paper.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:08:37.037281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-11T14:08:36.801359Z digest=sha256:48df79e3c7202771a792312771602cdbe8ca5e52652d5ddd4e1ff31a8c01f874

Observation 0af520a0-f9d2-4d3e-adfc-e940073ddea5 · inbound

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer cites this paper.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:31.331517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:31.331517Z digest=sha256:01b142f6ecd0cc58a05445a59fec7f37dd4bf0a85078ef05292f6def52abb0d4

Observation 2e1d1b14-e900-4af9-9b4e-44852958871f · inbound

PhyGDPO: Physics-Aware Groupwise Direct Preference Optimization for Physically Consistent Text-to-Video Generation cites this paper.

PhyGDPO: Physics-Aware Groupwise Direct Preference Optimization for Physically Consistent Text-to-Video Generation Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T13:24:44.583633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:24:44.583633Z digest=sha256:5fa5ac485b4c56af46572e6cd01b42d43dea648bcecadf3e188c4b700f98b19e

Observation 3baab651-d424-4f57-8e34-64e567e52b9d · inbound

LTX-2: Efficient Joint Audio-Visual Foundation Model cites this paper.

LTX-2: Efficient Joint Audio-Visual Foundation Model Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:06:20.602944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-13T07:06:20.470686Z digest=sha256:11ae848555cff345b6c6588497b649f9bc077b9d8bbe37c6a330162ea77313ea

Observation a1c7b072-6116-4d0d-9a7a-7cef17bd31ae · inbound

WeEdit: A Dataset, Benchmark and Glyph-Guided Framework for Text-centric Image Editing cites this paper.

WeEdit: A Dataset, Benchmark and Glyph-Guided Framework for Text-centric Image Editing Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T18:25:53.388773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:25:53.388773Z digest=sha256:202ae3e841b35abec2dfecc75bc8fdab95776be20e327bea992d81429c0ebb85

Observation 8cedf145-7646-4f11-97d4-ed771cb4ced8 · inbound

Your Pre-trained Diffusion Model Secretly Knows Restoration cites this paper.

Your Pre-trained Diffusion Model Secretly Knows Restoration Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:25:54.311194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T19:08:08.266895Z digest=sha256:a8aecb59175c337202476b4edd1ef8dca402bb538eebdedaad4e4fe08f0534e8

Observation 01e589f1-73b8-4d5b-a956-6b2e0ed37672 · inbound

PoM: A Linear-Time Replacement for Attention with the Polynomial Mixer cites this paper.

PoM: A Linear-Time Replacement for Attention with the Polynomial Mixer Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:05:49.276835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T19:20:52.810648Z digest=sha256:dae12f26e1fc4a61d25ef2e6d85ed4c9a3b61e276804f5149fe0421309b1cd38

Observation 78f93a64-378d-4863-af34-f65f86ad570f · inbound

Rethinking Cross-Layer Information Routing in Diffusion Transformers cites this paper.

Rethinking Cross-Layer Information Routing in Diffusion Transformers Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-21T05:29:39.351506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-21T05:29:19.111071Z digest=sha256:bca469883ed768325f615038081b4dc5a85b854c514f93d5cb8542367a174a9f

Observation 3974ee18-589f-4e11-a768-666e71b0b1e5 · inbound

Rethinking Cross-Layer Information Routing in Diffusion Transformers cites this paper.

Rethinking Cross-Layer Information Routing in Diffusion Transformers Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:54:58.006086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T17:50:22.258541Z digest=sha256:cb8c818bdcfe2a3d73f2556d65c748189a48fdef3c7fc47e8ddc24bf3af45837

Observation b484cbb6-46a1-4905-a953-82437073dd91 · inbound

iTryOn: Mastering Interactive Video Virtual Try-On with Spatial-Semantic Guidance cites this paper.

iTryOn: Mastering Interactive Video Virtual Try-On with Spatial-Semantic Guidance Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 67

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T05:13:58.480079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-21T05:11:34.441784Z digest=sha256:4617c4fa06d0999e4b6b2968d269a63c714b17e0cc91b15a38359371aefd8f06

Observation 99f43335-dc78-45c4-9a5f-c92e163eaf42 · inbound

PAI-Studio: Cinematic Video Background Replacement with Camera-Aware Motion cites this paper.

PAI-Studio: Cinematic Video Background Replacement with Camera-Aware Motion Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-01T21:16:14.711300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-28T17:14:23.336848Z digest=sha256:b181ef58f19578d78a75f7a16de4a3d09f1adaf21fe9a9f19a10ee41650cb71c

Observation 80f432bc-4c07-4351-a531-359f6798c77f · inbound

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation cites this paper.

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T03:06:29.553020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-28T10:23:43.501656Z digest=sha256:8e5ee85a750be07829cd9aeb435bbba1a0a04504c236812ffeb784a845e63b91

Observation ebfe7e57-5f25-46fc-9c68-ec5c72bedbd4 · inbound

DiffusionBench: On Holistic Evaluation of Diffusion Transformers cites this paper.

DiffusionBench: On Holistic Evaluation of Diffusion Transformers Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 50

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T16:59:58.191470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-26T00:06:11.951205Z digest=sha256:80c3c63e8bb725b932d1817b7895825e28e36e627ed77725a7dbfa6d7548ef74

Observation e74f25ad-366e-4769-ad71-e43ae8dab661 · inbound

SciIR: A Large-scale Training Dataset and Benchmark for Scientific Image Reasoning Generation cites this paper.

SciIR: A Large-scale Training Dataset and Benchmark for Scientific Image Reasoning Generation Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-06-30T06:54:20.738061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T06:48:22.594002Z digest=sha256:83e55f1de6d4d455861f65aa40f2a7394fa82a85223a597a3d3c9d17397aa503

Observation 317e4733-0934-4af5-9322-41a0ecdb7d7c · inbound

On the Limitations of Cross-Lingual Consistency in Multilingual Text-to-image Generation cites this paper.

On the Limitations of Cross-Lingual Consistency in Multilingual Text-to-image Generation Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T12:27:44.185380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:27:44.185380Z digest=sha256:5e06d0080ff44fd98ef3151deff672826ead1cb2d206fbfc625c08e90eab3dd2