Pith. sign in

Paper Citation Record · LEDGER

Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

As of 11 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 40 inbound Pith citation observations for arXiv:2405.05945.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2405.05945 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 40 of 40 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T20:55:01.928899Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T16:59:58.189803Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 1a8be5e9-fde3-4620-85d7-e12ba2b9effc · inbound

HunyuanVideo: A Systematic Framework For Large Video Generative Models cites this paper.

HunyuanVideo: A Systematic Framework For Large Video Generative Models Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-23T07:42:43.487031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:270c932a55b4b62c92273c4ffc9523afe4d0c62f47dc635acd52cf6ff2f59be4

Observation b0514003-a314-4573-b676-3ebcc310c0e8 · inbound

Mind the Time: Temporally-Controlled Multi-Event Video Generation cites this paper.

Mind the Time: Temporally-Controlled Multi-Event Video Generation Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T20:55:01.928899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:55:01.928899Z digest=sha256:3138826b2713739103e42ad2cf73ef9e6fb742ebbecf15ca9111932d60a098d1

Observation 8b34c3ec-711c-49a9-bd46-a26009a8df44 · inbound

MotionStone: Decoupled Motion Intensity Modulation with Diffusion Transformer for Image-to-Video Generation cites this paper.

MotionStone: Decoupled Motion Intensity Modulation with Diffusion Transformer for Image-to-Video Generation Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T20:20:36.493002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:20:36.493002Z digest=sha256:85b2f7cefe18ff7aeaf2f5187d88a41ab7582442703cf968ed1883c44a27eb37

Observation f5969d1a-10eb-459c-8dac-467337f136ca · inbound

SnapGen: Taming High-Resolution Text-to-Image Models for Mobile Devices with Efficient Architectures and Training cites this paper.

SnapGen: Taming High-Resolution Text-to-Image Models for Mobile Devices with Efficient Architectures and Training Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:01.614056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:01.614056Z digest=sha256:45bfb7e189c24c3ca763be76f2e6cd9f0cd6a7713deb18d8d97e5b258ddc3cc8

Observation b5e1b2c8-5007-4e4c-a5a6-da577476013d · inbound

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation cites this paper.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T15:43:00.877098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:43:00.877098Z digest=sha256:fbc03a0d13457ed1209907d599d999c2d648a9f49109307397f4a32cbf76c7c5

Observation 60c5fffb-27ba-46fc-ba1a-9f59906940ed · inbound

Video Diffusion Transformers are In-Context Learners cites this paper.

Video Diffusion Transformers are In-Context Learners Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:19.798497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:19.798497Z digest=sha256:45eb79ad89b015c8762f3333bd9191e2cf4d8252aed41c22157f56d72a300fd0

Observation 24708269-1164-4a1c-8351-c328662ed805 · inbound

Causal Diffusion Transformers for Generative Modeling cites this paper.

Causal Diffusion Transformers for Generative Modeling Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T14:20:49.289133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:20:49.289133Z digest=sha256:5118e1052192f7fee0efc89586270c4e8206aba2de39a517898944916538b359

Observation 0d33ef00-2426-439d-9090-5595f082775d · inbound

CLEAR: Conv-Like Linearization Revs Pre-Trained Diffusion Transformers Up cites this paper.

CLEAR: Conv-Like Linearization Revs Pre-Trained Diffusion Transformers Up Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T10:54:30.795049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:54:30.795049Z digest=sha256:03eaa980c88d4650196c338055273a46f037b471780eeabc767c70c7ea05a400

Observation c5a220a6-cc78-4066-bc0d-e24c36358e57 · inbound

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement cites this paper.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:35.030224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:35.030224Z digest=sha256:46324a5f1bf3a13fa692c79f239fb753ca39ccf4a395c81c739976bbb88fa32e

Observation 70ca07cd-d460-41aa-aafc-a56c2210b7c7 · inbound

Dual Diffusion for Unified Image Generation and Understanding cites this paper.

Dual Diffusion for Unified Image Generation and Understanding Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T23:02:07.643640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:02:07.643640Z digest=sha256:aa9a692e2e39179d33e63ec953c0f4d48653818fc72e686bbcbc6ce84771169d

Observation 7f2264ac-3a17-4ecb-93cb-f017f9639c8d · inbound

Reconstruction vs. Generation: Taming Optimization Dilemma in Latent Diffusion Models cites this paper.

Reconstruction vs. Generation: Taming Optimization Dilemma in Latent Diffusion Models Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T22:34:14.232420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:34:14.232420Z digest=sha256:c556841cd7912d8f1abd465f0ca8764249767752a4efd9391ecf9ba17e7fe212

Observation 01546f54-6fc4-4311-8884-8941a3753d23 · inbound

Ingredients: Blending Custom Photos with Video Diffusion Transformers cites this paper.

Ingredients: Blending Custom Photos with Video Diffusion Transformers Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T22:25:25.600000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:25:25.600000Z digest=sha256:622c5da3d1b3157376494e64d2baa574cb2b023cc7400470485371fa1c2b35f4

Observation 86cbbaa3-b068-402a-ae93-7afc7168ddb0 · inbound

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models cites this paper.

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T15:31:35.500088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:31:35.500088Z digest=sha256:b1eb135fb4e723fc70bf2dc4cfb08f9c1e999ebc288c3f74273c8a6bc3e44aa1

Observation 0e61a512-9c5d-48db-b995-acd0aeb50426 · inbound

DiffSplat: Repurposing Image Diffusion Models for Scalable Gaussian Splat Generation cites this paper.

DiffSplat: Repurposing Image Diffusion Models for Scalable Gaussian Splat Generation Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T10:56:30.925033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:56:30.925033Z digest=sha256:fd88a8510e32d1759c96118bb6467bcdb1fbfc9c222e50bcf5a24b4d2f14e495

Observation cf5811c7-1ecb-4db6-856d-72512fd18b65 · inbound

Lumina-Video: Efficient and Flexible Video Generation with Multi-scale Next-DiT cites this paper.

Lumina-Video: Efficient and Flexible Video Generation with Multi-scale Next-DiT Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T14:26:46.667065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:26:46.667065Z digest=sha256:07465276003542e878e1d3927fa7882a7caf665523064c9dbef2e1f37fcd0bbc

Observation f0b70c9f-20b8-4842-9c61-371730ec43de · inbound

Enhance-A-Video: Better Generated Video for Free cites this paper.

Enhance-A-Video: Better Generated Video for Free Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T12:32:46.422166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:32:46.422166Z digest=sha256:c655693f6a09c0f401535a59400e56ac44f59de11739507f8b97562aa4c94514

Observation e6debad6-641c-428b-8ca2-12cf1f884dcf · inbound

RealCam-I2V: Real-World Image-to-Video Generation with Interactive Complex Camera Control cites this paper.

RealCam-I2V: Real-World Image-to-Video Generation with Interactive Complex Camera Control Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T19:40:21.192239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T19:40:21.192239Z digest=sha256:f6734649a7f2e7dd793e391795b48a7579a1d4a00ad8fe8af81f19e969692b6e

Observation 369b99fb-2f79-4006-ae23-413dac14c3a7 · inbound

R2I-Bench: Benchmarking Reasoning-Driven Text-to-Image Generation cites this paper.

R2I-Bench: Benchmarking Reasoning-Driven Text-to-Image Generation Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:50.794369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:50.794369Z digest=sha256:f146aa9cc436a7ea1822bf66b2cb9076fd4b02950d05d2f941536d970ceb8d3a

Observation becaf3ad-2e97-42fc-9820-e757bec27070 · inbound

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation cites this paper.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.030846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.030846Z digest=sha256:ecdc40f5d0596dfdeb42452bede999361cc7cc42b1fcd0d3b8c53990b4c28cfb

Observation 4e0b2401-591c-44d4-8424-d67df269a7a7 · inbound

The Hidden Cost of an Image: Quantifying the Energy Consumption of AI Image Generation cites this paper.

The Hidden Cost of an Image: Quantifying the Energy Consumption of AI Image Generation Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T23:36:10.416115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:36:10.416115Z digest=sha256:2ac6a90471e63786874c1fe78477bf3150eb22d6dfa19392c20332ac63a4ccfc

Observation 68f82bf5-177a-43f2-bd05-a159abd0926b · inbound

Consistent Zero-shot 3D Texture Synthesis Using Geometry-aware Diffusion and Temporal Video Models cites this paper.

Consistent Zero-shot 3D Texture Synthesis Using Geometry-aware Diffusion and Temporal Video Models Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T22:41:38.549537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:41:38.549537Z digest=sha256:c4772c4fa2b6719372af00394c15b40e327cc549a8fbd435a84bd172b1543958

Observation 507a892c-920f-4d4f-8c5d-79a184db2557 · inbound

UniMC: Taming Diffusion Transformer for Unified Keypoint-Guided Multi-Class Image Generation cites this paper.

UniMC: Taming Diffusion Transformer for Unified Keypoint-Guided Multi-Class Image Generation Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T20:26:37.854469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:26:37.854469Z digest=sha256:d38850810ef73fa79ddeaa9b524291b75b20100a4f43ceb03ec7a052483c039c

Observation f36ad844-76c0-432f-88db-c24ea48c0fe3 · inbound

FreeCus: Free Lunch Subject-driven Customization in Diffusion Transformers cites this paper.

FreeCus: Free Lunch Subject-driven Customization in Diffusion Transformers Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T15:42:23.696320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:42:23.696320Z digest=sha256:c722a4a3095498243367c1441bbe9d385334e71d38639f27271985ec0e8c8dc9

Observation f39e90c7-6473-421f-8a07-88b379c89963 · inbound

SADA: Stability-guided Adaptive Diffusion Acceleration cites this paper.

SADA: Stability-guided Adaptive Diffusion Acceleration Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 1971

Resolution
unresolved
no resolver link, observed 2026-08-06T15:06:47.862588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:06:47.862588Z digest=sha256:4a4320e85a7c5650c8700526c9ea6f2c9ecd29251172061095ef555dbd3fa0af

Observation 5ea0a60b-59a7-460a-b24c-a29539916253 · inbound

Transition Models: Rethinking the Generative Learning Objective cites this paper.

Transition Models: Rethinking the Generative Learning Objective Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T10:19:54.338843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:19:54.338843Z digest=sha256:a7a44f54ae148fd7e5dc55d2649316cacd31867c3494ea30026fb94b362aa59a

Observation e22bf4ad-90c0-485e-9b86-15ba0e62355b · inbound

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation cites this paper.

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:04.960072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:04.960072Z digest=sha256:9364caccd31421978c63cd54f298a0f1f639164e7d26e62c762798ded43c7520

Observation 9615c93f-3338-4e9e-9b36-a96ec39ad54c · inbound

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer cites this paper.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:08:37.037281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-11T14:08:36.801359Z digest=sha256:1853abef8b34a5dd34119a3283dd4570bd077e8c3f5feaccf040bd7d06120e06

Observation 0af520a0-f9d2-4d3e-adfc-e940073ddea5 · inbound

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer cites this paper.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:31.331517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:31.331517Z digest=sha256:1c9ed634e69bff1a6ecc74405cc5e716f2186e31b1d6c899c04c2a5f7d557348

Observation 2e1d1b14-e900-4af9-9b4e-44852958871f · inbound

PhyGDPO: Physics-Aware Groupwise Direct Preference Optimization for Physically Consistent Text-to-Video Generation cites this paper.

PhyGDPO: Physics-Aware Groupwise Direct Preference Optimization for Physically Consistent Text-to-Video Generation Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T13:24:44.583633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:24:44.583633Z digest=sha256:a050457cde8ba66da53867d386c07877d7fe170261d588b1e45371577a029fc5

Observation 3baab651-d424-4f57-8e34-64e567e52b9d · inbound

LTX-2: Efficient Joint Audio-Visual Foundation Model cites this paper.

LTX-2: Efficient Joint Audio-Visual Foundation Model Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:06:20.602944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T07:06:20.470686Z digest=sha256:1b07d6e61846bc27b86ce35a95cabd355f6f9eafffbbeffdde5a43488b2f2c85

Observation a1c7b072-6116-4d0d-9a7a-7cef17bd31ae · inbound

WeEdit: A Dataset, Benchmark and Glyph-Guided Framework for Text-centric Image Editing cites this paper.

WeEdit: A Dataset, Benchmark and Glyph-Guided Framework for Text-centric Image Editing Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T18:25:53.388773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:25:53.388773Z digest=sha256:f50e8f261c066a06d5955b12382303cde64f9465f7bcee7d39c7c9c5caf78afa

Observation 8cedf145-7646-4f11-97d4-ed771cb4ced8 · inbound

Your Pre-trained Diffusion Model Secretly Knows Restoration cites this paper.

Your Pre-trained Diffusion Model Secretly Knows Restoration Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:25:54.311194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T19:08:08.266895Z digest=sha256:f0abc4f18f251452926579ea63edb812275c2e7061c78d0fc583a1689be58269

Observation 01e589f1-73b8-4d5b-a956-6b2e0ed37672 · inbound

PoM: A Linear-Time Replacement for Attention with the Polynomial Mixer cites this paper.

PoM: A Linear-Time Replacement for Attention with the Polynomial Mixer Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:05:49.276835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T19:20:52.810648Z digest=sha256:708d1287e33fee530666b899d8fd6e8c6fea483c8a46ffb152d85fc63af2f054

Observation 78f93a64-378d-4863-af34-f65f86ad570f · inbound

Rethinking Cross-Layer Information Routing in Diffusion Transformers cites this paper.

Rethinking Cross-Layer Information Routing in Diffusion Transformers Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-21T05:29:39.351506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-21T05:29:19.111071Z digest=sha256:ddebb8ae3bfe95be7beeb6e92ced2dba5be59eb625ce1ebb8021d1f04780628f

Observation 3974ee18-589f-4e11-a768-666e71b0b1e5 · inbound

Rethinking Cross-Layer Information Routing in Diffusion Transformers cites this paper.

Rethinking Cross-Layer Information Routing in Diffusion Transformers Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:54:58.006086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-30T17:50:22.258541Z digest=sha256:5714d127b512ba0d537c6c2d93bd287d3fd1594e7da7591db7f624830dd7e179

Observation b484cbb6-46a1-4905-a953-82437073dd91 · inbound

iTryOn: Mastering Interactive Video Virtual Try-On with Spatial-Semantic Guidance cites this paper.

iTryOn: Mastering Interactive Video Virtual Try-On with Spatial-Semantic Guidance Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 67

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T05:13:58.480079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-21T05:11:34.441784Z digest=sha256:5b2718fabf89b89c1e1547ee8af6bd69f0cc818526b2849b66e97628a25cc486

Observation 99f43335-dc78-45c4-9a5f-c92e163eaf42 · inbound

PAI-Studio: Cinematic Video Background Replacement with Camera-Aware Motion cites this paper.

PAI-Studio: Cinematic Video Background Replacement with Camera-Aware Motion Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-01T21:16:14.711300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-28T17:14:23.336848Z digest=sha256:be394d46275ea6dadeee74f8da9583ae3bdf47a65727613a41ddfce8fc1bac53

Observation 80f432bc-4c07-4351-a531-359f6798c77f · inbound

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation cites this paper.

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T03:06:29.553020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-28T10:23:43.501656Z digest=sha256:d73ad927a85491a9b987cb196c1989629c860a3912ea159b776c09608526384e

Observation ebfe7e57-5f25-46fc-9c68-ec5c72bedbd4 · inbound

DiffusionBench: On Holistic Evaluation of Diffusion Transformers cites this paper.

DiffusionBench: On Holistic Evaluation of Diffusion Transformers Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 50

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T16:59:58.191470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-26T00:06:11.951205Z digest=sha256:4aedae1f414fc0c6274c45c26d4986bfb81a881b75538b84fd009cb0b701e3a1

Observation e74f25ad-366e-4769-ad71-e43ae8dab661 · inbound

SciIR: A Large-scale Training Dataset and Benchmark for Scientific Image Reasoning Generation cites this paper.

SciIR: A Large-scale Training Dataset and Benchmark for Scientific Image Reasoning Generation Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-06-30T06:54:20.738061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-30T06:48:22.594002Z digest=sha256:6ce2c52b861b597b6620876dd62c84687a266505a58b0f633eaa2e8de5dd4c62