Pith. sign in

Paper Citation Record · LEDGER

Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

As of 20 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 100 inbound Pith citation observations for arXiv:2405.08748.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2405.08748 v1

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-16T14:58:37.383749Z

measured 141 of 141 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 100 of 123 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:40:31.264007Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

41 of 41 outbound references displayed

  • verified exact15
  • verified fuzzy25
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

1
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 0c574b84-5f3a-4d03-a150-467b23204e7d · outbound

This paper cites https://www.midjourney.com/home.

Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding https://www.midjourney.com/home

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T14:58:37.495138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T14:58:37.383749Z digest=sha256:34f92be6c914ac7ec309cf69a6d11c82398501f90da47b34f03744ace975b934

Observation 039c4d4d-be1b-4460-9f15-a82fceac5df1 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-16T14:58:37.407084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T14:58:37.383749Z digest=sha256:653dc89397bc807189cd371d46fda14223f27a9be7f9beab33993137e20e421b

Observation 40c7ac0e-f72c-4415-a366-1a4b84bb0193 · outbound

This paper cites eDiff-I: Text-to-Image Diffusion Models with an Ensemble of Expert Denoisers.

Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding eDiff-I: Text-to-Image Diffusion Models with an Ensemble of Expert Denoisers

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-16T14:58:37.412582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T14:58:37.383749Z digest=sha256:9b8a3a394953b6896453416986975f370ce466fcb9979dafb317c2e30844f172

Observation 8593d81b-356a-4f68-bda6-e9926667a283 · outbound

This paper cites All are worth words: A vit backbone for diffusion models.

Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding All are worth words: A vit backbone for diffusion models

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T14:58:37.524125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T14:58:37.383749Z digest=sha256:0fc657d20c336de7a4b9283128b2ad9a92008906c6569753ef08f6dc0dee1af0

Observation 6f64cbf0-2d0b-42a1-9e50-5cca60d37f1a · outbound

This paper cites Improving image generation with better captions.

Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding Improving image generation with better captions

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T14:58:37.527389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T14:58:37.383749Z digest=sha256:f5bd358f7f10c0f024b5d978ee290f3881a8efce7dcc7ca7513d2ba288a76544

Observation 40350f7a-cdac-4ba3-a06b-6a626f54c1f9 · outbound

This paper cites Muse: Text-to-image generation via masked generative transformers.

Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding Muse: Text-to-image generation via masked generative transformers

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T14:58:37.530183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T14:58:37.383749Z digest=sha256:d9e4415e5d8001592b0d6d735460a67dab6d21815c869e3349a1d823f7dc557f

Observation 581e6b9a-8d09-4f9d-8b70-64f55df30ad4 · outbound

This paper cites Pixart-\alpha: Fast training of diffusion transformer for photorealistic text-to-image synthesis.

Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding Pixart-\alpha: Fast training of diffusion transformer for photorealistic text-to-image synthesis

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T14:58:37.532988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T14:58:37.383749Z digest=sha256:50526e77a61cbdfdca01307af77233d5408a17bbe6def3d372ba3a0d656ad043

Observation d7ee4fbd-c76c-454f-919d-0c74c54748df · outbound

This paper cites Flashattention: Fast and memory-efficient exact attention with io-awareness.

Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding Flashattention: Fast and memory-efficient exact attention with io-awareness

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T14:58:37.535720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T14:58:37.383749Z digest=sha256:07778b399ce9522d0ced33dc28181bc88090dc19efe1cce7e17a69c603ad8b92

Observation 0723abeb-487a-49fd-aebf-02d570f19eed · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding An image is worth 16x16 words: Transformers for image recognition at scale

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T14:58:37.538916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T14:58:37.383749Z digest=sha256:6a7b8fac8da3b704d4af9bee7609c9aeefe3b81ac6ab3b29dded9918fcd31621

Observation 06ea00d7-df55-4979-8d97-edf34b1c8b75 · outbound

This paper cites Scaling Rectified Flow Transformers for High-Resolution Image Synthesis.

Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding Scaling Rectified Flow Transformers for High-Resolution Image Synthesis

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-16T14:58:37.424102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T14:58:37.383749Z digest=sha256:d0d3e4eb0995ae9212899c44c10ce90d82fe7712ebab4386ff9c91a43b3b0ef6

Observation 239459fe-7563-48fd-a297-eca4d5de87d5 · outbound

This paper cites Making LLaMA SEE and Draw with SEED Tokenizer.

Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding Making LLaMA SEE and Draw with SEED Tokenizer

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-16T14:58:37.478355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T14:58:37.383749Z digest=sha256:9e7c4b27e95e534c6ee3b3df106cb052a57ca81e01c617d8e43581ccbaa98e2b

Observation fc659712-5713-49dc-8a1e-11a9fa8ebe90 · outbound

This paper cites Matryoshka diffusion models.

Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding Matryoshka diffusion models

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T14:58:37.542066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T14:58:37.383749Z digest=sha256:ba7f94783efe9c31cca96ef40af6b40caf949c8f1fbb909a3b3292dddbef37ed

Observation a562c28b-18a9-4b2f-ae70-b6f40eb956c4 · outbound

This paper cites Query-key normalization for transformers.

Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding Query-key normalization for transformers

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T14:58:37.545151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T14:58:37.383749Z digest=sha256:16a07ed154f2b7261e2f0c3e683ee4a95e05b38bd07e1eb4edb3c86702f2f4a4

Observation e738f32f-cb26-4a95-abd1-213b077cb13f · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilibrium.

Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding Gans trained by a two time-scale update rule converge to a local nash equilibrium

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T14:58:37.548073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T14:58:37.383749Z digest=sha256:6957c78ee16da7ddfb7498b35984cef70ad8fc19aa6d5846c8bcf7f6d34398bb

Observation aab99592-907a-43ce-8e5b-57bfea5fa980 · outbound

This paper cites DialogGen: Multi-modal Interactive Dialogue System for Multi-turn Text-to-Image Generation.

Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding DialogGen: Multi-modal Interactive Dialogue System for Multi-turn Text-to-Image Generation

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-16T14:58:37.443989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T14:58:37.383749Z digest=sha256:ad1bff838b2075b33f577297441a6e15c6330662d263f052e553079f832f6caa

Observation c828b109-e8ce-4c3f-acb1-6a46c0f85c15 · outbound

This paper cites Playground v2.5: Three Insights towards Enhancing Aesthetic Quality in Text-to-Image Generation.

Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding Playground v2.5: Three Insights towards Enhancing Aesthetic Quality in Text-to-Image Generation

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-16T14:58:37.450305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T14:58:37.383749Z digest=sha256:548f1a9249c428121503dae9a0fb5e7d6d4cdb37b8aeb258dc66420f46a59aee

Observation 906aa91c-7c34-497a-9f21-d74968aeaabe · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T14:58:37.550615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T14:58:37.383749Z digest=sha256:8eca3a00884f575e1b73dcc8bd393fb85a765c39a7bc9375c00300026279ea63

Observation ce80a498-3bb0-424a-997e-2e225d8a2be4 · outbound

This paper cites Swinv2-imagen: Hierarchical vision transformer diffusion models for text-to-image generation.

Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding Swinv2-imagen: Hierarchical vision transformer diffusion models for text-to-image generation

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T14:58:37.554109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T14:58:37.383749Z digest=sha256:74c8f12566adde3f1d53659db8c063791f7ecc7a1fdffb1660e29a8c4a691a0f

Observation cf3631ec-65a9-4c21-8fa0-8c3f833a8f41 · outbound

This paper cites Microsoft coco: Common objects in context.

Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding Microsoft coco: Common objects in context

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T14:58:37.556896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T14:58:37.383749Z digest=sha256:697c019cb004d4de3e234b4f3bde2b18e616fb558d5b08fbc05bda4faa68fbef

Observation 0357996c-8a73-4404-a89a-9b856e3d90b5 · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding Improved Baselines with Visual Instruction Tuning

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-16T14:58:37.418068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T14:58:37.383749Z digest=sha256:d282116fc687cdb4462fb6c60a98963ca3ae4da345c620fec256849a39df2b95

Observation 49898400-25a7-4fab-a029-f035f9f43716 · outbound

This paper cites Instaflow: One step is enough for high-quality diffusion-based text-to-image generation.

Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding Instaflow: One step is enough for high-quality diffusion-based text-to-image generation

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T14:58:37.559561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T14:58:37.383749Z digest=sha256:3d2e93fb3961fad6f6b82025532b692a9999207d35a5484d75459db37ba199a8

Observation f1f8c7b2-f33b-44f0-8ca5-ec49c773b450 · outbound

This paper cites Latent Consistency Models: Synthesizing High-Resolution Images with Few-Step Inference.

Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding Latent Consistency Models: Synthesizing High-Resolution Images with Few-Step Inference

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-16T14:58:37.430827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T14:58:37.383749Z digest=sha256:c0178fd1f9f3cc9583b7b78ff317532d33fbe91e40e28335cf42ccd608d672b0

Observation 92525b43-496a-4e6a-b0fb-3ad6866b5bd9 · outbound

This paper cites Scalable diffusion models with transformers.

Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding Scalable diffusion models with transformers

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T14:58:37.562527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T14:58:37.383749Z digest=sha256:d58e20ef76e1c7c02eb8d606940c68e95ad7e91233f1953e84a50ae9076fb417

Observation 4b421b8b-cce4-43ef-9a75-aa9814a74e6c · outbound

This paper cites Sdxl: Improving latent diffusion models for high-resolution image synthesis.

Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding Sdxl: Improving latent diffusion models for high-resolution image synthesis

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T14:58:37.565049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T14:58:37.383749Z digest=sha256:bc160b7f52d1a8adf466b2cb89662e88df2b726a229d73cef83d5dc50c33b774

Observation e044e6b6-5b34-4e22-a1ee-91824dfd3ba5 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding Learning transferable visual models from natural language supervision

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T14:58:37.491793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T14:58:37.383749Z digest=sha256:1e4d33045dd1d4ad1f0c2c8a173ca2a9cd354a96b9a2b84ac78dd850f899da59

Observation ff7379f4-2eff-4f58-baf3-6277d9280564 · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.

Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding Exploring the limits of transfer learning with a unified text-to-text transformer

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T14:58:37.521112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T14:58:37.383749Z digest=sha256:ad72dc4910b50cb190e584241ecb92effec2e9534dbbddceeb58a2197d31e03d

Observation 0817e3f4-34ad-4a5f-874d-fb924a198d12 · outbound

This paper cites Zero: Memory optimizations toward training trillion parameter models.

Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding Zero: Memory optimizations toward training trillion parameter models

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T14:58:37.498331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T14:58:37.383749Z digest=sha256:9646a04699b8305c4fcff5c51a41de441ecfad29d6c6f052734cd20c4e10bd29

Observation 6f26e00a-fc3d-4085-9615-d1c07f0055b0 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding High-resolution image synthesis with latent diffusion models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T14:58:37.501856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T14:58:37.383749Z digest=sha256:65a268b38e91e3e9e1086be14fe14b299f97619e2865cca1af30851a36cb5922

Observation e2deabf2-87a2-4bd1-bc43-1c41295c135b · outbound

This paper cites Photorealistic text-to-image diffusion models with deep language understanding.

Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding Photorealistic text-to-image diffusion models with deep language understanding

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T14:58:37.505419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T14:58:37.383749Z digest=sha256:d337c148ff4760ddcb327faef0b81c609896704e6a4fc4f4d62afdf400bf0835

Observation e12772ff-067d-4b1e-b59f-3caea0ed7bf1 · outbound

This paper cites Progressive distillation for fast sampling of diffusion models.

Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding Progressive distillation for fast sampling of diffusion models

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T14:58:37.508949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T14:58:37.383749Z digest=sha256:e50a9326b164c3a96da20ca6c1daa01f9818f10d62aed1645772efeddad02a76

Observation 40df8d60-f4ca-468b-b68b-8d4529eddc54 · outbound

This paper cites Adversarial Diffusion Distillation.

Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding Adversarial Diffusion Distillation

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-16T14:58:37.436697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T14:58:37.383749Z digest=sha256:379a7f163a114d41dca979faf4c0c99c0eb08796d31fbaae2299ebc4e2f18912

Observation 496b6f5b-db62-4586-8d14-79b01dd9ae0f · outbound

This paper cites Roformer: Enhanced transformer with rotary position embedding.

Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding Roformer: Enhanced transformer with rotary position embedding

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T14:58:37.512083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T14:58:37.383749Z digest=sha256:83e5a1679518dc2602aace5efa6b9d0eb6e66daa22942aed6ecf83e507592755

Observation cf757cd1-032d-4489-9283-539769e0af20 · outbound

This paper cites Attention is all you need.

Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding Attention is all you need

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T14:58:37.514892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T14:58:37.383749Z digest=sha256:7e63de36dd456119019b60f967b23cf6dd1ac476432003b861fdf83c8ce606c9

Observation 9c83a454-109f-4355-af2b-7413fc309d7d · outbound

This paper cites PAI-Diffusion: Constructing and Serving a Family of Open Chinese Diffusion Models for Text-to-image Synthesis on the Cloud.

Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding PAI-Diffusion: Constructing and Serving a Family of Open Chinese Diffusion Models for Text-to-image Synthesis on the Cloud

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-16T14:58:37.455502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T14:58:37.383749Z digest=sha256:558dacbd8249a09932baf1aac7c1c69685d6c966cd652bb132505785e015a3ff

Observation 082d8ffd-e82a-4384-9e7f-0bc07055dbf8 · outbound

This paper cites NExT-GPT: Any-to-Any Multimodal LLM.

Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding NExT-GPT: Any-to-Any Multimodal LLM

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-16T14:58:37.460606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T14:58:37.383749Z digest=sha256:0b9fc441542b585e947748d13bd45ce7af627cc6042418464b49ff99c28843f6

Observation 17dcc32c-0090-4f4f-b188-e71e1e898e35 · outbound

This paper cites Taiyi-Diffusion-XL: Advancing Bilingual Text-to-Image Generation with Large Vision-Language Model Support.

Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding Taiyi-Diffusion-XL: Advancing Bilingual Text-to-Image Generation with Large Vision-Language Model Support

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-16T14:58:37.465038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T14:58:37.383749Z digest=sha256:1232d331c70c1db471290d30c38a06ce832abcca8b3a04da3a0ed83968b0193e

Observation e03bf071-645a-47ad-8781-513f76be546c · outbound

This paper cites UFOGen: You Forward Once Large Scale Text-to-Image Generation via Diffusion GANs.

Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding UFOGen: You Forward Once Large Scale Text-to-Image Generation via Diffusion GANs

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-16T14:58:37.469901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T14:58:37.383749Z digest=sha256:5c909a267eb985bcec6f3280bca1108afd4b9ebcf3c66567a46b4c74224239e6

Observation d2c356f3-c195-426d-8230-ba8200344824 · outbound

This paper cites Mastering Text-to-Image Diffusion: Recaptioning, Planning, and Generating with Multimodal LLMs.

Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding Mastering Text-to-Image Diffusion: Recaptioning, Planning, and Generating with Multimodal LLMs

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-16T14:58:37.473771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T14:58:37.383749Z digest=sha256:0dd297c168e830f8c1c38a2582e62b1f60dba8a8fec8f62dd8ba3af04261b715

Observation 5bfd222e-5971-46f7-b3aa-f0b6e370dbaa · outbound

This paper cites Altdiffusion: A multilingual text-to-image diffusion model.

Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding Altdiffusion: A multilingual text-to-image diffusion model

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T14:58:37.517888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T14:58:37.383749Z digest=sha256:7c6ce4bcc8b4b6ba2fefe0022b9f3c9674f5ef1021a6275e55b91dd67d66f7fc

Observation b7dc518b-cf5d-429f-8d5c-aa5a690bb261 · outbound

This paper cites One-step Diffusion with Distribution Matching Distillation.

Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding One-step Diffusion with Distribution Matching Distillation

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T14:58:37.483479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T14:58:37.383749Z digest=sha256:dca5c50c9cdffe94c44eb89c6c836d1f56d8f0687cb76b67c010fddc7ddb5642

Observation a7276b04-4dc2-40a0-9297-790122fa7487 · outbound

This paper cites CapsFusion: Rethinking Image-Text Data at Scale.

Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding CapsFusion: Rethinking Image-Text Data at Scale

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-16T14:58:37.488355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T14:58:37.383749Z digest=sha256:0cc49aa944105d1af9e50d0a47aae36145890a2df81a7c3fd2bfdd51a4a83600

Pith citing papers

Observation e95e4235-bd61-4cc4-98c2-271aa5e4a753 · inbound

Emu3: Next-Token Prediction is All You Need cites this paper.

Emu3: Next-Token Prediction is All You Need Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-16T14:58:37.566044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-11T10:56:06.418360Z digest=sha256:375c40aef03212abacf5f0e7b8eccfde19e241ba5d7553b2df45d39aacd56a85

Observation b784b988-ed8e-4877-8bf6-3f1c26089a6a · inbound

SANA: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformers cites this paper.

SANA: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformers Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T14:58:37.566044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T00:56:50.009149Z digest=sha256:598241b40b59273c6b3e4ddb9857f66982e00aaede2e37922db6ef61f73c4ea7

Observation 981345a0-6025-41c6-a7b9-3e4e260114b9 · inbound

High-Resolution Image Synthesis via Next-Token Prediction cites this paper.

High-Resolution Image Synthesis via Next-Token Prediction Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-12T14:55:42.108525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:55:42.108525Z digest=sha256:3d89bc87263248ad1bfbe9ec6755d29b136e6574586fe4415e2b8dfa6a203d74

Observation 35a26a3b-db9c-4b16-a921-46dc957d7ccf · inbound

Text-to-Image Synthesis: A Decade Survey cites this paper.

Text-to-Image Synthesis: A Decade Survey Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 123

Resolution
unresolved
no resolver link, observed 2026-08-12T13:33:38.760346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:33:38.760346Z digest=sha256:c7338e8401ed7b27de6e4d70436e57e5cd61c9bffbb2e863484609ba117c7ba0

Observation 3153ac62-9931-449c-a407-586072d518d3 · inbound

One Diffusion to Generate Them All cites this paper.

One Diffusion to Generate Them All Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T13:19:54.633105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:19:54.633105Z digest=sha256:e5d653246a784092bf5400c5d7f45c201c51c44e0d1ae10641ee034e93c4586b

Observation 0358a844-7806-4255-9a7c-cc027c1d2bf3 · inbound

Efficient Multi-modal Large Language Models via Visual Token Grouping cites this paper.

Efficient Multi-modal Large Language Models via Visual Token Grouping Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T12:25:00.345323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:25:00.345323Z digest=sha256:14eda2006245d1c488b51c669dbc1aedcbdbaa22f34680382dd25b33cf7c1452

Observation 41b9b3fb-433a-4e58-a9b7-1e3ab7cd8412 · inbound

Open-Sora Plan: Open-Source Large Video Generation Model cites this paper.

Open-Sora Plan: Open-Source Large Video Generation Model Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-23T08:42:45.165303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-23T08:38:27.946746Z digest=sha256:a8302dbfa9dc3673120a1892883b59034564f2b25a62c9b7f847f9913d165668

Observation f1b82155-c6c3-4532-b83a-08dd778c6193 · inbound

HoloDrive: Holistic 2D-3D Multi-Modal Street Scene Generation for Autonomous Driving cites this paper.

HoloDrive: Holistic 2D-3D Multi-Modal Street Scene Generation for Autonomous Driving Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T04:27:47.864360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:27:47.864360Z digest=sha256:64c990fdbf4a364eae88d8927aaeb577b31e75ef06b6fb0fe15e2d5b756c258b

Observation 7765ed2f-9fb4-482c-a343-69def11b1db7 · inbound

Pinco: Position-induced Consistent Adapter for Diffusion Transformer in Foreground-conditioned Inpainting cites this paper.

Pinco: Position-induced Consistent Adapter for Diffusion Transformer in Foreground-conditioned Inpainting Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T22:10:42.746365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:10:42.746365Z digest=sha256:58da1646ec4a69bcf57840e0d1dafa67a13b6eb76e79f4c033225c8993301426

Observation 7d696d16-2c28-4696-ba80-897ca68eead4 · inbound

SnapGen: Taming High-Resolution Text-to-Image Models for Mobile Devices with Efficient Architectures and Training cites this paper.

SnapGen: Taming High-Resolution Text-to-Image Models for Mobile Devices with Efficient Architectures and Training Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:01.695752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:01.695752Z digest=sha256:bc6d192594fa0043b1120b8ac7b4a0a170eac555deee46ec97e6a0dc4d5f1c09

Observation 29af3adc-1aeb-4752-b769-eb686e7c6d92 · inbound

Zigzag Diffusion Sampling: Diffusion Models Can Self-Improve via Self-Reflection cites this paper.

Zigzag Diffusion Sampling: Diffusion Models Can Self-Improve via Self-Reflection Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T15:36:11.704510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:36:11.704510Z digest=sha256:38acb6baad74967fb88bf2f4121298d279f66b3ee37b8bde4c546b7faab1d419

Observation cb023b8d-cfb7-4994-8ec8-83b8fd4eff4d · inbound

Efficient Scaling of Diffusion Transformers for Text-to-Image Generation cites this paper.

Efficient Scaling of Diffusion Transformers for Text-to-Image Generation Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T14:11:00.567171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:11:00.567171Z digest=sha256:6d7ce89fcdddcf4e60f1200b25951ea7e05989217272100d792fa38cc163094c

Observation cbb84cfd-0e08-491d-b4e3-b9c6eb9be24d · inbound

F-Bench: Rethinking Human Preference Evaluation Metrics for Benchmarking Face Generation, Customization, and Restoration cites this paper.

F-Bench: Rethinking Human Preference Evaluation Metrics for Benchmarking Face Generation, Customization, and Restoration Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T13:28:09.360430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:28:09.360430Z digest=sha256:fc0f42c15a085b58eb6ff72d07c4409b06a3e7129e49ffcbb187acf976d97acd

Observation c8803063-afa4-44ac-98f4-725ef1f37e45 · inbound

CLEAR: Conv-Like Linearization Revs Pre-Trained Diffusion Transformers Up cites this paper.

CLEAR: Conv-Like Linearization Revs Pre-Trained Diffusion Transformers Up Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T10:54:30.876054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:54:30.876054Z digest=sha256:e266602ca74e8ce758151f6c333bcfb83544b35140d84e55906048521da0be93

Observation 9abed91e-9e08-447c-94ad-77d2fa0eafaf · inbound

PromptLA: Towards Integrity Verification of Black-box Text-to-Image Diffusion Models cites this paper.

PromptLA: Towards Integrity Verification of Black-box Text-to-Image Diffusion Models Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T11:19:55.532992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:19:55.532992Z digest=sha256:1edafb66fca8decb6985ca36b74463c4b09de96adb38aacdabdaa5a150b10811

Observation 6decca13-abc0-427c-a26b-d644163b295b · inbound

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation cites this paper.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T05:02:21.029310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:02:21.029310Z digest=sha256:9cc3933d20ab22e4a5a37e800648125d03102a1c7ec95aebccb6c944e2e85d29

Observation f5a4f1d5-8e79-4e79-bb7a-dcb878fd2cb5 · inbound

UNIC-Adapter: Unified Image-instruction Adapter with Multi-modal Transformer for Image Generation cites this paper.

UNIC-Adapter: Unified Image-instruction Adapter with Multi-modal Transformer for Image Generation Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T04:24:20.032698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:24:20.032698Z digest=sha256:cfde020315590ef7c5acbfeba6ba0d7189f9987fd0f865f51b90b09067541f67

Observation f407df20-bd64-4035-920a-d02445d15108 · inbound

SeedVR: Seeding Infinity in Diffusion Transformer Towards Generic Video Restoration cites this paper.

SeedVR: Seeding Infinity in Diffusion Transformer Towards Generic Video Restoration Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T22:33:42.710583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:33:42.710583Z digest=sha256:6ecb61af032b838e178a9c3eb50f31ccb8a50f402e154dbb28a28731da3b553d

Observation 007b7b0d-6cae-4683-9d25-75725bc7c622 · inbound

Enhancing Image Generation Fidelity via Progressive Prompts cites this paper.

Enhancing Image Generation Fidelity via Progressive Prompts Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T20:52:46.053945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:52:46.053945Z digest=sha256:f84735dba2b920bf4a587e185f3f3c863f6dbe1b7a8c9b1682dddc9ccac07d78

Observation 82d4b05c-a5fa-4a76-8362-e9ebb2c09444 · inbound

AnyStory: Towards Unified Single and Multiple Subject Personalization in Text-to-Image Generation cites this paper.

AnyStory: Towards Unified Single and Multiple Subject Personalization in Text-to-Image Generation Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T20:01:21.485520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:01:21.485520Z digest=sha256:8f0c006439145548e15ee51a7b604d8f6f4c55e0abe89a71d7ec46172b2cdbc7

Observation 4287647a-e996-43af-9b67-db4af3afc96b · inbound

T2ISafety: Benchmark for Assessing Fairness, Toxicity, and Privacy in Image Generation cites this paper.

T2ISafety: Benchmark for Assessing Fairness, Toxicity, and Privacy in Image Generation Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T17:05:50.691701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:05:50.691701Z digest=sha256:2c61fddf7bec8de8fde42f1c24d7c2d09c83ec903a9c8ca22bed8905a36a9e7f

Observation abb4556a-db91-4677-9123-e5c6d201c538 · inbound

EchoVideo: Identity-Preserving Human Video Generation by Multimodal Feature Fusion cites this paper.

EchoVideo: Identity-Preserving Human Video Generation by Multimodal Feature Fusion Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T16:12:00.121311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:12:00.121311Z digest=sha256:5bc359f423b6dd2a53254e799d4a5dc0df6cf9c1d5571a52af2ecb4aa49290c2

Observation 1b77b361-f091-4e10-b15f-9c6f63666604 · inbound

Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling cites this paper.

Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-16T14:58:37.566044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-11T08:14:52.890145Z digest=sha256:b0757edbe48a5a1ec334643b8e54498fdc9def295c607dfa695fe9899d02adce

Observation 613fc111-a701-4619-92a3-5df42363d9a0 · inbound

SANA 1.5: Efficient Scaling of Training-Time and Inference-Time Compute in Linear Diffusion Transformer cites this paper.

SANA 1.5: Efficient Scaling of Training-Time and Inference-Time Compute in Linear Diffusion Transformer Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-09T23:37:12.585450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T23:37:12.585450Z digest=sha256:caad0e95b42f31ce94ab623f8835489d623dde721ca298998083f13649809993

Observation 8bc4c100-3750-4c4e-a460-f2075022732a · inbound

UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding cites this paper.

UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T19:32:31.875847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:32:31.875847Z digest=sha256:df0def4fab76bdc5eb2efeda99589b33642fb21ccb16d1245f4a87fdf4f3588a

Observation ea13ae8c-cd82-4109-969a-4b4c3f4fa44f · inbound

VFX Creator: Animated Visual Effect Generation with Controllable Diffusion Transformer cites this paper.

VFX Creator: Animated Visual Effect Generation with Controllable Diffusion Transformer Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T17:13:02.774524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:13:02.774524Z digest=sha256:fc61dbe2ff2f434a54a7e80937e5accb12fca7571bc83d6efab47867fa26676d

Observation f728abfe-1ac2-47eb-b4a0-b36b74f59ed7 · inbound

Matrix3D: Large Photogrammetry Model All-in-One cites this paper.

Matrix3D: Large Photogrammetry Model All-in-One Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-08T11:56:49.457187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T11:56:49.457187Z digest=sha256:c4de1adb374fe490ecb7c7ac775bc441ffe58e711714a0ef18dc2266793fbb99

Observation 93a71d0d-a75f-49d1-ac8b-be50586c74c0 · inbound

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model cites this paper.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-19T08:02:23.310121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:073c86ade8965b4a044db1eddccaddcf88fb52c92e6db7571a52d074c27a0f88

Observation 3ea81b16-5537-4f3f-8adb-e6216c3153de · inbound

RectifiedHR: Enable Efficient High-Resolution Synthesis via Energy Rectification cites this paper.

RectifiedHR: Enable Efficient High-Resolution Synthesis via Energy Rectification Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-23T01:35:19.497719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-23T01:33:49.962427Z digest=sha256:0d1daace6d7a63b53c42a35a514a347c737a19c0321065f737276129a0bcb7b6

Observation 76dcd2c9-854b-48f0-8874-9b2212109e86 · inbound

VACE: All-in-One Video Creation and Editing cites this paper.

VACE: All-in-One Video Creation and Editing Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-16T14:58:37.566044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T00:53:53.855965Z digest=sha256:2a06a20b6ec3fe8d2b2dc327a7f4ab1fc3e88a7c03048d11d22819d9adf779ea

Observation 619bf7e2-d1fb-4938-8c63-eea9a23e6b4e · inbound

Seedream 2.0: A Native Chinese-English Bilingual Image Generation Foundation Model cites this paper.

Seedream 2.0: A Native Chinese-English Bilingual Image Generation Foundation Model Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-17T08:27:36.309996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-17T08:27:36.242416Z digest=sha256:b71979ae41ce8a1ba0a5a57e1afd730407c58f1ab457feec75a6b82343e68673

Observation 7ac6b9e6-7cf5-4f49-9e6d-cb4c928b463f · inbound

Instruction-augmented Multimodal Alignment for Image-Text and Element Matching cites this paper.

Instruction-augmented Multimodal Alignment for Image-Text and Element Matching Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T12:40:26.523312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:40:26.523312Z digest=sha256:d4473625f1b5766942ff0d004a6989a091e9bdb5318c74cdb6860bb818d27b27

Observation dc9f41c5-f5f1-47ab-b673-6f68df039d55 · inbound

Understanding Attention Mechanism in Video Diffusion Models cites this paper.

Understanding Attention Mechanism in Video Diffusion Models Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T12:40:31.264007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:40:31.264007Z digest=sha256:c6e334d69ae82af12de5aaa9bd6b898ddcd6472146e89f469d8960cefc36d375

Observation 19f56009-acaa-4905-9dd4-7fb6034a54fe · inbound

SUDO: Enhancing Text-to-Image Diffusion Models with Self-Supervised Direct Preference Optimization cites this paper.

SUDO: Enhancing Text-to-Image Diffusion Models with Self-Supervised Direct Preference Optimization Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T11:49:17.767084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:49:17.767084Z digest=sha256:67a7e665ab779809bd04a545c7758c6c45ae1dd2b686aa4ffa8ea53d39312417

Observation 2dea6047-9081-4053-ae64-92d6faf54234 · inbound

From Reflection to Perfection: Scaling Inference-Time Optimization for Text-to-Image Diffusion Models via Reflection Tuning cites this paper.

From Reflection to Perfection: Scaling Inference-Time Optimization for Text-to-Image Diffusion Models via Reflection Tuning Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T11:15:34.933445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:15:34.933445Z digest=sha256:b033164ad20e1c5adfee6ab8b5567693154411ca06858bd6b0c48bcc4565d14c

Observation 5ccc9090-6d8b-4dcb-bc33-982fb70a9314 · inbound

Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation cites this paper.

Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-05-17T07:24:04.875390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-17T07:24:04.460276Z digest=sha256:d31b7628004969a9a56c084f8b32f3742d8766b78eecab71d3ed1c221906c35b

Observation e11c36e8-7e6e-4681-8621-271b32ffcb65 · inbound

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation cites this paper.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:46.313528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:46.313528Z digest=sha256:f86e4c24f28b3c5c03604cf2288c8c2831f251489631e0faa82f2eda21b9caf8

Observation 4cd9816c-ceae-437d-a5f7-7a33db577638 · inbound

Hunyuan-Game: Industrial-grade Intelligent Game Creation Model cites this paper.

Hunyuan-Game: Industrial-grade Intelligent Game Creation Model Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:31.425024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:31.425024Z digest=sha256:e268d6a9d1ce2c576b82cc12a7e5ab3067b8426eac9de33b550e792f3e874a61

Observation 9bbb5788-a59c-4612-a250-14e6e2bce774 · inbound

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation cites this paper.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:53.968678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:53.968678Z digest=sha256:b016f0aaf1015c316d97c84c70f17535e8cdacbb03609c52f2f3ec915c97cf15

Observation 25088bc4-a1f9-4f6b-90a1-acd6d0d13a56 · inbound

Harnessing Caption Detailness for Data-Efficient Text-to-Image Generation cites this paper.

Harnessing Caption Detailness for Data-Efficient Text-to-Image Generation Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:13.390675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:13.390675Z digest=sha256:7ed43003668a03014d132ccab81e3a0b50673eb1f5308de65b2d3ce5c8edf81e

Observation 3a1a8f78-555b-489e-b3db-a88b273631c5 · inbound

OmniConsistency: Learning Style-Agnostic Consistency from Paired Stylization Data cites this paper.

OmniConsistency: Learning Style-Agnostic Consistency from Paired Stylization Data Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:08.431458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:08.431458Z digest=sha256:36a9f1f01993820ea84067ff62ed3ba9389d56a14a711acf06da9a7255a5f4d0

Observation 7464fcef-16b5-4120-bebc-5a2d6fff5501 · inbound

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models cites this paper.

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:34.926538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:34.926538Z digest=sha256:fb5ca85f77688c75276d98042574f7f8dfebca0862b2d092feb53c96aeb1ba91

Observation a4c39549-f37a-4bdc-b515-ceb6be0a7050 · inbound

Multimodal LLM-Guided Semantic Correction in Text-to-Image Diffusion cites this paper.

Multimodal LLM-Guided Semantic Correction in Text-to-Image Diffusion Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:06:47.850143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:06:47.850143Z digest=sha256:dc9a653dc10a351b48740cba52521bc080a5d4b3cc37afe37b36aac54a4cbcab

Observation 586e70c8-b0c4-4609-b585-1b9f4cbb586b · inbound

FUDOKI: Discrete Flow-based Unified Understanding and Generation via Kinetic-Optimal Velocities cites this paper.

FUDOKI: Discrete Flow-based Unified Understanding and Generation via Kinetic-Optimal Velocities Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 130

Resolution
unresolved
no resolver link, observed 2026-08-07T14:05:06.764006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:05:06.764006Z digest=sha256:c7177954eb6e7a5dee578bd844bf97ab0146ef0c97cb15339c94bb0315b8032c

Observation cfd03ea0-5c34-4716-b360-59d3162780ea · inbound

AniCrafter: Customizing Realistic Human-Centric Animation via Avatar-Background Conditioning in Video Diffusion Models cites this paper.

AniCrafter: Customizing Realistic Human-Centric Animation via Avatar-Background Conditioning in Video Diffusion Models Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:10.567223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:10.567223Z digest=sha256:02cd9be2346efdf508d57c99279315c8f36a9e23de74e4a957aecd4412981770

Observation 2077f18a-effd-48eb-be04-092b0ee7208c · inbound

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation cites this paper.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:16.037523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:16.037523Z digest=sha256:d9d7c3a2183e270f9363ce986f4482823230920b469c61fc41c601f26cbd3611

Observation 0ba3e372-9816-449e-8acf-d4a831eb2ab6 · inbound

GenSpace: Benchmarking Spatially-Aware Image Generation cites this paper.

GenSpace: Benchmarking Spatially-Aware Image Generation Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:21:23.182573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:21:23.182573Z digest=sha256:5d86384da7c6dab576e6495873ccee358c8175a76f8569f900c9a4dd49e5f6d4

Observation 90fc9349-56ef-4074-ae6e-9e5335767cba · inbound

NoiseAR: AutoRegressing Initial Noise Prior for Diffusion Models cites this paper.

NoiseAR: AutoRegressing Initial Noise Prior for Diffusion Models Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T11:51:33.213194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:51:33.213194Z digest=sha256:aab0f45a3270aa25582d1f0e014564ca38e0353be42c41115234fbb27e799ef9

Observation 33884d4d-d15d-40be-8aad-eb28daf44fed · inbound

OmniV2V: Versatile Video Generation and Editing via Dynamic Content Manipulation cites this paper.

OmniV2V: Versatile Video Generation and Editing via Dynamic Content Manipulation Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T11:39:52.770522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:39:52.770522Z digest=sha256:7b73149a4a4f2beaae92d677021b98a763827564eb1e93550af222e9370e80f6

Observation ce85b1c3-0fda-4b7a-b908-926aba5f9c5f · inbound

Marrying Autoregressive Transformer and Diffusion with Multi-Reference Autoregression cites this paper.

Marrying Autoregressive Transformer and Diffusion with Multi-Reference Autoregression Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T04:52:55.047546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:52:55.047546Z digest=sha256:b9e77866c09e91630e86cc61f2ba450db402e2ac7e0d5507bae0c7ac35bc5102

Observation 4158dac0-cec8-4cd3-bf99-534b47953d49 · inbound

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer cites this paper.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T00:40:59.660776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:40:59.660776Z digest=sha256:0d9442be8a57560ec9faa0b6eba8f52f4154d1f0ca4a3a9e7c76c4e3a31646b9

Observation b75c20ac-9734-4f78-8c5f-350876d716c1 · inbound

COME: Adding Scene-Centric Forecasting Control to Occupancy World Model cites this paper.

COME: Adding Scene-Centric Forecasting Control to Occupancy World Model Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T00:41:30.764087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:41:30.764087Z digest=sha256:c361875c5f809b588d3c2aae85f45258facee76588862220116bdbbeebf127bd

Observation 7ba124ae-8510-4cb4-890d-62a2b53796e0 · inbound

ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation cites this paper.

ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T23:28:05.645405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:28:05.645405Z digest=sha256:2e930cf01f5d7b11ad3c1db430c5fb55a4434009151204d9921408714d2a7f0d

Observation c9305dfa-6c32-4c91-ab1e-c02c83693a8d · inbound

OmniGen2: Towards Instruction-Aligned Multimodal Generation cites this paper.

OmniGen2: Towards Instruction-Aligned Multimodal Generation Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-05-19T07:52:10.879158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:c11e9c8514eb822db2dbc6891bdea4f561471859c9860cac43621ae6734a9044

Observation ca526121-fa56-4e51-8f05-80d22f967955 · inbound

Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations cites this paper.

Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:10.710650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:10.710650Z digest=sha256:4ac5a2fdeb1b7608d2b00d5e5a73b8f7c208adf6ba7f39e45a64dd3004ff1b08

Observation a179ed4e-2cdc-4e92-a040-54e4580ed319 · inbound

Ctrl-Z Sampling: Scaling Diffusion Sampling with Controlled Random Zigzag Explorations cites this paper.

Ctrl-Z Sampling: Scaling Diffusion Sampling with Controlled Random Zigzag Explorations Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T22:59:14.172131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:59:14.172131Z digest=sha256:a87ffeb8d84a97d1b8f835ce116851ede4e7237d13b37efeaec9ea256190140f

Observation 83f93069-964f-43f1-bc48-b538e56401bb · inbound

Med-Art: Diffusion Transformer for 2D Medical Text-to-Image Generation cites this paper.

Med-Art: Diffusion Transformer for 2D Medical Text-to-Image Generation Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T22:53:13.531583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:53:13.531583Z digest=sha256:d5688139f7405ffd6fff38234270436a33c68b7afeaadb2379ebbe5043141693

Observation 887d158b-933f-4a91-9be9-639b82c323dd · inbound

DFVEdit: Conditional Delta Flow Vector for Zero-shot Video Editing cites this paper.

DFVEdit: Conditional Delta Flow Vector for Zero-shot Video Editing Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T22:42:45.233124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:42:45.233124Z digest=sha256:9fd5916fefe40e5e54541ee53a83c46c21cf72e7a40e622b8a4f424cd35a6d66

Observation 1011fda0-9765-4d35-9bc8-af0f014b05e4 · inbound

On the Feasibility of Poisoning Text-to-Image AI Models via Adversarial Mislabeling cites this paper.

On the Feasibility of Poisoning Text-to-Image AI Models via Adversarial Mislabeling Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T22:23:42.916538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:23:42.916538Z digest=sha256:cf1bddf099dc215d0e3ceaf23d7bff1f975f390df77a539512e3fea97e97b870

Observation 01613cf8-1f4b-4777-9245-2199ba8b590b · inbound

UniMC: Taming Diffusion Transformer for Unified Keypoint-Guided Multi-Class Image Generation cites this paper.

UniMC: Taming Diffusion Transformer for Unified Keypoint-Guided Multi-Class Image Generation Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T20:26:37.875779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:26:37.875779Z digest=sha256:53c3ea316d83abcde00d236117b5efcda82f8c8e8b7ef35492434d1c7dd00a26

Observation 8e175b62-cf9c-4935-acea-0f58edf05a1d · inbound

CoT-lized Diffusion: Let's Reinforce T2I Generation Step-by-step cites this paper.

CoT-lized Diffusion: Let's Reinforce T2I Generation Step-by-step Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T19:55:02.899597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:55:02.899597Z digest=sha256:04b4ddf04f345aa54bedde5157178d0819cd3560c305019bd69a7a9169322330

Observation 131fe098-44f5-4301-a768-3fb169eca97f · inbound

CharaConsist: Fine-Grained Consistent Character Generation cites this paper.

CharaConsist: Fine-Grained Consistent Character Generation Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T17:10:55.041081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:10:55.041081Z digest=sha256:ca7f99f328db31e350f16d36507e68eb986bb30e039cb1e21ccc8b229080bb01

Observation cc134c95-a660-44d6-be19-d5fa8236b61e · inbound

AnimeColor: Reference-based Animation Colorization with Diffusion Transformers cites this paper.

AnimeColor: Reference-based Animation Colorization with Diffusion Transformers Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T13:47:46.235477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:47:46.235477Z digest=sha256:e60d5438c3f6ed88083730caab9ed4eb56699d047f496e0d0479e34e28c7c508

Observation f39972e5-3f49-4445-ae4a-8fa12728a4a4 · inbound

T2I-Copilot: A Training-Free Multi-Agent Text-to-Image System for Enhanced Prompt Interpretation and Interactive Generation cites this paper.

T2I-Copilot: A Training-Free Multi-Agent Text-to-Image System for Enhanced Prompt Interpretation and Interactive Generation Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T17:47:25.587462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:47:25.587462Z digest=sha256:42c5f48599531fdcbde8b8de9ec7370b8219fa5081299804d8c85f0655a23d8d

Observation 0a12d2be-2083-4d49-a18a-e96e4b528ae4 · inbound

ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation cites this paper.

ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T05:59:15.539301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:59:15.539301Z digest=sha256:2c00296f504a4c05a792b961dca0f43ab3be045d3f748d0a56709e2324f7e18f

Observation d6a96f96-3294-407b-ac8c-dd8960487d9b · inbound

A Summer Meridional Subsurface Temperature Dipole Mode in the South China Sea cites this paper.

A Summer Meridional Subsurface Temperature Dipole Mode in the South China Sea Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T04:44:57.916258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:44:57.916258Z digest=sha256:d3ab965556b38e7378460b901aabfcffb9040f945dd764b664922d10892fbbda

Observation e7854012-5878-4145-a2f3-f0b127f40fdd · inbound

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation cites this paper.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:13.554546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:13.554546Z digest=sha256:44d52b704fa58b266c313e0efbfaee271e82ebc27f7e49ef29c034b96483a5ba

Observation cf68c54a-d28b-4b4e-887a-decbab4b14fc · inbound

Transition Models: Rethinking the Generative Learning Objective cites this paper.

Transition Models: Rethinking the Generative Learning Objective Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T10:19:54.393763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:19:54.393763Z digest=sha256:2532d3e9788c308f3871adc96416841934b617e9c58cef3fd7956e22bbe791dd

Observation 9ca1c93a-5adc-46ee-935f-5a453af8e7e3 · inbound

Home-made Diffusion Model from Scratch to Hatch cites this paper.

Home-made Diffusion Model from Scratch to Hatch Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T04:34:36.974733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:34:36.974733Z digest=sha256:7b2168a7eac9af54dc73a4b05bf30c5c741c0849439f357848975f3cce672c1b

Observation 2aea1fa6-9b79-49d5-af67-7cdfd5a622c4 · inbound

Directly Aligning the Full Diffusion Trajectory with Fine-Grained Human Preference cites this paper.

Directly Aligning the Full Diffusion Trajectory with Fine-Grained Human Preference Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T22:58:08.801305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:58:08.801305Z digest=sha256:632a72f27b4c456347bdd287742c6db7dec9d97a59d1b00a187d83c7dff88f6e

Observation 0356551c-13d7-480c-8ac5-f129a797c09f · inbound

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark cites this paper.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.745641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.745641Z digest=sha256:10175908f1e85cdf38873cfe1f4c0028a56099b56e10b9b917e947cd524df6c2

Observation f70f34df-e7cb-4993-9ec6-9e54491e60ec · inbound

Compute Only 16 Tokens in One Timestep: Accelerating Diffusion Transformers with Cluster-Driven Feature Caching cites this paper.

Compute Only 16 Tokens in One Timestep: Accelerating Diffusion Transformers with Cluster-Driven Feature Caching Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T16:00:04.618874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:00:04.618874Z digest=sha256:fb12e7499ed0be1a0703f462fd0622329adefcf215b7bebd2404f38525636f04

Observation 4560414f-06f5-4d1a-aada-954181a4b0ea · inbound

HunyuanImage 3.0 Technical Report cites this paper.

HunyuanImage 3.0 Technical Report Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-16T14:58:37.566044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T02:02:32.806844Z digest=sha256:8746f82da605e62024fc507d8a6abad80488088ec25ac8f2ee298ed06348a4dc

Observation d4482470-c6fe-4ea1-985e-0cb3fa272087 · inbound

HunyuanImage 3.0 Technical Report cites this paper.

HunyuanImage 3.0 Technical Report Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T14:42:14.415987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:42:14.415987Z digest=sha256:f86653a72faefbcf8cc479b73dc5c1ce96ca1bdd2e5c77d1935a9e2b769eea7d

Observation 37ccd311-6343-46c1-921a-b5e96caee0ab · inbound

Few-Shot Synthetic Image Attribution: Identifying Unseen Generators with Limited Samples cites this paper.

Few-Shot Synthetic Image Attribution: Identifying Unseen Generators with Limited Samples Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T13:44:08.667828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:44:08.667828Z digest=sha256:631b473d6ed680cdcdf15b65847d421d64468f63701d154f8c6ce19b2d742070

Observation 59025a7d-7553-453e-b225-0d57ba994af1 · inbound

Uniworld-V2: Reinforce Image Editing with Diffusion Negative-aware Finetuning and MLLM Implicit Feedback cites this paper.

Uniworld-V2: Reinforce Image Editing with Diffusion Negative-aware Finetuning and MLLM Implicit Feedback Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-21T18:01:19.828310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-21T18:01:19.748677Z digest=sha256:e3afc944fce8a7e5b435ab0179aded0b2519795c44cf0a2c388a197a5e6cbc48

Observation 001e6a19-564b-40af-aef8-eb00406fd66b · inbound

Phase-Aligned RoPE for Mixed-Resolution Diffusion Transformer cites this paper.

Phase-Aligned RoPE for Mixed-Resolution Diffusion Transformer Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-03T20:30:31.643191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:30:31.643191Z digest=sha256:7710f529de1b506a256311c90766067eb53bd31bb88722ab297523696cfa8530

Observation c8596c85-a782-40fc-af9d-409f28cb433c · inbound

PixelDiT: Pixel Diffusion Transformers for Image Generation cites this paper.

PixelDiT: Pixel Diffusion Transformers for Image Generation Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-05-17T04:31:31.175892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-17T04:30:07.417197Z digest=sha256:e77d40442ad1d7f4d55ad249e79c7617db4545f21c2b91b29a50dca7644b18ee

Observation 5a541b64-c887-40fd-9858-c92028ea12a4 · inbound

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer cites this paper.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-16T14:58:37.566044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-11T14:08:36.801359Z digest=sha256:95ac16c1fe8ce7b71e4dcce915bdff69b43261c51f84baf35e52f7f909643df2

Observation 31537b14-4724-42bc-9266-d28b5e4755c2 · inbound

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer cites this paper.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:32.749977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:32.749977Z digest=sha256:97d018a6021fc37c94c0a60a5bbb182b4fbd411b411fe8c14abcf8a0507da44c

Observation c0bbf535-9216-4c07-8dea-6120b2872110 · inbound

Probing and Leveraging Video Diffusion Transformer Features for Robust Point Tracking cites this paper.

Probing and Leveraging Video Diffusion Transformer Features for Robust Point Tracking Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-03T14:24:49.959310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:24:49.959310Z digest=sha256:0805b2947cca3bc6e42f816a96e2320763803a09f14f408153f5f8984d0911b7

Observation 86bdbf9d-1baa-4439-b1af-bf37f8db312f · inbound

Guiding Token-Sparse Diffusion Models cites this paper.

Guiding Token-Sparse Diffusion Models Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-03T12:50:57.283172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:50:57.283172Z digest=sha256:d1bb5c0dcf4fb8ec49165bf20502b08b8b184e0df399ba77be14eca81d27b13b

Observation caba5c44-c215-4525-95ee-5dcf5f2a0fab · inbound

TexTailor: Inference-Time Textual Guidance Tailoring for Multimodal Diffusion Transformers cites this paper.

TexTailor: Inference-Time Textual Guidance Tailoring for Multimodal Diffusion Transformers Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T12:39:55.224838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:39:55.224838Z digest=sha256:b2aaf579f30d089ed749bb8086c3ae68cb72981cb5eb45f8515a6bf1cdcefab9

Observation 21edd1ef-12eb-47a5-9ea5-5c5afc398638 · inbound

The Algorithmic Gaze of Image Quality Assessment: An Audit and Trace Ethnography of the LAION-Aesthetics Predictor cites this paper.

The Algorithmic Gaze of Image Quality Assessment: An Audit and Trace Ethnography of the LAION-Aesthetics Predictor Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-16T14:58:37.566044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T14:03:00.513248Z digest=sha256:74c1046e19d071c5032d806e46d4765b7eccf9f89ed17b0bcb48bffb7dea0958

Observation 59d9d0a1-507c-4d17-a774-c773e8bfe673 · inbound

Happy Young Women, Grumpy Old Men? Emotion-Driven Demographic Biases in Synthetic Face Generation cites this paper.

Happy Young Women, Grumpy Old Men? Emotion-Driven Demographic Biases in Synthetic Face Generation Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-03T09:52:19.476716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:52:19.476716Z digest=sha256:e23ec5bd3ed05072cd369c73a39ac1f96377edde4a12edb765a451b91ffa9296

Observation 2345c93a-837e-4460-9f5c-7cc8a2a04a41 · inbound

OmniFysics: Towards Physical Intelligence Evolution via Omni-Modal Signal Processing and Network Optimization cites this paper.

OmniFysics: Towards Physical Intelligence Evolution via Omni-Modal Signal Processing and Network Optimization Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 77

Resolution
verified exact
arxiv_id, observed 2026-05-16T14:58:37.566044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T07:09:46.254851Z digest=sha256:478f9ab9645926c57dd523bee3109d84fd98103825f4532daf6a068c6e09935d

Observation c127b862-4de1-4325-8945-a6b1fdd850c9 · inbound

When Numbers Speak: Aligning Textual Numerals and Visual Instances in Text-to-Video Diffusion Models cites this paper.

When Numbers Speak: Aligning Textual Numerals and Visual Instances in Text-to-Video Diffusion Models Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-16T14:58:37.566044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T18:39:26.420463Z digest=sha256:3820d67ec922825ef1f1ed91a1bf084060bcc0d8cd5873d977204cfd9ad66a87

Observation e1e45613-ed44-4f95-b3ec-b929758b557b · inbound

CineAGI: Character-Consistent Movie Creation through LLM-Orchestrated Multi-Modal Generation and Cross-Scene Integration cites this paper.

CineAGI: Character-Consistent Movie Creation through LLM-Orchestrated Multi-Modal Generation and Cross-Scene Integration Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-16T14:58:37.566044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T05:08:37.971891Z digest=sha256:82a914bf6407ce6a4c90b0e22ba44333ba240ec789f0e6c0d0cb78ebbd1dd274

Observation 6ff9371d-2cf4-4076-a9de-44dc210e2828 · inbound

Diffusion Templates: A Unified Plugin Framework for Controllable Diffusion cites this paper.

Diffusion Templates: A Unified Plugin Framework for Controllable Diffusion Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-16T14:58:37.566044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T04:24:53.023047Z digest=sha256:bbeb08c8359c402a9f716e80f58fbef9dd8ecb9b945736b3b3d94fa1a50dc686

Observation 25a3dff3-6d9e-4f6a-9200-69e4c2d47ff8 · inbound

ACPO: Anchor-Constrained Perceptual Optimization for Diffusion Models with No-Reference Quality Guidance cites this paper.

ACPO: Anchor-Constrained Perceptual Optimization for Diffusion Models with No-Reference Quality Guidance Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-16T14:58:37.566044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-07T14:05:05.560294Z digest=sha256:774b5c3ea5e33748bfb9d33042ace763e5a66ef33c91faef9cdfd2fcad6d8b5e

Observation ae66cecb-b76a-418f-85d1-e2da5783fb5d · inbound

Beyond Fixed Formulas: Data-Driven Linear Predictor for Efficient Diffusion Models cites this paper.

Beyond Fixed Formulas: Data-Driven Linear Predictor for Efficient Diffusion Models Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-16T14:58:37.566044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-07T13:51:23.367393Z digest=sha256:be8c87272e51a0f2c73a134b569c36370046eda49624c5fddce3380eac61fdf5

Observation 9717243b-4342-4726-a454-7c5b71be061e · inbound

Leveraging Verifier-Based Reinforcement Learning in Image Editing cites this paper.

Leveraging Verifier-Based Reinforcement Learning in Image Editing Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-16T14:58:37.566044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-07T08:00:33.307429Z digest=sha256:4f7c4c7ca0b081d736d4b86cf71c5e7e1b004047c3181152b473047347885f76

Observation c6b71185-2754-4f18-9e0b-4ec7af0773b4 · inbound

Leveraging Verifier-Based Reinforcement Learning in Image Editing cites this paper.

Leveraging Verifier-Based Reinforcement Learning in Image Editing Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-21T09:14:05.955534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-21T09:11:02.183133Z digest=sha256:d8e8713bc0956b89f91d8b22aba5c99634fd0a885a831611314174c6ca8da4f8

Observation 6e4765f9-5092-4f08-b420-40ba1ec391cd · inbound

Fashion130K: An E-commerce Fashion Dataset for Outfit Generation with Unified Multi-modal Condition cites this paper.

Fashion130K: An E-commerce Fashion Dataset for Outfit Generation with Unified Multi-modal Condition Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-16T14:58:37.566044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-12T03:56:43.640874Z digest=sha256:f820cd777993b58ccab483d61d1c4172dee1d5b7d534d97e800dc4d370ba7589

Observation 48430f21-1021-4096-b0ec-2ca5c20ff89a · inbound

Fashion130K: An E-commerce Fashion Dataset for Outfit Generation with Unified Multi-modal Condition cites this paper.

Fashion130K: An E-commerce Fashion Dataset for Outfit Generation with Unified Multi-modal Condition Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-16T14:58:37.566044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T21:56:10.499014Z digest=sha256:ed30b4086bd9daa692d0cf36d0a372f0b14403865add8b73d3422068a8fac2db

Observation b4877599-743e-4b84-b4c6-fa145bfe051b · inbound

Qwen-Image-2.0 Technical Report cites this paper.

Qwen-Image-2.0 Technical Report Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-16T14:58:37.566044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-12T04:21:18.312613Z digest=sha256:3d2e98f87f816498e1c0b1d07d9abd1326135bc3611b9d3cbcef659c3eb27b9c

Observation 9d5a4e95-03ff-4132-8a32-06f8e89223cb · inbound

HiDream-O1-Image: A Natively Unified Image Generative Foundation Model with Pixel-level Unified Transformer cites this paper.

HiDream-O1-Image: A Natively Unified Image Generative Foundation Model with Pixel-level Unified Transformer Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T14:58:37.566044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-13T07:30:53.939221Z digest=sha256:28d8ac477a9a2165ac4fba1a93ffae24f7508012154ae4d0b854c37c7dcd17ac

Observation 732da0c5-a6ad-416c-b3d1-10d99fddbe1c · inbound

Asymmetric Flow Models cites this paper.

Asymmetric Flow Models Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-16T14:58:37.566044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T19:28:21.625879Z digest=sha256:188e7687a75a2529667980614cc7294aef53b2759c6403e8b6597ce52000a4d7

Observation f164c029-7467-4a9b-bf47-e53501ccf378 · inbound

Asymmetric Flow Models cites this paper.

Asymmetric Flow Models Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:15:47.401196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T22:07:44.850763Z digest=sha256:614d0668f98ee2b7355341ed60e3213321c7a860122f155f006548ad74ccf258

Observation 1a4b0297-079b-4eb6-b530-1bcf3220cd70 · inbound

ElasticDiT: Efficient Diffusion Transformers via Elastic Architecture and Sparse Attention for High-Resolution Image Generation on Mobile Devices cites this paper.

ElasticDiT: Efficient Diffusion Transformers via Elastic Architecture and Sparse Attention for High-Resolution Image Generation on Mobile Devices Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 40

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T20:03:43.903399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-20T20:00:27.987481Z digest=sha256:c6e96c70106df559a37ddea81c71231561e827cd7dfc68460b1858a39943dcd6