Pith. sign in

Paper Citation Record · LEDGER

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

As of 5 August 2026, this Paper Citation Record lists 100 of 100 outbound references and 100 inbound Pith citation observations for arXiv:2511.22699.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2511.22699 v5

Coverage vector

measured 100 of 100 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T19:47:32.986689Z

measured 200 of 200 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 100 of 153 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T21:47:20.091819Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

100 of 100 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved100
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 28fae302-c7f4-43ae-a7e9-4fbb34c873d5 · outbound

This paper cites Pytorch 2: Faster machine learning through dynamic python bytecode transformation and graph compilation.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer Pytorch 2: Faster machine learning through dynamic python bytecode transformation and graph compilation

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:30.475102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:30.475102Z digest=sha256:c422723d19af4792b616cb2a37be9290b5064679798d3398ae0446da26ec9086

Observation cde753cb-7cd0-4565-b9ec-cfa708e46925 · outbound

This paper cites Qwen2.5-VL Technical Report.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer Qwen2.5-VL Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:30.562252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:30.562252Z digest=sha256:3aa06e1c0e58ae68cf6122dd1e6b3fd9e39f555d7298abc3621b52c78cff47b6

Observation 5de9258b-3273-4120-a708-48d4d271cd6b · outbound

This paper cites Imagen 3.arXiv preprint arXiv:2408.07009, 2024.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer Imagen 3.arXiv preprint arXiv:2408.07009, 2024

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:30.646391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:30.646391Z digest=sha256:8074fc0cbf0b22a10ae7b4f769ef4f05c217b5d742245c4ed4f649c9d6596746

Observation 3f74fa1a-f1c6-4c0c-ab2d-0aee17c7ba9f · outbound

This paper cites Improving image generation with better captions.Computer Science.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer Improving image generation with better captions.Computer Science

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:30.700289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:30.700289Z digest=sha256:11abfeaa72c0712b9327a952959eb26c155ba4c2208235982e2d694ee3bcd8b4

Observation 05e2d8b6-848d-4e48-bb2d-afeba793fc31 · outbound

This paper cites Instructpix2pix: Learning to follow im- age editing instructions.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer Instructpix2pix: Learning to follow im- age editing instructions

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:30.729213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:30.729213Z digest=sha256:eaaec565c73fa3329ecafb9de51974d21eba9809fa2f8d4163dad84f011445cb

Observation 441b816e-9fc6-4508-9df1-b6cc05f13de3 · outbound

This paper cites Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901, 2020.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901, 2020

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:30.809386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:30.809386Z digest=sha256:c641595a591de94b8ac7101bb5e76fc00e2c630b49fa973c76c76a724f3bd34d

Observation e87f4a65-4f4d-42fc-8eb7-0e0436d52bb8 · outbound

This paper cites Hidream-i1: An open-source high-efficient image generative foundation model.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer Hidream-i1: An open-source high-efficient image generative foundation model

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:30.849396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:30.849396Z digest=sha256:e3fea449a9a304e7bceaee5c13cf97570ebb40fb564b568c1dec7c8f04d42af4

Observation 45847ee5-85b5-46d0-a19c-43a7daade548 · outbound

This paper cites HunyuanImage 3.0 Technical Report.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer HunyuanImage 3.0 Technical Report

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:30.877338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:30.877338Z digest=sha256:94aaf9a289aac5c93a24bf8a0cf31121a73a29fc25bbb219eb09847d8d032b9f

Observation 245f5d62-2469-4ceb-8d6f-5b2ee348c88f · outbound

This paper cites OneIG-Bench: Omni-dimensional Nuanced Evaluation for Image Generation.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer OneIG-Bench: Omni-dimensional Nuanced Evaluation for Image Generation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:30.934626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:30.934626Z digest=sha256:219d140132f6c3eabba504d45189dc2d306d411a8e6bb929c17cfaeefcc8bb10

Observation 5f576ab7-7a74-42e1-8cf8-69bfd11d612f · outbound

This paper cites Textdiffuser-2: Unleashing the power of language models for text rendering.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer Textdiffuser-2: Unleashing the power of language models for text rendering

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:30.990759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:30.990759Z digest=sha256:da9fedb35805e0d3699f69db4f6e16e2f2cc210dcb4c52876eb3f8ae892febae

Observation 29386781-04fa-4eef-8e9a-82c97c78ff74 · outbound

This paper cites BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:31.024621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:31.024621Z digest=sha256:5e67f12695ada0bf58c8f915ae3ef1ac7da23543a44aaa942e689268433f5920

Observation 3e54ed50-4cc9-4c0d-b555-97c2e576da47 · outbound

This paper cites Pixart-𝜎: Weak-to-strong training of diffusion transformer for 4k text-to-image generation.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer Pixart-𝜎: Weak-to-strong training of diffusion transformer for 4k text-to-image generation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:31.057684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:31.057684Z digest=sha256:ba1080e3aada84e8be34bffe6b7f3c0b22a260b1cb1f486cfefe91b7d2fcacc8

Observation ba09161e-606e-4c5a-b74b-49f3b23ef8b6 · outbound

This paper cites Pixart-𝛼: Fast training of diffusion transformer for photo- realistic text-to-image synthesis.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer Pixart-𝛼: Fast training of diffusion transformer for photo- realistic text-to-image synthesis

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:31.091651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:31.091651Z digest=sha256:80fb5e2a04938ece88dda2b96f5f97aa8006cf3c247e8864a50cbab5a6146fcc

Observation f8562c6d-748d-455c-983f-683e9c065ecb · outbound

This paper cites Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:31.126783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:31.126783Z digest=sha256:af6db73b1fe8df1f443a847c5d2b0664d4278a4b4b0bcc4b95e43aa0c3ac6415

Observation 373e2210-19fc-4629-b3b0-f8710f8983a8 · outbound

This paper cites Emerging Properties in Unified Multimodal Pretraining.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer Emerging Properties in Unified Multimodal Pretraining

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:31.160663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:31.160663Z digest=sha256:10bbde7e3cf3ac67eecf64459e4d3af879858207881111a3f394ab648cf2fbd6

Observation 126de000-b4a2-4adb-ad18-ba6e62d6b30f · outbound

This paper cites Cogview: Mastering text-to-image generation via transformers.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer Cogview: Mastering text-to-image generation via transformers

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:31.198772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:31.198772Z digest=sha256:7bf6a4dbd8cbcbcdea066ec5b137783a085041ec91940d25808081eb7029b8c2

Observation 6e659890-00c2-4f33-a20c-2198ba661e01 · outbound

This paper cites Textcrafter: Accurately rendering multiple texts in complex visual scenes.arXiv preprint arXiv:2503.23461, 2025.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer Textcrafter: Accurately rendering multiple texts in complex visual scenes.arXiv preprint arXiv:2503.23461, 2025

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:31.258509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:31.258509Z digest=sha256:c6f38740f7a585d4193b9f69d4658bcf95c3d0ed424311c2dce377a51f6c7bd7

Observation 53875729-1899-4cd1-90d6-e6f71402d439 · outbound

This paper cites Scaling rectified flow transformers for high-resolution image synthesis.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer Scaling rectified flow transformers for high-resolution image synthesis

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:31.312423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:31.312423Z digest=sha256:37acdbbaafd2b48d41843444277fe00e8be8a1b3d8ccd18db8139d6cf4a3a63e

Observation 6b6f9400-4ee2-466a-9de1-97b015cb413c · outbound

This paper cites FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:31.327318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:31.327318Z digest=sha256:aba704fb018a3301dbd8bb38c5761feacc1e00a89cb4c3c7e88c820070729e76

Observation 0af520a0-f9d2-4d3e-adfc-e940073ddea5 · outbound

This paper cites Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:31.331517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:31.331517Z digest=sha256:1c9ed634e69bff1a6ecc74405cc5e716f2186e31b1d6c899c04c2a5f7d557348

Observation 261c5bf2-d0d4-42ee-ac7a-bfbfcf415f0e · outbound

This paper cites Seedream 3.0 Technical Report.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer Seedream 3.0 Technical Report

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:31.335779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:31.335779Z digest=sha256:88ad194cb2fb4e4875f9a0de525d70aabcd5a325b492578e3d9199c82a669fa9

Observation cfd7ea82-2d80-4696-9395-a33b2ee4b16c · outbound

This paper cites X-Omni: Reinforcement Learning Makes Discrete Autoregressive Image Generative Models Great Again.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer X-Omni: Reinforcement Learning Makes Discrete Autoregressive Image Generative Models Great Again

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:31.339897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:31.339897Z digest=sha256:763d45210f40a341d106beda0956d0d7dd3afef3438f12845caa8073e168683c

Observation 680d4670-3c35-4ffc-ba01-a85f0445ecbd · outbound

This paper cites Geneval: An object-focused framework for evaluating text-to-image alignment.Advances in Neural Information Processing Systems, 36:52132– 52152, 2023.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer Geneval: An object-focused framework for evaluating text-to-image alignment.Advances in Neural Information Processing Systems, 36:52132– 52152, 2023

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:31.343751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:31.343751Z digest=sha256:c29dfb13db952f5c1d6a145c3d84b0fc8a09f68dc985d877b4d446169564514c

Observation 17f4bde7-669c-4abb-b77a-2825ae4b551b · outbound

This paper cites Dynamic few-shot visual learning without forgetting.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer Dynamic few-shot visual learning without forgetting

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:31.347436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:31.347436Z digest=sha256:9bd40760d9b817ecf3e0e3c06aa8225829bb1c95d4cb35ce9b296589396eee64

Observation c26176ae-26c0-477f-b2e7-8a1b698acdcb · outbound

This paper cites Gemini 2.5 flash & 2.5 flash image model card.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer Gemini 2.5 flash & 2.5 flash image model card

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:31.362336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:31.362336Z digest=sha256:0d093fd591989937d7f9685dcd7831888408fe6958783c4ead57809cfe72bc6b

Observation 6d1afd41-2448-41c7-a0df-bede08690930 · outbound

This paper cites Imagen 4 model card.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer Imagen 4 model card

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:31.413630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:31.413630Z digest=sha256:ca1c61b83ca2705dddfe32a187f5823b6bda7880913e9590aec0a1f7c21da694

Observation 2b28f84b-170d-4a92-9565-b3b7a9771eb1 · outbound

This paper cites Nano banana pro.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer Nano banana pro

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:31.567733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:31.567733Z digest=sha256:36cc069ff204e316b70b9e8fc2a5e7d0b67eaf10437bc1020659ffe88c4a2042

Observation d860859d-0ede-4f11-a591-f323e7dc2c55 · outbound

This paper cites Infinity: Scaling bitwise autoregressive modeling for high-resolution image synthesis.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer Infinity: Scaling bitwise autoregressive modeling for high-resolution image synthesis

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:31.722054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:31.722054Z digest=sha256:c4b22cd2e24aa5d8024d5b4de69c7d120a7f869de8805252a3864a2bee075664

Observation 36508c3d-f09f-4e79-a116-2b97fa94d97c · outbound

This paper cites Classifier-free diffusion guidance.Advances in Neural Information Processing Systems Workshops (NeurIPS Workshops), 2021.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer Classifier-free diffusion guidance.Advances in Neural Information Processing Systems Workshops (NeurIPS Workshops), 2021

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:31.797853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:31.797853Z digest=sha256:2393dd6bd7f4b325b5b541ff6f5b12acaeae472089ad072cc6c3342fff842db8

Observation 247dc9e9-a06c-46c1-9fbd-754140edaeff · outbound

This paper cites ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:31.943799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:31.943799Z digest=sha256:13a0a06c8cf6d311ebf5b79eb790ee27dfd1e9a286d897948cd59676f1925f64

Observation 6bb10234-dd0d-4a45-b7b6-9b94bdb865e0 · outbound

This paper cites Distribution Matching Distillation Meets Reinforcement Learning.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer Distribution Matching Distillation Meets Reinforcement Learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:32.038741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:32.038741Z digest=sha256:38851e0409859f6a7cd237bdfa16c4c37d63acb8a02dac74915e5cf004ecda50

Observation d6789c7b-cc47-400e-abb4-3762a195fc9c · outbound

This paper cites Analyzing and improving the training dynamics of diffusion models.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer Analyzing and improving the training dynamics of diffusion models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:32.129517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:32.129517Z digest=sha256:3906a7efe559298f741ee80cea7431b831140d9ba16b950a00fc2b9484369a70

Observation 3822905b-a2fb-455a-bf28-33e0d8a27e68 · outbound

This paper cites Kolors 2.0.https://app.klingai.com/cn/, 2025.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer Kolors 2.0.https://app.klingai.com/cn/, 2025

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:32.243783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:32.243783Z digest=sha256:9b4d6c665e54f37a89319abd8fc7432c8c28a021660fa25f7a891ad6a329eebc

Observation 86b1d17d-62e2-48b1-aa7a-25a3f2fb1522 · outbound

This paper cites Flux.https://github.com/black-forest-labs/flux, 2023.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer Flux.https://github.com/black-forest-labs/flux, 2023

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:32.330782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:32.330782Z digest=sha256:12f1d63aca9aa35092884decdf10dac882ed45c01a727efed0d466545d9f89ab

Observation 8dd85b46-19dc-4cd6-978d-41d6d0daa65f · outbound

This paper cites FLUX.2: State-of-the-Art Visual Intelligence.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer FLUX.2: State-of-the-Art Visual Intelligence

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:32.382835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:32.382835Z digest=sha256:818b29b772b65970a706787a7521a6178b094a8eb920e7654d8977820e432ff3

Observation 8757f701-eafa-4306-b768-20a719972bba · outbound

This paper cites FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:32.488207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:32.488207Z digest=sha256:87050ef6a2d2fb70abdd2b3f32cba9bc4cdc2615b99a09c47367710e7f3bdffd

Observation 06d3c611-4751-4bb0-a00d-c84263c49369 · outbound

This paper cites Gpu server rental pricing.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer Gpu server rental pricing

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:32.631497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:32.631497Z digest=sha256:f098dc4ee830862924200402e4163cfa30f03eb49bfa9a8d34301028ab141898

Observation e406c874-4fcf-416f-a346-fe3c80f85494 · outbound

This paper cites Playground v2.5: Three Insights towards Enhancing Aesthetic Quality in Text-to-Image Generation.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer Playground v2.5: Three Insights towards Enhancing Aesthetic Quality in Text-to-Image Generation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:32.718177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:32.718177Z digest=sha256:8c49358c993140efae74581181efbf442424fccbbe89fd95334a7f1de58239ee

Observation 9562f21d-da7d-45f2-8a5f-e8bad6940781 · outbound

This paper cites Ragdiffusion: Faithful cloth generation via external knowledge assimilation.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer Ragdiffusion: Faithful cloth generation via external knowledge assimilation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:32.746065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:32.746065Z digest=sha256:0ce456e5684633648f7e00229dd5d8dbfdf196aa3005dbe1e623ad40bed6ccad

Observation 31537b14-4724-42bc-9266-d28b5e4755c2 · outbound

This paper cites Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:32.749977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:32.749977Z digest=sha256:4226694f87daaaf919f3b05fdd2a638ce3f84c79c82539a58b6d36f8040664dd

Observation 32545311-bce1-4a11-846e-f03ab50aa75c · outbound

This paper cites Visualcloze: A universal image generation framework via visual in-context learning.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer Visualcloze: A universal image generation framework via visual in-context learning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:32.754194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:32.754194Z digest=sha256:972f70da92e61a56a9f26ff7ddbc94a6c486dfd85a10218d84e78882ed88f098

Observation ad822751-6882-4b7c-86f5-4f244de43800 · outbound

This paper cites Uniworld-V2: Reinforce Image Editing with Diffusion Negative-aware Finetuning and MLLM Implicit Feedback.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer Uniworld-V2: Reinforce Image Editing with Diffusion Negative-aware Finetuning and MLLM Implicit Feedback

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:32.758069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:32.758069Z digest=sha256:58121c5c25de9e515b3e51e0dee906df5ca09289ded21b838f2cdb53d7eb41da

Observation 52956ea3-9be0-4a8d-b1d4-e64d0ac1a586 · outbound

This paper cites UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:32.762169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:32.762169Z digest=sha256:6a12a880f3b13f5f6941138932139710af8d4725b4cb0c793af986cc4060f02e

Observation dd5b2ed7-9bf0-49a2-989d-8f9df31cceb9 · outbound

This paper cites Flow Matching for Generative Modeling.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer Flow Matching for Generative Modeling

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:32.766217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:32.766217Z digest=sha256:5b29d26d3b19e21cbf0bb8ecd1f1144e12fd4f44755024a3cb3f34e7135cb7cc

Observation de4402de-2521-403c-ade7-51d505165ab5 · outbound

This paper cites Decoupled dmd: Cfg augmentation as the spear, distribu- tion matching as the shield.arXiv preprint, 2025.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer Decoupled dmd: Cfg augmentation as the spear, distribu- tion matching as the shield.arXiv preprint, 2025

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:32.770481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:32.770481Z digest=sha256:63cc92c643889acccb4065a343392e32503e8520dc0b636e0894ea9bd95bc7aa

Observation 8df3ea5e-949a-4a5d-b433-07aee08b746c · outbound

This paper cites Flow-GRPO: Training Flow Matching Models via Online RL.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer Flow-GRPO: Training Flow Matching Models via Online RL

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:32.774192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:32.774192Z digest=sha256:6fcc81ffa3d85a6383aeece55ca2f5240975e7889ef6b7a29b09bc8b017f0b77

Observation 06e0ce66-b787-4070-8781-989d799cca07 · outbound

This paper cites Step1X-Edit: A Practical Framework for General Image Editing.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer Step1X-Edit: A Practical Framework for General Image Editing

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:32.778169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:32.778169Z digest=sha256:8bb25a05dce965289822e38aa17369693c821c34f61f22105af229392929e1d8

Observation 43f98f48-9d2c-4168-837b-19f2d91210be · outbound

This paper cites Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:32.782064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:32.782064Z digest=sha256:3ff32362eaecaa1b912b3e2871fb17fb557b39dd81858b34e24c4670fcc88697

Observation c21bfeb1-2a69-48f2-8e4d-de3b6d031cfa · outbound

This paper cites OmniCaptioner: One Captioner to Rule Them All.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer OmniCaptioner: One Captioner to Rule Them All

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:32.786122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:32.786122Z digest=sha256:645fe40dd65743c00804143727275b84f124fb1f590bf1b8e5b7332d3752861a

Observation bb5e7e7b-9b35-40bb-a5c1-b13a311833ac · outbound

This paper cites Cosine normalization: Using cosine similarity instead of dot product in neural networks.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer Cosine normalization: Using cosine similarity instead of dot product in neural networks

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:32.789992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:32.789992Z digest=sha256:fb7f5abed0bbc4c222cbdba967180ff5dd2a3c1d7d54ccd1bcc3e759b2768b2a

Observation af18e0f6-a6c1-4d71-b93a-a757630e669e · outbound

This paper cites JanusFlow: Harmonizing Autoregression and Rectified Flow for Unified Multimodal Understanding and Generation.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer JanusFlow: Harmonizing Autoregression and Rectified Flow for Unified Multimodal Understanding and Generation

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:32.793928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:32.793928Z digest=sha256:127351c5bc91fbc86df118875e0f29c468bc35bbe562e4a3f763f54fdb025fc2

Observation 1112a812-c971-43fb-9035-4022c6ef2a02 · outbound

This paper cites Midjourney v7.https://www.midjourney.com/home, 2025.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer Midjourney v7.https://www.midjourney.com/home, 2025

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:32.797901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:32.797901Z digest=sha256:c26ad67d6e844f43f6a1daadbf0cc4629a9a48e1f8219834f6f292d3db08cc36

Observation b21d78b7-2663-4049-86cc-3971a91ca126 · outbound

This paper cites Enhancing few-shot image classification with cosine transformer.IEEE Access, 11:79659–79672, 2023.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer Enhancing few-shot image classification with cosine transformer.IEEE Access, 11:79659–79672, 2023

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:32.801626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:32.801626Z digest=sha256:1a37bce27440cb5c7f381b7dea9f62cf035c72fd6097ca1851f9280a0ab9d8c3

Observation 402a0527-8c22-46cc-8e70-2bedd5770fe9 · outbound

This paper cites Cagra: Highly parallel graph construction and approximate nearest neighbor search for gpus.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer Cagra: Highly parallel graph construction and approximate nearest neighbor search for gpus

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:32.805453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:32.805453Z digest=sha256:b6296e64226ce7c5a229de71bdda4542760f5cd0ec5ec0f60d8b86254b2ae7b4

Observation 4d93624c-7506-4252-b846-08656cca8f73 · outbound

This paper cites Gpt-image-1.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer Gpt-image-1

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:32.808944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:32.808944Z digest=sha256:1484f18e91f9286f192f0979deadaa5ab50ee365fbee06f603d085aba22f4d37

Observation f05a983a-cde9-48d8-a1e4-6b712b5bb37e · outbound

This paper cites The pagerank citation ranking: Bringing order to the web.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer The pagerank citation ranking: Bringing order to the web

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:32.812444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:32.812444Z digest=sha256:d65ee3c9460712d679dd033b95aa73ce992d30c02ca8fd500a1edf427bc6ab24

Observation 80639104-f252-4422-b3fc-54545848c8c7 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:32.816279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:32.816279Z digest=sha256:1b4953486610be1c9952f34ba65971ecf1002ab2a208e4e7841cc9133e2e9e29

Observation 083082b4-0ea8-4d8a-b712-96b7160d0bd5 · outbound

This paper cites Lumina-Image 2.0: A Unified and Efficient Image Generative Framework.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer Lumina-Image 2.0: A Unified and Efficient Image Generative Framework

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:32.820285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:32.820285Z digest=sha256:1cf73d717657cda9e7ef695f7d4416014c2766f2d3d2912bcc8c16f3402732e4

Observation 9badaaa4-7339-4a94-9e1b-475e03d93446 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.Advances in neural information processing systems, 36:53728–53741, 2023.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer Direct preference optimization: Your language model is secretly a reward model.Advances in neural information processing systems, 36:53728–53741, 2023

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:32.823876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:32.823876Z digest=sha256:8253e3989b30a2e45021b8c029b68deb0e719608d5789290b95ed0b8a955cc26

Observation add8b61f-7163-4506-9a3e-a5cf06f6eeb6 · outbound

This paper cites cuGraph - RAPIDS Graph Analytics Library.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer cuGraph - RAPIDS Graph Analytics Library

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:32.827296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:32.827296Z digest=sha256:92e2f55ab158640416aaf3f2af178d9529f8cc6ac0c29edfea1bf59de637359b

Observation 0fd77d68-dfdb-47b6-8f9d-b4bc66c38e4d · outbound

This paper cites Recraft v3.https://www.recraft.ai/docs/recraft-models/recraft-V3, 2024.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer Recraft v3.https://www.recraft.ai/docs/recraft-models/recraft-V3, 2024

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:32.830910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:32.830910Z digest=sha256:49afc566da20369184f1ecc87c3db4ee6d0577d92a42165f72a8915160f85401

Observation 7c9581a8-9a96-42f2-9379-6db1f75516cb · outbound

This paper cites The probabilistic relevance framework: Bm25 and beyond.Foundations and Trends® in Information Retrieval, 3(4):333–389, 2009.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer The probabilistic relevance framework: Bm25 and beyond.Foundations and Trends® in Information Retrieval, 3(4):333–389, 2009

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:32.834387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:32.834387Z digest=sha256:2092509c69d914a53cb839b3f279aa9d5ba546a57ae54394a4b34b9a94397872

Observation de6e366d-b1cb-4d6d-b8e0-6bbf2d0ca9d4 · outbound

This paper cites High- resolution image synthesis with latent diffusion models.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer High- resolution image synthesis with latent diffusion models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:32.838076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:32.838076Z digest=sha256:df068c620748ef8e1f4190b341cecdf1c98cc9c20a39286f448623e6917855cb

Observation 7fd5d0e7-076e-4197-8209-27b54b88993e · outbound

This paper cites Seedream 4.0: Toward Next-generation Multimodal Image Generation.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer Seedream 4.0: Toward Next-generation Multimodal Image Generation

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:32.841640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:32.841640Z digest=sha256:1810a5017a7331c8d374a23c67e6ac3e534a2b26ff04984ad7b48ded640ebb7c

Observation 001f9e5e-79db-4125-bd69-62f8a157e1ba · outbound

This paper cites Flashattention-3: Fast and accurate attention with asynchrony and low-precision.Advances in Neural Information Processing Systems, 37:68658–68685, 2024.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer Flashattention-3: Fast and accurate attention with asynchrony and low-precision.Advances in Neural Information Processing Systems, 37:68658–68685, 2024

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:32.845771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:32.845771Z digest=sha256:a484c87dffe51c113ce3d071282b9ddae4633c3b2bb578cbd414c5b609a067ad

Observation 6cbba010-adce-477b-94fd-5ab51a938109 · outbound

This paper cites an unresolved cited work.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer Unresolved cited work

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:32.849321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:32.849321Z digest=sha256:b34873819860c9e0aaa50262b39c2ee26c14f3f14f1081e57a75e01a992c80fb

Observation bf6655a5-17b8-4c07-85e5-ba6c4dff28ca · outbound

This paper cites Flux.1 krea [dev].https://github.com/krea-ai/flux-krea, 2025.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer Flux.1 krea [dev].https://github.com/krea-ai/flux-krea, 2025

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:32.852837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:32.852837Z digest=sha256:0a9560b28bfc8e855547be7c143f4e2e48d0b8a44a4e99980ec2ce30f5c3624d

Observation b3f2e89b-804a-4105-81dc-a7ec64ad74c0 · outbound

This paper cites From louvain to leiden: guaranteeing well-connected communities.Scientific reports, 9(1):1–12, 2019.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer From louvain to leiden: guaranteeing well-connected communities.Scientific reports, 9(1):1–12, 2019

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:32.856326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:32.856326Z digest=sha256:c030960af6e7f7ee3df76c30138d3c773c6037930497d60b2060d2409f835914

Observation 8547a68f-3bc8-436a-b350-d3b1ebfe4763 · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:32.859900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:32.859900Z digest=sha256:f1711f058c375d17862e5a915e7530292788f8e7c0988204c937d359c356a4da

Observation 9d7a79c0-80d7-4f36-a560-e97452fa4c77 · outbound

This paper cites Anytext: Multilingual visual text generation and editing.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer Anytext: Multilingual visual text generation and editing

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:32.863821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:32.863821Z digest=sha256:61d21bc1bb0261a732cb6d2a19bcd24071dcc5f9c155e2d721d60d63084801e9

Observation 2c059e89-b1c9-4c53-9b29-e88f08179e47 · outbound

This paper cites Automatic Data Curation for Self-Supervised Learning: A Clustering-Based Approach.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer Automatic Data Curation for Self-Supervised Learning: A Clustering-Based Approach

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:32.867409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:32.867409Z digest=sha256:c4bfc7806ea87f3f9880e7e440a0b4c129b661b59ace7ea9773c473d6fd2694a

Observation ae2eea76-621f-4287-94a7-8e76b1acd82e · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer Emu3: Next-Token Prediction is All You Need

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:32.871256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:32.871256Z digest=sha256:e5546acf4627efaf1504c5dc3049f0253fe473f7f4e99706b48ec17faeff7f25

Observation 6fb69a08-9e45-47f3-a5e9-b4272ca191fb · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.Advances in neural information processing systems, 35:24824–24837, 2022.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer Chain-of-thought prompting elicits reasoning in large language models.Advances in neural information processing systems, 35:24824–24837, 2022

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:32.875546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:32.875546Z digest=sha256:1f162752a05b81023bf5dc8dd2f264ad03ff8cd99642780a19dea4bea0c1e61d

Observation c225f343-de65-4a46-a361-411e08e52309 · outbound

This paper cites TIIF-Bench: How Does Your T2I Model Follow Your Instructions?.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer TIIF-Bench: How Does Your T2I Model Follow Your Instructions?

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:32.879231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:32.879231Z digest=sha256:3fb81a38d6332b88ac2d4fa05ab8bc0ee9c07975461f388946f7de68741f983e

Observation 04c73a54-6442-49b1-b6ad-8f40308d6c62 · outbound

This paper cites Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:32.882907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:32.882907Z digest=sha256:29b33076fb8a65e04b26eebdfac2f90240050853b5f9dc98e33c1c9f14663d04

Observation f6cbc9a7-7c97-4211-9433-48b1576430ea · outbound

This paper cites Qwen-Image Technical Report.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer Qwen-Image Technical Report

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:32.886603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:32.886603Z digest=sha256:3d86fef549974906d8a2765fbc842b6a9ff61df1b12fcb974ded34584d406c11

Observation 591e6a73-4ffb-41ff-97be-97f818aa624b · outbound

This paper cites Janus: Decoupling visual encoding for unified multimodal understanding and generation.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer Janus: Decoupling visual encoding for unified multimodal understanding and generation

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:32.890313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:32.890313Z digest=sha256:e0b96c4bc30ff1ca26c9cffa123252847af012382c92d89b66bd4517a6cdbcb7

Observation 51eaf0c8-952c-4956-a7cf-6ffc95dbe6bf · outbound

This paper cites OmniGen2: Towards Instruction-Aligned Multimodal Generation.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:32.893816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:32.893816Z digest=sha256:bf81c4cac417c205ed4cf52e4bb2fd7dc0856bf25c53800c305108e8f04ffef6

Observation ea22e517-559e-489c-a0c5-3dc9a51399a4 · outbound

This paper cites LightGen: Efficient Image Generation through Knowledge Distillation and Direct Preference Optimization.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer LightGen: Efficient Image Generation through Knowledge Distillation and Direct Preference Optimization

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:32.897508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:32.897508Z digest=sha256:620700fd79b714b54b9d2fbcc15ef8be4a134167bfd790a02a1c997a1a775161

Observation e5a2cb71-57c7-450b-94f0-6568e6b5f5e6 · outbound

This paper cites Omnigen: Unified image generation.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer Omnigen: Unified image generation

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:32.901206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:32.901206Z digest=sha256:cb7939832064eb83cbf257a5d66b593afeb70bd3f8007d241ab4741f035d238b

Observation 34af0cec-fee1-41fa-bc57-d9178a7b71b2 · outbound

This paper cites Sana 1.5: Efficient scaling of training-time and inference- time compute in linear diffusion transformer.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer Sana 1.5: Efficient scaling of training-time and inference- time compute in linear diffusion transformer

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:32.905000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:32.905000Z digest=sha256:06b5fcdcca0dd7e0c99103af475e8fb4da0bde9ca3c0982fd87f12592b345a49

Observation 414f53da-d353-4115-8a61-2d07dd8fdfba · outbound

This paper cites Show-o: One single transformer to unify multimodal understanding and generation.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer Show-o: One single transformer to unify multimodal understanding and generation

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:32.908697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:32.908697Z digest=sha256:4138652b0ed3cfd2ee01df2c952f2687e9e54dc3502e5de3908ad3060a7e1ba0

Observation 5b6c705a-04c7-4a30-b649-78ad33a04a3b · outbound

This paper cites Show-o2: Improved Native Unified Multimodal Models.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer Show-o2: Improved Native Unified Multimodal Models

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:32.912339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:32.912339Z digest=sha256:622769a5301f818f8c96ffb3ef1a006e38c2bd16a00ca6d2f40edf05b8932009

Observation 407fb0a1-7253-40a0-8e5b-c93341c6631c · outbound

This paper cites VisionReward: Fine-Grained Multi-Dimensional Human Preference Learning for Image and Video Generation.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer VisionReward: Fine-Grained Multi-Dimensional Human Preference Learning for Image and Video Generation

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:32.916121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:32.916121Z digest=sha256:9b6aa2e72b8ad0775a5348a440464464dabba6a197ffd6af8c69b32af2b53bc5

Observation 51cdbaa3-842d-4ff5-9420-88e348106128 · outbound

This paper cites Qwen3 Technical Report.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer Qwen3 Technical Report

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:32.919923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:32.919923Z digest=sha256:4a3c7461984f712ef380578bc447afe5113bf6a12542d281f9ecb2df353b541f

Observation 9313398e-7389-40e6-9162-02bc5618da1d · outbound

This paper cites Chinese CLIP: Contrastive Vision-Language Pretraining in Chinese.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer Chinese CLIP: Contrastive Vision-Language Pretraining in Chinese

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:32.923852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:32.923852Z digest=sha256:c8a6beea31ebcc4952955c151fa6f44cd25018d4f64ae0c43defaffe55ac5e37

Observation b95ce2d4-deaf-43d9-a338-bffeab7b07e9 · outbound

This paper cites ImgEdit: A Unified Image Editing Dataset and Benchmark.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer ImgEdit: A Unified Image Editing Dataset and Benchmark

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:32.927991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:32.927991Z digest=sha256:2a08f4a8828723b3eb1aee6aa1f1de951e2473973564d4e5d5dd20a524e08e03

Observation 2f78b4cf-82b7-4a6b-9b14-2ebe4c44aa91 · outbound

This paper cites Improved distribution matching distillation for fast image synthesis.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer Improved distribution matching distillation for fast image synthesis

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:32.931947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:32.931947Z digest=sha256:7c6004c9df2fa9de67e4b7bef05fbfbec45b312081391cf35577c894af860fda

Observation 19d8feed-c4f9-443c-ac5e-2e95158b2cf6 · outbound

This paper cites One-step diffusion with distribution matching distillation.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer One-step diffusion with distribution matching distillation

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:32.936030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:32.936030Z digest=sha256:9197fb7a5ae7b3a89291008195dc9e642ba7d6680ab2d698148828e69ff353f6

Observation 89dffe26-4af2-4f31-9c22-7e00364e6146 · outbound

This paper cites Anyedit: Mastering unified high-quality image editing for any idea.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer Anyedit: Mastering unified high-quality image editing for any idea

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:32.939949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:32.939949Z digest=sha256:1b70d276f86ca6ccef41353f2dc46e3f49502f1fbd0b840ea4f564ae8bfb5a8b

Observation daf94073-59ef-4700-97cb-aef9c5aa98e9 · outbound

This paper cites Root mean square layer normalization.Advances in Neural Informa- tion Processing Systems, 32, 2019.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer Root mean square layer normalization.Advances in Neural Informa- tion Processing Systems, 32, 2019

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:32.943774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:32.943774Z digest=sha256:cd09405ecada485de42c6e6986d142e55f371595879b6a290435f53979ed718b

Observation 2f39c923-2d12-4713-ba8e-37fb47a77507 · outbound

This paper cites Magicbrush: A manually annotated dataset for instruction-guided image editing.Advances in Neural Information Processing Systems, 36:31428–31449, 2023.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer Magicbrush: A manually annotated dataset for instruction-guided image editing.Advances in Neural Information Processing Systems, 36:31428–31449, 2023

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:32.947772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:32.947772Z digest=sha256:8255088456f3d39dac219f42a3510a5a5a074d9bccd222f7868dc2dadc5a5d81

Observation 724d8e93-be20-455a-9c1d-538edb4c1b6a · outbound

This paper cites Waver: Wave Your Way to Lifelike Video Generation.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer Waver: Wave Your Way to Lifelike Video Generation

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:32.951504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:32.951504Z digest=sha256:16be495ec24b622c0170b75cbc78055bba4afed927d0c1b5c0274e26a50ec963

Observation d577266e-387f-46a8-8ce8-25dad117d4df · outbound

This paper cites In-Context Edit: Enabling Instructional Image Editing with In-Context Generation in Large Scale Diffusion Transformer.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer In-Context Edit: Enabling Instructional Image Editing with In-Context Generation in Large Scale Diffusion Transformer

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:32.955470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:32.955470Z digest=sha256:8bedd656885202013c79f4693c8529fb5e013df13c59d2c43eab249d3486da8b

Observation d8c04b0f-276f-44b8-818f-4219b120d5f2 · outbound

This paper cites Ultraedit: Instruction-based fine-grained image editing at scale.Advances in Neural Information Processing Systems, 37:3058–3093, 2024.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer Ultraedit: Instruction-based fine-grained image editing at scale.Advances in Neural Information Processing Systems, 37:3058–3093, 2024

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:32.959458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:32.959458Z digest=sha256:9b8442ce7f69e868dfbb0c444ab2f981855244e0d95e2bc0c6d01ce712b0287e

Observation e5b84ebc-1b0e-4bc5-bdc0-b8b55e55a319 · outbound

This paper cites PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:32.963177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:32.963177Z digest=sha256:5096cb327bff73a14d942bb1541e98e9515ec9ec10ae8f6a50a9245131690a23

Observation 5cc1cd5d-82f7-4f75-8817-df4ba1599eb1 · outbound

This paper cites CogView3: Finer and Faster Text-to-Image Generation via Relay Diffusion.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer CogView3: Finer and Faster Text-to-Image Generation via Relay Diffusion

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:32.968338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:32.968338Z digest=sha256:4763b895365cb732be49b859db37e72005ab65db79e03dcacdcc9a8612f0d242

Observation 7f3e05e9-da72-47ae-a9d2-eae80e7dac7b · outbound

This paper cites 3DIS-FLUX: simple and efficient multi-instance generation with DiT rendering.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer 3DIS-FLUX: simple and efficient multi-instance generation with DiT rendering

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:32.977475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:32.977475Z digest=sha256:0c92d551942734fd10e7517194c43e0a1e710f77c986efca07994842bb7aa136

Observation a9d8bf48-453a-46bb-adc0-368c4766b394 · outbound

This paper cites Lumina-next: Making lumina-t2x stronger and faster with next-dit.Advances in Neural Information Processing Systems (NeurIPS), 2024.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer Lumina-next: Making lumina-t2x stronger and faster with next-dit.Advances in Neural Information Processing Systems (NeurIPS), 2024

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:32.982043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:32.982043Z digest=sha256:a3ee4640a4d148150ee2048c12d13c40fea3d9156976c7ad3e0834b943ae1a2d

Observation 6a802147-c552-40de-9f64-5ad920717acd · outbound

This paper cites Turbo”字样,胸口白色小字清晰的写着“MAI.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer Turbo”字样,胸口白色小字清晰的写着“MAI

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:32.986689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:32.986689Z digest=sha256:e3208b86453a765845a9f1a2818723b77b93f398b426b0da5b11ab33b99288f8

Pith citing papers

Observation c81ea765-ef6b-477a-a082-dd27135b2b35 · inbound

Distribution Matching Distillation Meets Reinforcement Learning cites this paper.

Distribution Matching Distillation Meets Reinforcement Learning Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-03T21:47:20.091819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:47:20.091819Z digest=sha256:50eb36be043e45776dc2dd3b002f5836ff1ef1f3f349a361f6bbc4c9680340b0

Observation 9689cddf-a710-43e7-a250-f549ef12df32 · inbound

Prompt Reinjection: Alleviating Prompt Forgetting in Multimodal Diffusion Transformers for Text-to-Image Generation cites this paper.

Prompt Reinjection: Alleviating Prompt Forgetting in Multimodal Diffusion Transformers for Text-to-Image Generation Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:04:10.565559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T13:00:22.875471Z digest=sha256:9205dda7edda23a1a81f2dc5575e6956cbd67a408585ecdf5369e854de09b354

Observation 61207d22-6a08-47c7-beb3-4a670fc9c49a · inbound

Prompt Reinjection: Alleviating Prompt Forgetting in Multimodal Diffusion Transformers for Text-to-Image Generation cites this paper.

Prompt Reinjection: Alleviating Prompt Forgetting in Multimodal Diffusion Transformers for Text-to-Image Generation Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T03:50:45.140207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:50:45.140207Z digest=sha256:ebe3672e98dbf78a288e9307cdbd86d6c1491bc0300aa0e6c3067f4c5208216a

Observation e47ed3b2-0d9a-45f7-a42c-d69329f70c11 · inbound

Beyond VLM-Based Rewards: Diffusion-Native Latent Reward Modeling cites this paper.

Beyond VLM-Based Rewards: Diffusion-Native Latent Reward Modeling Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-25T06:45:26.193206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-25T06:41:33.493927Z digest=sha256:0b4011770c9c5d9d72944dc44c584f19925b4b29f30c9ee2a57dc79a7e2f0a61

Observation fd15ab86-c4cc-4ca0-a4f3-28920028d1b0 · inbound

FAIL: Flow Matching Adversarial Imitation Learning for Image Generation cites this paper.

FAIL: Flow Matching Adversarial Imitation Learning for Image Generation Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-02T23:59:13.166878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:59:13.166878Z digest=sha256:6ebc993f381f7b139a8dab0626332579c247cefa96d7a084ab0e00707bf32f5c

Observation f0c45f18-6d24-4c4f-8610-5b361a36f76d · inbound

RelaxFlow: Text-Driven Amodal 3D Generation cites this paper.

RelaxFlow: Text-Driven Amodal 3D Generation Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-15T14:36:31.220132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:36:31.220132Z digest=sha256:6f987ae274101ae8d350553150bb858ad31b67922b7a6f35f80a89a9929af21e

Observation 7ad30ae1-7a54-4881-9abe-9c7863b9657c · inbound

Spectral and Trajectory Regularization for Diffusion Transformer Super-Resolution cites this paper.

Spectral and Trajectory Regularization for Diffusion Transformer Super-Resolution Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 78

Resolution
unresolved
no resolver link, observed 2026-07-15T13:54:54.362524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T13:54:54.362524Z digest=sha256:c5f58e285fed60d52d5fbc8b1453b8869dee721ca9d9f86d2a180b6b797a5f40

Observation 3f06452c-899d-4cbe-8af5-7c5b96768746 · inbound

Na\"ive PAINE: Lightweight Text-to-Image Generation Improvement with Prompt Evaluation cites this paper.

Na\"ive PAINE: Lightweight Text-to-Image Generation Improvement with Prompt Evaluation Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-14T22:28:35.981853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T22:28:35.981853Z digest=sha256:74e9556b5c68fd527c7072e915febcd5c3bad247d9fc30e268e6c313dc05ba6a

Observation f7f73945-cecc-42ca-ae2d-8aa94eb462d6 · inbound

Gen-Searcher: Reinforcing Agentic Search for Image Generation cites this paper.

Gen-Searcher: Reinforcing Agentic Search for Image Generation Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-14T21:19:28.174364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T21:18:10.087258Z digest=sha256:89574ce07664d77b464251166f99b88c6f66a3f3c59aa6bf235e83f00be29997

Observation 5dabeb3f-6b33-458e-a361-9663a75f3a2d · inbound

Gen-Searcher: Reinforcing Agentic Search for Image Generation cites this paper.

Gen-Searcher: Reinforcing Agentic Search for Image Generation Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-25T06:35:25.098083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-25T06:31:13.232880Z digest=sha256:cabafacb256c5712af82db15fabe2a05ec1f6d2f5e0c64359e9ccc673bccbead

Observation 8b77b497-933e-4086-9e0d-f277a852add0 · inbound

The Darkside-20k Data Acquisition System cites this paper.

The Darkside-20k Data Acquisition System Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-15T11:41:34.003655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T11:41:34.003655Z digest=sha256:301b3eda5a728710951d4147e9dc03897c5ad7f6c11ce9f590dc4742a24001ff

Observation 3f2f45c2-1057-43f1-be25-518e96e59c0a · inbound

Can Nano Banana 2 Replace Traditional Image Restoration Models? An Evaluation of Its Performance on Image Restoration Tasks cites this paper.

Can Nano Banana 2 Replace Traditional Image Restoration Models? An Evaluation of Its Performance on Image Restoration Tasks Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-13T20:13:13.384687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T20:11:44.590571Z digest=sha256:6643afb83ce02d5089857d4ee4f16bd97677d3f5b7bd8b138b783af3cebb6437

Observation f996419c-c92a-4bdb-8b41-27907867f2e5 · inbound

Beyond Few-Step Inference: Accelerating Video Diffusion Transformer Model Serving with Inter-Request Caching Reuse cites this paper.

Beyond Few-Step Inference: Accelerating Video Diffusion Transformer Model Serving with Inter-Request Caching Reuse Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T14:08:37.771343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:50:20.499116Z digest=sha256:c3388f91b037ab7dda9ee3b89ba00c77d50d70a06e9f6a6c4b3fb1c9a3766e69

Observation 1a7e33da-a28a-4c72-b3fe-9c75e168f906 · inbound

Beyond Standard Benchmarks: A Systematic Audit of Vision-Language Model's Robustness to Natural Semantic Variation Across Diverse Tasks cites this paper.

Beyond Standard Benchmarks: A Systematic Audit of Vision-Language Model's Robustness to Natural Semantic Variation Across Diverse Tasks Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T14:08:37.771343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T20:26:42.128350Z digest=sha256:66c8377b229bebacb30a12321234333546ce138ad4db4f085f178efba7141a76

Observation d75b1bd7-cfe4-4999-99be-cd9f949c808d · inbound

SurFITR: A Dataset for Surveillance Image Forgery Detection and Localisation cites this paper.

SurFITR: A Dataset for Surveillance Image Forgery Detection and Localisation Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:08:37.771343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:23:48.168061Z digest=sha256:e5c633744413ab3c1ba6eeb9c4d4f817afa6fbc3d75ff6004fd8af267e6bd511

Observation e4c7aac1-ebda-46a6-a70f-503b84790f99 · inbound

FlowGuard: Towards Lightweight In-Generation Safety Detection for Diffusion Models via Linear Latent Decoding cites this paper.

FlowGuard: Towards Lightweight In-Generation Safety Detection for Diffusion Models via Linear Latent Decoding Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:08:37.771343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T17:03:35.420451Z digest=sha256:3f4a9ab017cb509ad3ddd50a44f00444162967b7002f01c4165cd1eb03ca2f93

Observation 48fdee11-6ed5-4f2b-b553-24007ea23e7a · inbound

On Semiotic-Grounded Interpretive Evaluation of Generative Art cites this paper.

On Semiotic-Grounded Interpretive Evaluation of Generative Art Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:08:37.771343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:35:53.179645Z digest=sha256:7a726b2e36cdc5215391bcf9ed3d914198f8ef4725bbc249159eb9a6a4315521

Observation 7e7cb078-fdec-46b6-ad90-1fd8751cb2ff · inbound

Large-Scale Universal Defect Generation: Foundation Models and Datasets cites this paper.

Large-Scale Universal Defect Generation: Foundation Models and Datasets Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:08:37.771343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T17:49:58.002432Z digest=sha256:dd57bf2e30c9213a3c1eb71f1ccb39a649bfaab5f906746df0f5e9e3baa27e85

Observation 92aac240-6ecb-478c-9dad-be700d893396 · inbound

CT-1: Vision-Language-Camera Models Transfer Spatial Reasoning Knowledge to Camera-Controllable Video Generation cites this paper.

CT-1: Vision-Language-Camera Models Transfer Spatial Reasoning Knowledge to Camera-Controllable Video Generation Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:08:37.771343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T16:53:54.166457Z digest=sha256:d455338f438d9f9848413b44c1601b06b11f50e6d0b83d901bc50935ee715732

Observation 099445d6-f0a5-4aad-b13c-c522b680c04d · inbound

The Second Challenge on Real-World Face Restoration at NTIRE 2026: Methods and Results cites this paper.

The Second Challenge on Real-World Face Restoration at NTIRE 2026: Methods and Results Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:08:37.771343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T16:14:44.509619Z digest=sha256:fa7ba8747a9d260322d032eb0a88c931179a9c3045636a02720587130a6a70a0

Observation d5935658-43c1-4359-a2b0-689c357093db · inbound

Continuous Adversarial Flow Models cites this paper.

Continuous Adversarial Flow Models Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T14:08:37.771343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:25:53.420119Z digest=sha256:b67dd5948073f7a73e1b627c610022e5f3d9337e3a95505f26a9b008ec0ca8fc

Observation b5ff860e-bbb6-4c15-a1fd-bd4fb9d41786 · inbound

Generative Refinement Networks for Visual Synthesis cites this paper.

Generative Refinement Networks for Visual Synthesis Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:08:37.771343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:41:23.057949Z digest=sha256:b364f049f0fa06fc79d48e49d57ad83c4e02e9f657173847b79942b079d1bd16

Observation b10757ed-f2b8-4fb0-884c-5e5a2b9ce54b · inbound

Generative Refinement Networks for Visual Synthesis cites this paper.

Generative Refinement Networks for Visual Synthesis Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-12T21:00:35.123495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T21:00:35.123495Z digest=sha256:8f06c5e33025240e404739df4f16138b221516c5771e34245aceb7d8bfca6aa8

Observation f9f8c86e-05db-46ce-a1dc-b0859e697aa2 · inbound

Any3DAvatar: Fast and High-Quality Full-Head 3D Avatar Reconstruction from Single Portrait Image cites this paper.

Any3DAvatar: Fast and High-Quality Full-Head 3D Avatar Reconstruction from Single Portrait Image Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:08:37.771343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T14:21:36.209440Z digest=sha256:ebcfdee5f3e95189950f90c8eba62d2987c7b46a7d39dcf6eba0c47b7513e767

Observation 77041e8f-db5e-49b2-9d15-3f1d466ba68f · inbound

The Fourth Challenge on Image Super-Resolution ($\times$4) at NTIRE 2026: Benchmark Results and Method Overview cites this paper.

The Fourth Challenge on Image Super-Resolution ($\times$4) at NTIRE 2026: Benchmark Results and Method Overview Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:08:37.771343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T11:41:11.253990Z digest=sha256:234368802f277f718ad99600ef973b53ffceba293864835c9911933af607e882

Observation b0892c14-ad66-425e-a09e-f4a8187004c6 · inbound

Guiding Distribution Matching Distillation with Gradient-Based Reinforcement Learning cites this paper.

Guiding Distribution Matching Distillation with Gradient-Based Reinforcement Learning Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T14:08:37.771343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T02:45:35.600729Z digest=sha256:56ed873b8c20d932146b5e7b7f6da1cc2e8238747532f6cb1ff8ed244f82a438

Observation 4ff27e17-cf46-4aa6-a733-f62ccdf10847 · inbound

Generative Texture Filtering cites this paper.

Generative Texture Filtering Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T14:08:37.771343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-10T03:16:37.509163Z digest=sha256:f5119c6f3c65fea6c91725e444b6e13cad796bb170ef9a017c3a8f2819eec72d

Observation dbf6b477-aadd-4885-a492-4d722f1dbd7e · inbound

CreatiParser: Generative Image Parsing of Raster Graphic Designs into Editable Layers cites this paper.

CreatiParser: Generative Image Parsing of Raster Graphic Designs into Editable Layers Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:08:37.771343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T03:19:10.670279Z digest=sha256:94d6b2cdc44770d5d4d73c47068d67aa8ea836c0453d083ad5924b2d8ab214ed

Observation 5428014a-1fe5-4524-bdc6-b1ef0aff91bd · inbound

CoInteract: Physically-Consistent Human-Object Interaction Video Synthesis via Spatially-Structured Co-Generation cites this paper.

CoInteract: Physically-Consistent Human-Object Interaction Video Synthesis via Spatially-Structured Co-Generation Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T14:08:37.771343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T03:14:45.834520Z digest=sha256:ec2fb63812a9cc267691add8335068c539407bda36342361ef41f7a8d90ece3c

Observation dd4be4e1-ee35-4e95-a56f-c341d9bdf020 · inbound

LLaDA2.0-Uni: Unifying Multimodal Understanding and Generation with Diffusion Large Language Model cites this paper.

LLaDA2.0-Uni: Unifying Multimodal Understanding and Generation with Diffusion Large Language Model Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:08:37.771343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T00:49:38.156237Z digest=sha256:707b35f47b5b0988a4ec37b945dde0419f042ad8ebf509aefda0174bf0251bea

Observation 27fca583-69d2-4df7-852a-6acb7e861bb8 · inbound

Context Unrolling in Omni Models cites this paper.

Context Unrolling in Omni Models Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-11T14:21:04.931128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-09T22:02:57.841111Z digest=sha256:7a54a5ecbae46dfa30acec4cbd073dd5f54a873d2da9e871f4e6891233f0552e

Observation e1a80a04-44b1-4eb5-a0b4-8f57ef3288c8 · inbound

Evaluating Remote Sensing Image Captions Beyond Metric Biases cites this paper.

Evaluating Remote Sensing Image Captions Beyond Metric Biases Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:08:37.771343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T01:06:35.604862Z digest=sha256:6b29747cd9c27ff2994a2f9b35d0df9605347f919a36ce7b79a3a7b75f73b6cb

Observation 3c9c3ed0-f207-4226-b77f-0dc902f4d138 · inbound

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation cites this paper.

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-11T21:41:18.866479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T04:31:26.325118Z digest=sha256:8e8f9428b74954dae8aafba15eaf3ec1bbe3e206c38a0974dae07d95740fc5d7

Observation a95c9959-6841-4ac8-b32b-bcfbbdf71840 · inbound

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation cites this paper.

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-20T23:43:51.133147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T23:41:25.275207Z digest=sha256:aef225bd040c9aa425d1bfd76eb1e769d7f85375a66b96138e3edf654264973e

Observation db9e4650-eadd-442f-80c2-32e8716eb0b9 · inbound

Exploring Time Conditioning in Diffusion Generative Models from Disjoint Noisy Data Manifolds cites this paper.

Exploring Time Conditioning in Diffusion Generative Models from Disjoint Noisy Data Manifolds Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-05-11T23:31:15.532397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:52:29.681159Z digest=sha256:aa3e74ffc7625e6c3f258709c42fea9b968cc84ec3576e109e3015aeeedac474

Observation 447a3ff7-4aba-47f8-90f5-f57ee42ad371 · inbound

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling cites this paper.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-12T10:21:27.778015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:84093323fc260053ff75a44bcf503f8fd12bc9fb3d2c8b635ccad1733fffa344

Observation f969dd79-c707-434e-b7f4-416d4c3c0c94 · inbound

IdentiFace: Multi-Modal Iterative Diffusion Framework for Identifiable Suspect Face Generation in Crime Investigations cites this paper.

IdentiFace: Multi-Modal Iterative Diffusion Framework for Identifiable Suspect Face Generation in Crime Investigations Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-11T15:36:09.875360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-09T19:33:24.911067Z digest=sha256:7e7ed886fb60b57e0dc743948ac86b25e3469fdba7342fe37f3a5ccacccabb85

Observation f7fb7cef-3eb9-4c96-b3cc-8f1de6789977 · inbound

DirectEdit: Step-Level Accurate Inversion for Flow-Based Image Editing cites this paper.

DirectEdit: Step-Level Accurate Inversion for Flow-Based Image Editing Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T14:08:37.771343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-08T19:02:52.090839Z digest=sha256:4f2afaea25766cc87a5fee5abb27f4d4e8c49c555781ae7781afe0b473d29fa0

Observation 3839a43e-4ad8-4af3-afde-3e5328c28376 · inbound

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation cites this paper.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:08:37.771343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T17:57:08.606559Z digest=sha256:ea80c4c43731faf6064cca6a11b99080f291cf9013668535d18b5ce7ea6345f4

Observation 3a047d48-282c-4393-929b-712f74b3d7de · inbound

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation cites this paper.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 76

Resolution
verified exact
local_arxiv, observed 2026-05-21T08:19:52.748945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:0821ad3a3a94c39458ce9fa65d87b2ba436a2bb7d5beef4389c4229cf3669df0

Observation b9e6f6f5-29df-494c-b074-86a36ec0e97a · inbound

LIVEditor-14B: Lightning Unified Video Editing via In-Context Sparse Attention cites this paper.

LIVEditor-14B: Lightning Unified Video Editing via In-Context Sparse Attention Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-01T00:25:09.309077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-07-01T00:24:30.573122Z digest=sha256:46f1abf7594ec5bb29d46990d8f3fb3fa496f8541e8740d0398c45452d44ccb7

Observation 38cc7c59-5d00-4bf4-a5d2-3eb91df2f4b2 · inbound

D-OPSD: On-Policy Self-Distillation for Continuously Tuning Step-Distilled Diffusion Models cites this paper.

D-OPSD: On-Policy Self-Distillation for Continuously Tuning Step-Distilled Diffusion Models Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 93

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T17:36:05.880328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T17:25:26.391582Z digest=sha256:06a6a07cabcd6cc0324f988852916aefff4ff13db47be4d5602d4aeddbb3e8db

Observation 3c512b63-7543-4e52-958a-3235cd013849 · inbound

D-OPSD: On-Policy Self-Distillation for Continuously Tuning Step-Distilled Diffusion Models cites this paper.

D-OPSD: On-Policy Self-Distillation for Continuously Tuning Step-Distilled Diffusion Models Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 94

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T23:19:13.387338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T23:18:35.390642Z digest=sha256:608201206f391532350b9b538c65c000ce41a2c90f273ab8c984c5845fb5e85d

Observation a8b77ce6-e2a8-43b4-8300-545bc9362a7e · inbound

D-OPSD: On-Policy Self-Distillation for Continuously Tuning Step-Distilled Diffusion Models cites this paper.

D-OPSD: On-Policy Self-Distillation for Continuously Tuning Step-Distilled Diffusion Models Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 94

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T23:45:07.774888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T23:44:10.302520Z digest=sha256:6885cad3240cdfc481c02b40200a20df860734abcd45fd5af4d1ed393eb07da9

Observation 55e44ee2-ca52-4881-943e-9a1c8d803520 · inbound

DynT2I-Eval: A Dynamic Evaluation Framework for Text-to-Image Models cites this paper.

DynT2I-Eval: A Dynamic Evaluation Framework for Text-to-Image Models Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-11T18:46:10.551454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T13:54:00.141439Z digest=sha256:6f4c46644ed9a616373765e6e838872d109f2ffddea27d18a18de777eeee9632

Observation c583c316-ad95-453d-8501-413389feb0b6 · inbound

Weblica: Scalable and Reproducible Training Environments for Visual Web Agents cites this paper.

Weblica: Scalable and Reproducible Training Environments for Visual Web Agents Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:08:37.771343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:25:44.578007Z digest=sha256:7a2f26be85b05410360df2acc2de57bdb4444e131f7a0994ff9a0c8a9e897b15

Observation d0aeee1d-3b38-4a49-b64a-d2f586c94d92 · inbound

SCOPE: Structured Decomposition and Conditional Skill Orchestration for Complex Image Generation cites this paper.

SCOPE: Structured Decomposition and Conditional Skill Orchestration for Complex Image Generation Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T14:08:37.771343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-11T02:27:36.104714Z digest=sha256:c515caa81f41c1e8f2cf9e4cb808eabb171b5de1ced06f10265d42db2a8af97b

Observation 3fcf03f2-dcb1-4a58-8713-a067cb85b8dd · inbound

What Concepts Lie Within? Detecting and Suppressing Risky Content in Diffusion Transformers cites this paper.

What Concepts Lie Within? Detecting and Suppressing Risky Content in Diffusion Transformers Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-12T06:06:26.554232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T04:35:39.370969Z digest=sha256:95de8a8948bed0fba4c28558e92b3eba924f6c60c92c1ed127e1f792fda4c7f5

Observation abdb0e5c-0307-4df3-83f4-b1c34a40fc55 · inbound

Qwen-Image-2.0 Technical Report cites this paper.

Qwen-Image-2.0 Technical Report Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-12T04:21:21.908923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T04:21:18.312613Z digest=sha256:6acdddafec8f68f0a3f52bd2dc4920622091fc8027d0f0550ce3c9f97be49379

Observation e9b3cfab-41d8-4c3d-af72-011b99644dcf · inbound

HiDream-O1-Image: A Natively Unified Image Generative Foundation Model with Pixel-level Unified Transformer cites this paper.

HiDream-O1-Image: A Natively Unified Image Generative Foundation Model with Pixel-level Unified Transformer Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 48

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T07:32:29.555288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T07:30:53.939221Z digest=sha256:55105bcacc6cf7e00e685111b0badba8ac1e594cd72dce55e4ed843807c203ec

Observation 4d395256-cd06-4085-ae0e-b61278a91c60 · inbound

RevealLayer: Disentangling Hidden and Visible Layers via Occlusion-Aware Image Decomposition cites this paper.

RevealLayer: Disentangling Hidden and Visible Layers via Occlusion-Aware Image Decomposition Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 79

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T07:47:32.694082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T07:46:50.540528Z digest=sha256:a62a4304976976d51ed660220f27005cd38d788e932ad2781b1270caa20d8628

Observation cb153ed3-bed7-4ca1-ac32-331d3ef8d85a · inbound

L2P: Unlocking Latent Potential for Pixel Generation cites this paper.

L2P: Unlocking Latent Potential for Pixel Generation Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-13T07:17:29.693414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T07:12:28.181595Z digest=sha256:aca721bf6bbb611504f4443af4bc45af0e82bef2c44e37247bc4e6fd713a993f

Observation 5dbc52be-3602-4f6d-9dd4-6e6d071ff3e5 · inbound

UniCustom: Unified Visual Conditioning for Multi-Reference Image Generation cites this paper.

UniCustom: Unified Visual Conditioning for Multi-Reference Image Generation Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-13T06:02:23.965122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T05:53:21.851578Z digest=sha256:9068701664d5ceb0409de5e8c040d097c7e3dfb2fcdb61cf238e47be6df87143

Observation 6cde5d40-9720-49ae-ac79-82fb15a9599b · inbound

UniCustom: Unified Visual Conditioning for Multi-Reference Image Generation cites this paper.

UniCustom: Unified Visual Conditioning for Multi-Reference Image Generation Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-14T22:03:03.380931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T22:00:01.349754Z digest=sha256:5ece0334d548ec872d640038a60176ebc75ada0c84924459bde5f6c1f65d4790

Observation 1595dcbd-9599-46a1-a445-a10c6da4c2a8 · inbound

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture cites this paper.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:17:18.556563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:51a3eadbe2bfcc5565b3c0dc5f45f1dd92b5f110122360beb9661a78ab02e0f0

Observation a4c8bf21-737f-4cc2-aec3-215709b58134 · inbound

Asymmetric Flow Models cites this paper.

Asymmetric Flow Models Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 75

Resolution
verified exact
local_arxiv, observed 2026-05-14T19:29:23.980182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T19:28:21.625879Z digest=sha256:74e2596f5b98d2c8252da0b506016a9829680843aba91fcd719f2af74796069b

Observation 82e6df93-5d93-43ce-9b5f-608c6c6eef25 · inbound

Asymmetric Flow Models cites this paper.

Asymmetric Flow Models Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 75

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:15:47.403870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T22:07:44.850763Z digest=sha256:07a3ee04c58c9dece2c1c170ca589a698621dacd57530d5037eb732cb20b3542

Observation 37b7dcd6-5027-43e7-88bb-71d9018296e6 · inbound

ImageAttributionBench: How Far Are We from Generalizable Attribution? cites this paper.

ImageAttributionBench: How Far Are We from Generalizable Attribution? Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 66

Resolution
verified exact
local_arxiv, observed 2026-05-14T19:39:23.941699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T19:38:41.659261Z digest=sha256:5f974da1757209d126da6ff4eb0a59829efa6e0de347ce79679a84d07421b0ca

Observation 60678dc0-9456-4475-83a5-6107738bb82e · inbound

Early Semantic Grounding in Image Editing Models for Zero-Shot Referring Image Segmentation cites this paper.

Early Semantic Grounding in Image Editing Models for Zero-Shot Referring Image Segmentation Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-14T19:47:53.663940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T19:44:36.511647Z digest=sha256:a33822668f1e9450039a1787220b379af3283cbbcdc51899eba3acee6445cd8b

Observation b760ebc0-0715-456c-bc1b-67cdc84f6ed0 · inbound

Skill-Aligned Annotation for Reliable Evaluation in Text-to-Image Generation cites this paper.

Skill-Aligned Annotation for Reliable Evaluation in Text-to-Image Generation Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-14T20:19:28.523947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T20:13:21.750639Z digest=sha256:89956fc328399925ec08bfd92388a9a1c5aee75571fa6e8cf49fab076a76202d

Observation 23a48b4a-a355-48e2-9816-b1a4461148fc · inbound

MiVE: Multiscale Vision-language features for reference-guided video Editing cites this paper.

MiVE: Multiscale Vision-language features for reference-guided video Editing Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-15T05:25:03.107055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T05:23:20.179387Z digest=sha256:8b5ca2fe2a36a78b7ed69b7e4ec91c86f948ff916cf993f49080ec91cde210f5

Observation aa9d0521-cb63-4926-b5ca-dd15af61dabe · inbound

MiVE: Multiscale Vision-language features for reference-guided video Editing cites this paper.

MiVE: Multiscale Vision-language features for reference-guided video Editing Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:25:46.287958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T21:38:25.600916Z digest=sha256:fad1966ff5f8c08ab616deb1a227efa1c8464c0c8b3f51f0c0ff58b47570b157

Observation a7a59463-21f3-462f-b3e0-321f4ebdceab · inbound

Unlocking Complex Visual Generation via Closed-Loop Verified Reasoning cites this paper.

Unlocking Complex Visual Generation via Closed-Loop Verified Reasoning Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-05-19T16:37:39.687982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T16:35:43.166697Z digest=sha256:471a39cc0605936b7a6f19a3919d47a1c6b8abbfa2e5908a3cd1410f0be3be58

Observation fa1d10c9-266c-4d35-9f88-e295a47ef820 · inbound

ChangeFlow -- Latent Rectified Flow for Change Detection in Remote Sensing cites this paper.

ChangeFlow -- Latent Rectified Flow for Change Detection in Remote Sensing Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T15:52:37.994981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T15:52:00.691858Z digest=sha256:f0c7d198856f577abcfeabba14734cc436194d1fbb082762311411e822cf9f8b

Observation be3b8e36-4603-411f-ba70-facfd708f8e4 · inbound

ChangeFlow -- Latent Rectified Flow for Change Detection in Remote Sensing cites this paper.

ChangeFlow -- Latent Rectified Flow for Change Detection in Remote Sensing Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T14:00:31.040146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:00:31.040146Z digest=sha256:a86dca296131d3ac0264c5bb017a7491a3228e42f6cb73680c883c296135e0bf

Observation de55417a-ad8e-4275-9140-d6e2c400a9d6 · inbound

Generative 3D Gaussians with Learned Density Control cites this paper.

Generative 3D Gaussians with Learned Density Control Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 74

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T22:33:48.274173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-20T22:32:45.251599Z digest=sha256:4dc3708c745eff620a07505658cec69139d93f52643524c354bfabf622bd13ef

Observation f460ee5a-99ca-4f56-8772-4d02b46ad0d8 · inbound

Sketch Then Paint: Hierarchical Reinforcement Learning for Diffusion Multi-Modal Large Language Models cites this paper.

Sketch Then Paint: Hierarchical Reinforcement Learning for Diffusion Multi-Modal Large Language Models Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-19T21:22:48.589217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T21:18:03.005508Z digest=sha256:bfea032b384d7e1bf4e5d2af9aa488ae33f938a5ec02631278de5d5382287982

Observation 45151883-4db4-4572-b394-baf24449adb6 · inbound

WinTok: A Win-Win Hybrid Tokenizer via Decomposing Visual Understanding and Generation with Transferable Tokens cites this paper.

WinTok: A Win-Win Hybrid Tokenizer via Decomposing Visual Understanding and Generation with Transferable Tokens Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T12:08:15.771619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T12:04:19.761430Z digest=sha256:41e64974cc89b2c5d44e1dbfb4a53bdb3603cb7b4ffd86d5a62654523f27e6eb

Observation 316a5af7-dff3-435d-a625-0c6a031864de · inbound

Aurora: Unified Video Editing with a Tool-Using Agent cites this paper.

Aurora: Unified Video Editing with a Tool-Using Agent Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-20T10:48:12.763848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T10:47:05.038308Z digest=sha256:5cc814ec44c76e6d01ba084fd32b609ced66b8fbe50ec62a89e79a1b2d38e712

Observation dbe23e63-4e81-44f5-8cfc-af82c5c626aa · inbound

TextAlign: Preference Alignment for Text Rendering with Hierarchical Rewards cites this paper.

TextAlign: Preference Alignment for Text Rendering with Hierarchical Rewards Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-05-20T07:03:06.383496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T06:59:27.578911Z digest=sha256:07edd00869837973e2f41fdf623cb9ea7e05d2a613375b348e92e8c6a6bd487d

Observation f92daffe-b3ff-42ef-9028-b32cebb42f6d · inbound

TextAlign: Preference Alignment for Text Rendering with Hierarchical Rewards cites this paper.

TextAlign: Preference Alignment for Text Rendering with Hierarchical Rewards Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-06-30T18:55:00.426710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T18:51:00.698045Z digest=sha256:e6893b6e45164dc10adad576c6696b640d8c812eea58633d566bdc8020e83aea

Observation 37248da7-a389-43e1-af83-b006d5a92ffb · inbound

PixVerve: Advancing Native UHR Image Generation to 100MP with a Large-Scale High-Quality Dataset cites this paper.

PixVerve: Advancing Native UHR Image Generation to 100MP with a Large-Scale High-Quality Dataset Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-20T05:23:03.844651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T05:19:30.372528Z digest=sha256:86098fe067c121c02dd8469b44d022ea6115af68cc4279488bd12965af59f7a0

Observation ea07c8b7-d30c-4016-b316-ad23bfb8c304 · inbound

Rethinking Cross-Layer Information Routing in Diffusion Transformers cites this paper.

Rethinking Cross-Layer Information Routing in Diffusion Transformers Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-21T05:29:39.442913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T05:29:19.111071Z digest=sha256:b83d1383c7b97e9b5e07bf3b3509211c78c0ab0d6aecfb312b8ef3b6422483e6

Observation 2b72e306-7a7a-4c3a-8aca-594a70baad0f · inbound

Rethinking Cross-Layer Information Routing in Diffusion Transformers cites this paper.

Rethinking Cross-Layer Information Routing in Diffusion Transformers Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-06-30T17:54:58.000219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T17:50:22.258541Z digest=sha256:e9023ff34b25b85bf3c2b884ad2522e3c75e6ea0026873c294e5c093a29a2914

Observation 93fe4d9c-ca9d-4f57-81ca-9437f9c2e4e9 · inbound

TextSculptor: Training and Benchmarking Scene Text Editing cites this paper.

TextSculptor: Training and Benchmarking Scene Text Editing Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-21T05:19:39.382658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T05:16:43.756525Z digest=sha256:0402dbd6cac258656f5a4dece282ec057c050810dd51ff511fce400fbb4a4f80

Observation 7b9b421a-11ce-4d42-82fc-8e380c91cbf4 · inbound

Semantic Granularity Navigation in Image Editing cites this paper.

Semantic Granularity Navigation in Image Editing Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-21T05:53:59.253255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T05:51:19.211778Z digest=sha256:4dfd75acd0e43470a340b5a3c32489c58c20674ad50a73be87b2b6e190dbad0e

Observation f0e0c0f8-f4cb-42a5-b874-ee8466ec0beb · inbound

Semantic Granularity Navigation in Image Editing cites this paper.

Semantic Granularity Navigation in Image Editing Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-06-30T17:14:56.723749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T17:13:18.594937Z digest=sha256:3185cba82cbe677a63532ceca7a596eb1c2a20570fda7fe16aca33e1e9e6b90a

Observation 13c4cd87-0da6-4b54-b76e-0c73654c64aa · inbound

MONET: A Massive, Open, Non-redundant and Enriched Text-to-image dataset cites this paper.

MONET: A Massive, Open, Non-redundant and Enriched Text-to-image dataset Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 107

Resolution
verified exact
local_arxiv, observed 2026-05-21T05:23:58.462333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T05:21:18.369534Z digest=sha256:bea48f368483c15da78e3758330ccb4f5616e0f31ddfef25492e1c0e5eab0ae3

Observation 63013c4a-e02c-4a8b-a597-70b05bca2e5f · inbound

Lens: Rethinking Training Efficiency for Foundational Text-to-Image Models cites this paper.

Lens: Rethinking Training Efficiency for Foundational Text-to-Image Models Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-22T09:34:46.672966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T09:34:14.596976Z digest=sha256:1a5eb2ba77e14a22f0550280de1bc7454618c06fc58c69461f7027e0788d9750

Observation 0b614969-39fc-42e9-b966-8695b3e27b3a · inbound

GenEvolve: Self-Evolving Image Generation Agents via Tool-Orchestrated Visual Experience Distillation cites this paper.

GenEvolve: Self-Evolving Image Generation Agents via Tool-Orchestrated Visual Experience Distillation Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-22T09:31:22.956361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T09:29:05.360018Z digest=sha256:77158fecb342c25203c08473cacfc71905578b4da76302d0ae8fe6a37c7fdac8

Observation 4e27a481-fc03-474b-8b43-0303f01f40a4 · inbound

GenEvolve: Self-Evolving Image Generation Agents via Tool-Orchestrated Visual Experience Distillation cites this paper.

GenEvolve: Self-Evolving Image Generation Agents via Tool-Orchestrated Visual Experience Distillation Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-25T05:40:24.016810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-25T05:38:33.676797Z digest=sha256:80ac404bfe622fdfa060ac28a8545c0848c8116df8e8d74451a976e38abb2411

Observation 556454ff-12bf-4794-89d3-88c733105e57 · inbound

VDE: Training-Free Accelerating Rectified Flow Model via Velocity Decomposition and Estimation cites this paper.

VDE: Training-Free Accelerating Rectified Flow Model via Velocity Decomposition and Estimation Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-25T05:00:22.644501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-25T04:57:03.913813Z digest=sha256:431962ae8934c87fabc48a1db0765d5611dfdb6e22a69a0e68d27d60c923a22b

Observation 720c0356-bc8b-4626-bbf2-b701df67983a · inbound

PiD: Fast and High-Resolution Latent Decoding with Pixel Diffusion cites this paper.

PiD: Fast and High-Resolution Latent Decoding with Pixel Diffusion Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-25T04:20:19.320175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-25T04:18:45.403718Z digest=sha256:b3ce951623bd387e2c096806f6cecd29e3ae7438cd91839bce9820417516a58a

Observation ea21e8d1-4834-4059-b13e-3185bd86d825 · inbound

ERNIE-Image Technical Report cites this paper.

ERNIE-Image Technical Report Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-06-30T00:14:04.688925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T22:55:06.711872Z digest=sha256:56347b5588b27b45a57b001691c65e8cd3e0a00ea8d437c0d2a1c6ae672d75fd

Observation 92ade3b5-999d-4ee1-a5f4-630b0460b8c9 · inbound

Reinforcing Few-step Generators via Reward-Tilted Distribution Matching cites this paper.

Reinforcing Few-step Generators via Reward-Tilted Distribution Matching Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 69

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:24:00.295379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T22:21:06.005125Z digest=sha256:a02d06c2f4131e7fdff34e2031a9c6029a31338a05b6f9a0638e406524d5b53e

Observation 1c9a2b82-a5a7-4ba5-bdc1-99e85f94e06a · inbound

RT-Lynx: Putting the GEMM Sparsity In a Right Way for Diffusion Models cites this paper.

RT-Lynx: Putting the GEMM Sparsity In a Right Way for Diffusion Models Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 60

Resolution
verified exact
local_arxiv, observed 2026-06-29T19:43:54.684926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T19:40:42.033793Z digest=sha256:542510065acd513af526d1e74ee7fd3308a5acc36b7a83ae09f1f9728f1b30f4

Observation 8fa55d78-00bc-496d-8704-a3ffb4032cf9 · inbound

FakeVLM-R1: Internalizing Physical Laws via CoT for Synthetic Image Detection cites this paper.

FakeVLM-R1: Internalizing Physical Laws via CoT for Synthetic Image Detection Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T08:03:14.172762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T07:59:44.436264Z digest=sha256:652872390c92805b6e6d49f2193824f3187544042358f157f81737872d3d487a

Observation 6c31b516-3052-4135-b574-9cec68c14a39 · inbound

GenClaw: Code-Driven Agentic Image Generation cites this paper.

GenClaw: Code-Driven Agentic Image Generation Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-06-29T07:43:13.965775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T07:36:54.292448Z digest=sha256:0f0d240d7dea4707bfb2f1eac26070cc304d4e2d5263222ea26b771783efce87

Observation 13eec3df-aae4-4317-9945-ad92c785e957 · inbound

Representation Forcing for Bottleneck-Free Unified Multimodal Models cites this paper.

Representation Forcing for Bottleneck-Free Unified Multimodal Models Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 60

Resolution
verified exact
local_arxiv, observed 2026-07-01T19:16:00.955045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T22:54:10.460872Z digest=sha256:f7e168f208fc1ca4f7cc9ede340c00b24d9777c61ad54630c83ef1c6eba14e28

Observation 1ea2d78c-ba30-4fbc-9249-11a3c2607b33 · inbound

Representation Forcing for Bottleneck-Free Unified Multimodal Models cites this paper.

Representation Forcing for Bottleneck-Free Unified Multimodal Models Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 62

Resolution
unresolved
no resolver link, observed 2026-07-12T15:31:57.426559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T15:31:57.426559Z digest=sha256:17c43d8ba55bda2c4070f5a4cb38e1e7605949805db6be844924d0bc8ef1fe5a

Observation 6ac07a1a-1c56-4107-af36-6adacb35cd8c · inbound

APE: Agentic Prompt Enhancer for Image Generation and Editing cites this paper.

APE: Agentic Prompt Enhancer for Image Generation and Editing Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-06-28T23:02:46.691177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T22:58:17.807988Z digest=sha256:cfc07f6ef950f5c941c8325f99a4f88d44a25b3d5c1b713e1c14870cbbf23b75

Observation 99549bc5-3916-4355-85bf-4c8b726b27d2 · inbound

MemoGen: Can Past Experience Improve Future Text-to-Image Generation? cites this paper.

MemoGen: Can Past Experience Improve Future Text-to-Image Generation? Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:36:27.152945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T10:48:18.630777Z digest=sha256:0891f0a870fdb65ae8ec79ae0387377024f4bc4b2969077aa5abcd8b72e16d13

Observation 84b95862-e056-43f9-b6d3-40817345acc6 · inbound

Do Models Share Safety Representations? Cross-Model Steering for Safe Visual Generation cites this paper.

Do Models Share Safety Representations? Cross-Model Steering for Safe Visual Generation Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-02T08:16:47.710013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T06:16:53.305354Z digest=sha256:6c8997050b65aff3d94964f3620dc7f4cb2a33e973165abc10b4b0aa68bb1909

Observation 00b2c270-c0c3-4a1f-822e-07dff5e6702f · inbound

TextWand: A Unified Framework for Scene Text Editing cites this paper.

TextWand: A Unified Framework for Scene Text Editing Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 29

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T12:36:57.307931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T01:56:46.514177Z digest=sha256:12718770eaac10bf0684229c7117a40ca77badc5fc639ef1d5d26a6eb08f3610

Observation 873e85cf-971b-4c71-8d0f-8c502e25b75b · inbound

Faithful, Enriched, and Precise: Benchmarking Natural-Science Illustration Generation by T2I models cites this paper.

Faithful, Enriched, and Precise: Benchmarking Natural-Science Illustration Generation by T2I models Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-07-02T11:46:55.812192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T03:00:00.288142Z digest=sha256:f03c91802f5769b7efe1b708a63ecd7b091a08f6ce0cf1e0dc50e6a2893a2685

Observation 3da1bde6-f4c2-476d-84de-f8287c8eef86 · inbound

NutriMLLM: Multimodal Large Language Models for Dietary Micronutrient Analysis cites this paper.

NutriMLLM: Multimodal Large Language Models for Dietary Micronutrient Analysis Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 22

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T00:17:28.839137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T17:22:29.806105Z digest=sha256:993068e03be7df7daf9464d88a2765cd6646f9f3e69ef59c4877838774c77081

Observation eff6429b-4930-413c-b119-94d6da010edd · inbound

Z-Reward: Beyond Scalar Rewards by Internalizing Reasoning into Score Distributions cites this paper.

Z-Reward: Beyond Scalar Rewards by Internalizing Reasoning into Score Distributions Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-07-03T00:07:27.947025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T17:30:57.001021Z digest=sha256:bb405affb7698ba46a03ca7ff1f96714a83fbb99feb5b7848150589c9327c4ea

Observation b4562b7f-ec49-4b61-aa2b-eedd15c64ef6 · inbound

Z-Reward: Beyond Scalar Rewards by Internalizing Reasoning into Score Distributions cites this paper.

Z-Reward: Beyond Scalar Rewards by Internalizing Reasoning into Score Distributions Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 51

Resolution
unresolved
no resolver link, observed 2026-07-15T10:53:37.186361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T10:53:37.186361Z digest=sha256:8e1564b0099d7744eb5673df1986e2ab070292fe845eca88a3241d15ef0f82e4

Observation 890f8b57-61be-47be-b68a-9f0b8abd0265 · inbound

Can Image Models Imagine Time? ImageTime: A Novel Benchmark for Probing Visual World Modeling Through Spatiotemporal Consistency cites this paper.

Can Image Models Imagine Time? ImageTime: A Novel Benchmark for Probing Visual World Modeling Through Spatiotemporal Consistency Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-07-03T04:57:38.065459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T13:34:25.079037Z digest=sha256:49929a21972962b4fe4cb0ef4efb0e86701430ccbe3b75032ed1c826a25fe7df

Observation a2cd6be4-646e-488c-8f30-afb48953ef69 · inbound

The 1st PortraitCraft Challenge: A CVPR 2026 Workshop Competition on Portrait Composition Understanding and Generation cites this paper.

The 1st PortraitCraft Challenge: A CVPR 2026 Workshop Competition on Portrait Composition Understanding and Generation Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-07-03T04:47:38.642928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T13:36:15.259428Z digest=sha256:6f4706bf65ca1566088f2b3de34a6fc26ed9dbb1d3e7dd8adc6c7462c070ba43