Pith. sign in

Paper Citation Record · LEDGER

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark

As of 7 August 2026, this Paper Citation Record lists 72 of 72 outbound references and 10 inbound Pith citation observations for arXiv:2509.09680.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.09680 v1

Coverage vector

measured 72 of 72 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T18:48:03.885237Z

measured 82 of 82 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T04:09:49.613004Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T06:54:20.702391Z

Reference resolution

72 of 72 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved72
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f94e70a3-b1a5-48c8-84da-86480cc79bb4 · outbound

This paper cites Qwen2.5-VL Technical Report.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Qwen2.5-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.563904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.563904Z digest=sha256:ac73fea4537ceb5ddb3be9bb06f76704877fc0ae400e5ea7f97b6d1bede40f1a

Observation 43c2b44d-78a7-4e5d-bd7a-5685ed2424ae · outbound

This paper cites Hrs-bench: Holistic, reliable and scalable benchmark for text-to-image models.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Hrs-bench: Holistic, reliable and scalable benchmark for text-to-image models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.570285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.570285Z digest=sha256:1b61802529ce87126732adf4a8365f9e6bc1c48c76cfcc1fb63329d789949b70

Observation 0e83f8f1-4be9-4566-8f41-c96b1d33e98a · outbound

This paper cites Flux, 2024.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Flux, 2024

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.574995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.574995Z digest=sha256:4693c0300389246ed789ff55388908b2190216823d790c46486bb843e1f22e9a

Observation 44bd8c79-4409-4e9f-91aa-a0c4607246f6 · outbound

This paper cites Flux.1 krea, 2025.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Flux.1 krea, 2025

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.579471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.579471Z digest=sha256:b92cf063bfc33119ad4cab1c902845e5dbffde47db78c3af978243cc30f18088

Observation 01df3abc-a748-4c1f-897a-12ea03a3e2d2 · outbound

This paper cites HiDream-I1: A High-Efficient Image Generative Foundation Model with Sparse Diffusion Transformer.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark HiDream-I1: A High-Efficient Image Generative Foundation Model with Sparse Diffusion Transformer

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.584164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.584164Z digest=sha256:7116c71254ad82f5e070d9d9dfa7e2eee2eabbda875fab04461531a6c4b6af31

Observation 8744a1d0-620d-4afb-8e54-d2144bda7a69 · outbound

This paper cites Conceptual 12m: Pushing web- scale image-text pre-training to recognize long-tail visual concepts.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Conceptual 12m: Pushing web- scale image-text pre-training to recognize long-tail visual concepts

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.588996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.588996Z digest=sha256:c055ccc548e665303b7a830134f220df776d73cdf77a7c0c4edc9d5b0bd8af5f

Observation 0d15c85a-743b-404b-a2bd-92fb6e79b5df · outbound

This paper cites Attend-and-excite: Attention- based semantic guidance for text-to-image diffusion models.ACM transactions on Graphics (TOG), 42 (4):1–10, 2023.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Attend-and-excite: Attention- based semantic guidance for text-to-image diffusion models.ACM transactions on Graphics (TOG), 42 (4):1–10, 2023

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.594301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.594301Z digest=sha256:b43066d46a9555aa6899d6fe8a8ccf8aa82fbbac46d6c9b962fffb0ba677ef68

Observation e03adb8d-a695-4f2b-a510-211206bdb279 · outbound

This paper cites Textdiffuser: Diffusion models as text painters.Advances in Neural Information Processing Systems, 36:9353–9387, 2023.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Textdiffuser: Diffusion models as text painters.Advances in Neural Information Processing Systems, 36:9353–9387, 2023

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.598554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.598554Z digest=sha256:eca33b9d2ac40b61c8197ffbf7304f6823f67f7d81ab63d392770e8e32036005

Observation d11eda94-b037-487d-b91b-fb9edf9641a0 · outbound

This paper cites BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.602753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.602753Z digest=sha256:f5fa21d37996935b7589f2aecbccdd4e78a66cddddcce94f0307976a9de65350

Observation d52c7b55-7cc0-4e1a-b710-fc45ed3298d6 · outbound

This paper cites PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.608245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.608245Z digest=sha256:696e19429a50f681bc6ce916434c2ed8b113c0e68d7ff5789e343559c3503349

Observation 61fb0654-ee82-4366-9b49-ebdf97ab9ccb · outbound

This paper cites PosterCraft: Rethinking High-Quality Aesthetic Poster Generation in a Unified Framework.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark PosterCraft: Rethinking High-Quality Aesthetic Poster Generation in a Unified Framework

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.613130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.613130Z digest=sha256:87e44890c378e2d28e9ca101d0ecd3fa81dcd365869e97e40bb6646b59b21804

Observation 493d2111-f5c4-4ee3-847f-f6256d730c5a · outbound

This paper cites Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.618107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.618107Z digest=sha256:202c2cb41a0659e6a0bb593ade0a1bfd554914a39d63eeca854ac783239eb9be

Observation 1e079fd0-bf99-4dbb-9b5a-b28863b6d373 · outbound

This paper cites Davidsonian Scene Graph: Improving Reliability in Fine-grained Evaluation for Text-to-Image Generation.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Davidsonian Scene Graph: Improving Reliability in Fine-grained Evaluation for Text-to-Image Generation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.622457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.622457Z digest=sha256:db4733f47f791cdaf9bc915f03444336149bc14eb6fc0cb778ed0187541eba4f

Observation 102199fd-c156-441e-b7d0-d8e8ef9efe48 · outbound

This paper cites Emerging Properties in Unified Multimodal Pretraining.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Emerging Properties in Unified Multimodal Pretraining

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.626833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.626833Z digest=sha256:bcf04dcd507d77da5927f1d0b21ba4ed81885ea2ba7ec0261768f3be9277a1cd

Observation 5aa4c4a5-d4bd-4dba-a56a-700e5605e989 · outbound

This paper cites GoT-R1: Unleashing Reasoning Capability of MLLM for Visual Generation with Reinforcement Learning.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark GoT-R1: Unleashing Reasoning Capability of MLLM for Visual Generation with Reinforcement Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.631140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.631140Z digest=sha256:44493710243f933e499d1e959185137c1322603dbbb599e29bb9bd6fca450f0c

Observation 9c2c9c9f-153d-4808-aad0-f3dc622ed302 · outbound

This paper cites Scaling rectified flow transformers for high-resolution image synthesis.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Scaling rectified flow transformers for high-resolution image synthesis

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.635989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.635989Z digest=sha256:78a796e3499dd459d7c5a784b1b0da16146b19994111d3e0b23e055d7deacc14

Observation c3c5cd98-926e-4ae3-afbd-c637bb033a3c · outbound

This paper cites PUMA: Empowering Unified MLLM with Multi-granular Visual Generation.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark PUMA: Empowering Unified MLLM with Multi-granular Visual Generation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.640677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.640677Z digest=sha256:7a166594be538ef5c2768581e915d41e76fd6f14e94aad70fb4770862951b0a1

Observation 15feb0bb-9423-4260-a030-ccd50ba0b956 · outbound

This paper cites GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.645811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.645811Z digest=sha256:d5e5a77cfe25996786abad761700d91c5eacce57a25cd968504e68e05d882bd5

Observation 0f6fadec-312e-44c2-b44d-623bec1bce2f · outbound

This paper cites Training-Free Structured Diffusion Guidance for Compositional Text-to-Image Synthesis.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Training-Free Structured Diffusion Guidance for Compositional Text-to-Image Synthesis

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.650851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.650851Z digest=sha256:f814a166f9bb33b0a697e316298f8e3cee21e9affa45fac558d97927bc921443

Observation 66c32426-e674-488e-8e1d-aaad6cc0f225 · outbound

This paper cites Commonsense-T2I Challenge: Can Text-to-Image Generation Models Understand Commonsense?.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Commonsense-T2I Challenge: Can Text-to-Image Generation Models Understand Commonsense?

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.655845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.655845Z digest=sha256:3f54511d47ba161ae4274a8a8cdb3ad9e98e92df82404dd29e526aa4dcdabd2f

Observation 4a359ab1-f411-42fc-9429-91ab5da67e97 · outbound

This paper cites Datacomp: In search of the next generation of multimodal datasets.Advances in Neural Information Processing Systems, 36:27092–27112, 2023.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Datacomp: In search of the next generation of multimodal datasets.Advances in Neural Information Processing Systems, 36:27092–27112, 2023

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.661172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.661172Z digest=sha256:9d75f51c9854dd94ee9efbd1af09b30e7d16f2c63cb56a1950fe1c661a091258

Observation 22dd52d8-9ef1-4103-bfb2-6d994cfa619d · outbound

This paper cites Seedream 3.0 Technical Report.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Seedream 3.0 Technical Report

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.665647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.665647Z digest=sha256:49c07e0d55264fdc0c4356683fee758cf90f949c00f3475c031e46d3ebcde0bc

Observation 1d147607-e9e8-4fc6-9e11-1252bb0c4383 · outbound

This paper cites Geneval: An object-focused framework for evaluating text-to-image alignment.Advances in Neural Information Processing Systems, 36:52132– 52152, 2023.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Geneval: An object-focused framework for evaluating text-to-image alignment.Advances in Neural Information Processing Systems, 36:52132– 52152, 2023

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.670397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.670397Z digest=sha256:0261c5cb96978e7361dbbd58b72d408a4953b415dad85463690e08299f05f082

Observation 20adb2b8-4782-4348-8e6b-348dc4ba1ab5 · outbound

This paper cites Seedream 2.0: A Native Chinese-English Bilingual Image Generation Foundation Model.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Seedream 2.0: A Native Chinese-English Bilingual Image Generation Foundation Model

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.674844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.674844Z digest=sha256:5d673eb267746c8d1be1346257fd05ae31e86d4868524d370b42d2d47fdeaf7b

Observation a7103d71-2262-4a62-a77e-d9d3114b67ba · outbound

This paper cites Gemini2.5-pro, 2025.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Gemini2.5-pro, 2025

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.679408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.679408Z digest=sha256:5066c2bd0927ff024a8a972ed35f92d68ca8798e1f5f4ba700cab942f384d6cd

Observation a4855d2b-abd7-4496-988e-e78d7d438bb5 · outbound

This paper cites Imagen4, 2025.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Imagen4, 2025

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.683844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.683844Z digest=sha256:eca1420fbb8cf0c9bfd4c110ede9a6b39ef69d3cabf8f561f566199c671265df

Observation 88f7b782-fd25-43b6-a00a-426d7ba02dfe · outbound

This paper cites Gemini2.5-flash-image, 2025.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Gemini2.5-flash-image, 2025

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.688106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.688106Z digest=sha256:007bb725311acc2767c46887a5ce949f03a0e65edaa377d4ab3445b072a80864

Observation f050f52b-e189-4d85-9826-0aa32acfc810 · outbound

This paper cites EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.692481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.692481Z digest=sha256:dc9c3d5e9effddb7f2bfbf18fc6fb84146096ed6f6b8f3972ff5b1b9f3c09bbb

Observation 6f6cefc9-a127-45e4-9762-eedb1cfcf7f9 · outbound

This paper cites CLIPScore: A Reference-free Evaluation Metric for Image Captioning.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark CLIPScore: A Reference-free Evaluation Metric for Image Captioning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.697231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.697231Z digest=sha256:42d52109fc774fe90f94a8fb770ac8b2a42cb3850770fc0558106614d7af2c54

Observation d3674176-1a2c-4c13-a57e-88a56dbbaee2 · outbound

This paper cites Denoising diffusion probabilistic models.Advances in neural information processing systems, 33:6840–6851, 2020.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Denoising diffusion probabilistic models.Advances in neural information processing systems, 33:6840–6851, 2020

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.701850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.701850Z digest=sha256:a20b942f692ca35ab97e292c9e690346730edcc8075d021f7b710a4595035723

Observation b0fb7470-f858-49c1-a777-2b6f6d7da733 · outbound

This paper cites Scaling up vision-language pre-training for image captioning.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Scaling up vision-language pre-training for image captioning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.706247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.706247Z digest=sha256:77b6c980f462df042549ac60329e24706502c8f04dec22b9e6c39c81cb5f43c0

Observation bb319ba8-5638-4b97-a767-a9b0cb625bb2 · outbound

This paper cites ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.710756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.710756Z digest=sha256:0c2eba5ee0718bb1f8102b2c1e000cf4aea168838c481bf383198ddf51db78c4

Observation d87c14be-6654-48a6-9516-ba005052566b · outbound

This paper cites Tifa: Accurate and interpretable text-to-image faithfulness evaluation with question answer- ing.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Tifa: Accurate and interpretable text-to-image faithfulness evaluation with question answer- ing

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.715302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.715302Z digest=sha256:2482c2936c3f4c9b7a0d851abb18d3eff3b771d1704f8bab5158fd3cc8238d80

Observation c7fce712-9928-48a5-b394-133de9993800 · outbound

This paper cites T2i-compbench: A comprehensive benchmark for open-world compositional text-to-image generation.Advances in Neural Information Processing Systems, 36:78723–78747, 2023.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark T2i-compbench: A comprehensive benchmark for open-world compositional text-to-image generation.Advances in Neural Information Processing Systems, 36:78723–78747, 2023

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.719868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.719868Z digest=sha256:9afa203cf2da12c69d4997d8db5e2c2fa2d6cbac52c1a16ec58694d171956bb3

Observation 7944bef0-bd39-4828-a8ff-83ebf1886063 · outbound

This paper cites Scaling up visual and vision-language representation learning with noisy text supervision.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Scaling up visual and vision-language representation learning with noisy text supervision

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.724083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.724083Z digest=sha256:65f8dcfe90b37e7d316689a62c229be76470e66201533840987f37492e87c697

Observation 19fe8c2c-f8bf-4602-a64b-228326cd7725 · outbound

This paper cites Genetic k-means algorithm.IEEE Transactions on Systems, Man, and Cybernetics, Part B (Cybernetics), 29(3):433–439, 1999.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Genetic k-means algorithm.IEEE Transactions on Systems, Man, and Cybernetics, Part B (Cybernetics), 29(3):433–439, 1999

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.728017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.728017Z digest=sha256:d48c72dc0ad6236da651561da3d9c2486d2468e636da9900ad905d41a9c8dc53

Observation 817fe4d2-c078-4b26-b7d9-ae1f0aec9d92 · outbound

This paper cites GenAI-Bench: Evaluating and Improving Compositional Text-to-Visual Generation.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark GenAI-Bench: Evaluating and Improving Compositional Text-to-Visual Generation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.732027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.732027Z digest=sha256:92f44d1d191038e6303cbeab446c8c969c078b2a9096cb7f516df3af1b9ef07d

Observation b6ce2eb6-4a67-476f-b610-2c536fc11828 · outbound

This paper cites Playground v2.5: Three Insights towards Enhancing Aesthetic Quality in Text-to-Image Generation.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Playground v2.5: Three Insights towards Enhancing Aesthetic Quality in Text-to-Image Generation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.736630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.736630Z digest=sha256:0cbf1b7be723a3081cd9d4eacd6d73cc7e94a82678cf9226dcaf2019c5fa77c0

Observation b371269d-14db-4774-babd-8ae7f7a9625d · outbound

This paper cites Oscar: Object-semantics aligned pre-training for vision-language tasks.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Oscar: Object-semantics aligned pre-training for vision-language tasks

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.741275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.741275Z digest=sha256:b7c57d60b01f1bab71939eb2736e3d1fd155bf7235d636a8e470cd2af9811881

Observation 0356551c-13d7-480c-8ac5-f129a797c09f · outbound

This paper cites Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.745641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.745641Z digest=sha256:ae9e715917c4c4f0b743bf2fad3ccf824eb96779473556c7a8d6d01d8ff9892f

Observation c24289ca-fdc0-419f-a700-62dc2927313b · outbound

This paper cites Nuwa-infinity: Autoregressive over autoregressive generation for infinite visual synthesis.Advances in Neural Information Processing Systems, 35:15420–15432, 2022.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Nuwa-infinity: Autoregressive over autoregressive generation for infinite visual synthesis.Advances in Neural Information Processing Systems, 35:15420–15432, 2022

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.749799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.749799Z digest=sha256:2fcb429e41e4c2a19f6b54a9e7afa7ab40043590519da06d9e710fe1fe57015f

Observation 2157b36c-7885-484e-a4f1-3c621f24ba8b · outbound

This paper cites Evaluating text-to-visual generation with image-to-text generation.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Evaluating text-to-visual generation with image-to-text generation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.754532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.754532Z digest=sha256:9017548d9ede8c535eeaebc2c7748455d3347048685bf498766101a5199051e8

Observation 8ca082a6-7773-467d-af2e-1bcf828cd6e2 · outbound

This paper cites Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.759187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.759187Z digest=sha256:8f2c9e0ac5cd6887f66ff1f01e3d890f4a7de5ef0ef0c3f4bb17336e25c3bc7d

Observation dda439af-19b4-429b-938c-c35248794e9a · outbound

This paper cites Summarizing emotions from text using plutchik’s wheel of emotions.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Summarizing emotions from text using plutchik’s wheel of emotions

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.764203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.764203Z digest=sha256:1a9aacc02b6b14974dac03898242dc61856850fcc042d26917bcecc994b8ae65

Observation 3d5ed760-fcb6-457c-a127-e274c368c6a3 · outbound

This paper cites Gpt-4.1, 2025.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Gpt-4.1, 2025

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.768998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.768998Z digest=sha256:81a0a0a5c3e4e23693561e4cf785bc1588bd6010cdd15772a17ae316caacac9d

Observation 14637733-d07d-4f0d-a3a7-94d010d9e8fa · outbound

This paper cites Gpt-image-1, 2025.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Gpt-image-1, 2025

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.773709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.773709Z digest=sha256:5e0711712b7ef3aa62a23bbb832f9595304756c06c2ef3faa04c880fdfd466d7

Observation b8bc7723-71da-420b-afb6-57ca6ea7c95e · outbound

This paper cites Dall·e 3, September 2023.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Dall·e 3, September 2023

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.778454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.778454Z digest=sha256:87f313b7a1c85dedd3a11ac294277c27016cad0cee96aa0be3761e314273f1f4

Observation 20cb7434-a7c0-466f-a31e-d8119da7e8d8 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.782559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.782559Z digest=sha256:d7226034293d067c1f1b586eee439fd9b3ae45c60b3f04b56b7f6f5bb8fb94e8

Observation cbbeac40-3f01-459a-84ed-0493b3223119 · outbound

This paper cites Lumina-Image 2.0: A Unified and Efficient Image Generative Framework.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Lumina-Image 2.0: A Unified and Efficient Image Generative Framework

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.787032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.787032Z digest=sha256:952b459bd8c66dd41938911ffedc6c37ffa1db8c1df0594058061746899bedd9

Observation e2299ed6-5290-44bf-9101-c6001ce104c8 · outbound

This paper cites High- resolution image synthesis with latent diffusion models.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark High- resolution image synthesis with latent diffusion models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.791588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.791588Z digest=sha256:73214173a90a3944cd8870afa427dd053345b71222bb60f844adf23c28d6d605

Observation 00b83a60-5eb1-4ff5-8eb9-0d3a3afa232f · outbound

This paper cites Laion-5b: An open large-scale dataset for training next generation image-text models.Advances in neural information processing systems, 35:25278–25294, 2022.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Laion-5b: An open large-scale dataset for training next generation image-text models.Advances in neural information processing systems, 35:25278–25294, 2022

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.795622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.795622Z digest=sha256:960d5a05269fa473b318147e22f6f5826da1b8f2bbadd5e4215e41cdef0a0b99

Observation bb70cb03-f76e-4543-bde6-2c52f6442a69 · outbound

This paper cites Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.799595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.799595Z digest=sha256:9a435ca55ff8947a21c4699b57b98f5dc9348af74d5f4e55432106b69b9f00eb

Observation a40919a6-6724-4fa7-b093-3f5f9d14a557 · outbound

This paper cites From Pixels to Prose: A Large Dataset of Dense Image Captions.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark From Pixels to Prose: A Large Dataset of Dense Image Captions

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.803853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.803853Z digest=sha256:4aa5b0635eca543f5b09970f65eee32928e64dcf6d4f054b278fe09efbfcd830

Observation 767ffcd6-8a63-472b-838c-284ae1be0a57 · outbound

This paper cites Stable diffusion 2.1, 2022.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Stable diffusion 2.1, 2022

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.808078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.808078Z digest=sha256:7592942c99f3b7ecc8a4583dab9025ba6b7559207f009259ff8aeb8720f4a18f

Observation c8ce336c-b8b7-4bfa-a370-63d1d229dd7f · outbound

This paper cites Stable diffusion 3, 2024.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Stable diffusion 3, 2024

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.812209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.812209Z digest=sha256:0f22181696af369c04428d0d710b5be2dbdc53eb083a923163d7790dda503850

Observation 6437c0c5-32ed-4980-8808-28a8ca0458c8 · outbound

This paper cites Stable diffusion 3.5, 2024.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Stable diffusion 3.5, 2024

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.816335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.816335Z digest=sha256:39dca45424a08f8b0065051d18cce097d3f174c9afbbec81a88a536f126dde1e

Observation 05a59200-9ff6-43f0-b1f9-ba3cd7497707 · outbound

This paper cites T2I-ReasonBench: Benchmarking Reasoning-Informed Text-to-Image Generation.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark T2I-ReasonBench: Benchmarking Reasoning-Informed Text-to-Image Generation

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.820245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.820245Z digest=sha256:bbc4abc2dd2e41cba8bcb838b820a838b7b40192e9c7ab162892b447577981f8

Observation 9d54b302-4927-4779-9b8d-1947f2a21fa7 · outbound

This paper cites Journeydb: A benchmark for generative image understanding.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Journeydb: A benchmark for generative image understanding

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.824740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.824740Z digest=sha256:2e8ffcfdcdad4dc302fa53f2f2ac6c54d3725fb654e0c8741cf6540a337cb62a

Observation d8dfc13f-ab11-4415-b4fb-afc6ddb6ecd0 · outbound

This paper cites AnyText: Multilingual Visual Text Generation And Editing.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark AnyText: Multilingual Visual Text Generation And Editing

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.828806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.828806Z digest=sha256:3d766e44511a0122c31f8a99e010d9fdb5e701a5d7abadba07cf910b78cf6aeb

Observation 9314c1d1-ec1f-4db1-867a-30d8c7dfebc2 · outbound

This paper cites Textatlas5m: A large-scale dataset for dense text image generation.arXiv preprint arXiv:2502.07870, 2025.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Textatlas5m: A large-scale dataset for dense text image generation.arXiv preprint arXiv:2502.07870, 2025

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.832885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.832885Z digest=sha256:a65baace43e8bd7b110dc4489d406e8b2527061a26b6cba0d1203a5fe6835c3f

Observation 7a952e4b-07f4-4ffa-92fe-5179ad069cae · outbound

This paper cites Nuwa: Visual synthesis pre-training for neural visual world creation.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Nuwa: Visual synthesis pre-training for neural visual world creation

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.836817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.836817Z digest=sha256:ba43d6b6bc2736fa20d8c679d2640d31335d8427521136f32146284f5eb89ebb

Observation 7d331a1d-46c9-4d4d-b0a4-e012fdae3b65 · outbound

This paper cites Qwen-Image Technical Report.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Qwen-Image Technical Report

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.840789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.840789Z digest=sha256:4e732d1f06c16988919ce91d49c353cda18ab303e28cb3040981ad57d0762614

Observation 41db8f31-f4b4-4a70-b193-2734420b0857 · outbound

This paper cites OmniGen2: Towards Instruction-Aligned Multimodal Generation.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.845193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.845193Z digest=sha256:7c8f44f82909cfac0b685b7d9093c671e0fc7715bda61e388de90b2d0302ec2c

Observation 938a92ab-78fd-47e7-a096-b670c3016288 · outbound

This paper cites Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.849652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.849652Z digest=sha256:f615ad0036ca1976d7c723f70ab6ec90abca2c2f5845259b6fc941b4f6e6de62

Observation d8d8bce3-6b48-4a3e-9891-050a87ba4cf2 · outbound

This paper cites Conceptmix: A com- positional image generation benchmark with controllable difficulty.Advances in Neural Information Processing Systems, 37:86004–86047, 2024.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Conceptmix: A com- positional image generation benchmark with controllable difficulty.Advances in Neural Information Processing Systems, 37:86004–86047, 2024

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.854018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.854018Z digest=sha256:e8b530f750d738e0816967a7fc454c5a241837dcf88377574f42419873dc219e

Observation 29c414a0-efcf-470c-a956-7d6eb10b54d1 · outbound

This paper cites SANA 1.5: Efficient Scaling of Training-Time and Inference-Time Compute in Linear Diffusion Transformer.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark SANA 1.5: Efficient Scaling of Training-Time and Inference-Time Compute in Linear Diffusion Transformer

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.858574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.858574Z digest=sha256:3d343fc3b921abd124afc59281ae0537fdb03c12821f166440070a9191c14ca8

Observation e6ece735-e814-416a-ba27-7dc34e34b37f · outbound

This paper cites Show-o2: Improved Native Unified Multimodal Models.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Show-o2: Improved Native Unified Multimodal Models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.862906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.862906Z digest=sha256:dd818426dadaca8a17889fac4e460ed6159fa87e8008c09d950ed0af3daa8559

Observation db94fd1b-4d1a-4e33-80da-4069954ad002 · outbound

This paper cites Qwen3 Technical Report.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Qwen3 Technical Report

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.867345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.867345Z digest=sha256:88be7134c89f64ef7bc481a75a873a1fef9e93d9ea2564c608f9befbf79faf55

Observation d6030455-2072-4d5c-8360-fa3b2aee2280 · outbound

This paper cites Scaling Autoregressive Models for Content-Rich Text-to-Image Generation.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.872017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.872017Z digest=sha256:796b492ead4232d27ba3834a02c3b451404e1d7192e277e4fb85847dc1bb7a69

Observation 5216cd83-fb07-494a-a7c0-8828e1a8916c · outbound

This paper cites Capsfusion: Rethinking image-text data at scale.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Capsfusion: Rethinking image-text data at scale

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.876743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.876743Z digest=sha256:bdfdd47fdd4c9f16e943414456484a7ab49bd66ec36d47c12053d00d09238b4d

Observation 700b9ab6-fd88-4774-99c9-fa3b8739461a · outbound

This paper cites VinVL: Revisiting Visual Representations in Vision-Language Models.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark VinVL: Revisiting Visual Representations in Vision-Language Models

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.880972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.880972Z digest=sha256:3b7df5e1957e7040165362df1de2166e3adb74f1d87deb8d6047ba2ba54c36a5

Observation 8da24999-f075-454c-88ba-faedb3fde794 · outbound

This paper cites Lumina-next: Making lumina-t2x stronger and faster with next-dit.Advances in Neural Information Processing Systems, 37:131278–131315, 2024.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Lumina-next: Making lumina-t2x stronger and faster with next-dit.Advances in Neural Information Processing Systems, 37:131278–131315, 2024

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.885237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.885237Z digest=sha256:97d71d85b2d6bc26ea433dc3fd6b861441c2411452e889e1acad532c863be09e

Pith citing papers

Observation a2d368cc-6e46-4376-b939-0cb33154a5af · inbound

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer cites this paper.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:08:37.030219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T14:08:36.801359Z digest=sha256:7f3f2a69fb88ba10e34898596615325f4a389418b6ac7824e7941a722503818f

Observation 6b6f9400-4ee2-466a-9de1-97b015cb413c · inbound

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer cites this paper.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:31.327318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:31.327318Z digest=sha256:aba704fb018a3301dbd8bb38c5761feacc1e00a89cb4c3c7e88c820070729e76

Observation fce86933-3f87-4417-89e9-eac9301d2ba8 · inbound

Guiding Token-Sparse Diffusion Models cites this paper.

Guiding Token-Sparse Diffusion Models FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T12:50:55.743597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:50:55.743597Z digest=sha256:8311a1ed39a9cbc1612168a83f5a930eee0a97a3fe2412478bb970c352017ccd

Observation 3dccef82-3cf6-4e2a-b349-6088b0cc0d1d · inbound

Self-Adversarial One Step Generation via Condition Shifting cites this paper.

Self-Adversarial One Step Generation via Condition Shifting FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:11:01.970606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:37:36.994169Z digest=sha256:65082a521b29232dcf3798bceb6992822fba165bc8ddd76305374ed6d18520fd

Observation d77b53bb-b502-4f90-a876-656d0aa4aa6a · inbound

LLaDA2.0-Uni: Unifying Multimodal Understanding and Generation with Diffusion Large Language Model cites this paper.

LLaDA2.0-Uni: Unifying Multimodal Understanding and Generation with Diffusion Large Language Model FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-10T00:49:48.344114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T00:49:38.156237Z digest=sha256:27251f3bcd34e74ebc1b5ace35193b7c38867d466fba697bcb1f32d750f61ba7

Observation df3ef91b-bdce-4eab-b933-b60bd3102b55 · inbound

InsightTok: Improving Text and Face Fidelity in Discrete Tokenization for Autoregressive Image Generation cites this paper.

InsightTok: Improving Text and Face Fidelity in Discrete Tokenization for Autoregressive Image Generation FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:13:30.594327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T02:10:18.004724Z digest=sha256:74f0e54400ca0617bd8a59fd1b7c5d997e95165d00f9840ec90d31c7bd7469df

Observation fe9d9296-2407-4e88-90a3-9ca455a86712 · inbound

Unlocking Complex Visual Generation via Closed-Loop Verified Reasoning cites this paper.

Unlocking Complex Visual Generation via Closed-Loop Verified Reasoning FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-19T16:37:39.647701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T16:35:43.166697Z digest=sha256:473a38432e2b8aedade4a8c719b58209b099a3908832e34ce013f075d0d6f511

Observation 6850081c-98f7-4258-831d-423ea5956fb1 · inbound

SciIR: A Large-scale Training Dataset and Benchmark for Scientific Image Reasoning Generation cites this paper.

SciIR: A Large-scale Training Dataset and Benchmark for Scientific Image Reasoning Generation FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T06:54:20.704034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T06:48:22.594002Z digest=sha256:fd32649dc5db25df3bc4aac946b17b5878ef354ddae48ef5b3aa611bd52ccf2d

Observation 0b8bbb34-8a38-4ddd-9730-8cd48d5abea2 · inbound

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric cites this paper.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-01T15:42:58.828756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:42:58.828756Z digest=sha256:b301855bea4714b965892cbe088e7eb8e5821a0d4287df047d22d30de6c9ef4d

Observation a87f0171-fe34-42dc-86dc-ead1f635dfe0 · inbound

CAPEval: A Decoupled Caption Evaluation across Understanding and Generation cites this paper.

CAPEval: A Decoupled Caption Evaluation across Understanding and Generation FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T04:09:49.613004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:09:49.613004Z digest=sha256:57e1bcf833da233643759ac17c204311925a5882b0012d38a95e287e37ca31f9