Pith. sign in

Paper Citation Record · LEDGER

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark

As of 19 August 2026, this Paper Citation Record lists 72 of 72 outbound references and 10 inbound Pith citation observations for arXiv:2509.09680.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.09680 v1

Coverage vector

measured 72 of 72 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T18:48:03.885237Z

measured 82 of 82 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T04:09:49.613004Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T06:54:20.702391Z

Reference resolution

72 of 72 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved72
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f94e70a3-b1a5-48c8-84da-86480cc79bb4 · outbound

This paper cites Qwen2.5-VL Technical Report.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Qwen2.5-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.563904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.563904Z digest=sha256:2aad2aa956add9123879d2374e86f1d55e350c33b7b5beb32fa991d9a0c94643

Observation 43c2b44d-78a7-4e5d-bd7a-5685ed2424ae · outbound

This paper cites Hrs-bench: Holistic, reliable and scalable benchmark for text-to-image models.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Hrs-bench: Holistic, reliable and scalable benchmark for text-to-image models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.570285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.570285Z digest=sha256:389ebfb81fce4b0c12377f96b57db8a2d879ac86f3f63b657d3d87aa061c181b

Observation 0e83f8f1-4be9-4566-8f41-c96b1d33e98a · outbound

This paper cites Flux, 2024.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Flux, 2024

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.574995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.574995Z digest=sha256:9dae5cb73a0b8f24fb140814c017bca2e4e6e2bdc54195e2b4ffc1bd4354b78f

Observation 44bd8c79-4409-4e9f-91aa-a0c4607246f6 · outbound

This paper cites Flux.1 krea, 2025.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Flux.1 krea, 2025

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.579471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.579471Z digest=sha256:aacb9968818051b8d55c87155eca44c95c962b066350a59b44422872b1bf3e13

Observation 01df3abc-a748-4c1f-897a-12ea03a3e2d2 · outbound

This paper cites HiDream-I1: A High-Efficient Image Generative Foundation Model with Sparse Diffusion Transformer.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark HiDream-I1: A High-Efficient Image Generative Foundation Model with Sparse Diffusion Transformer

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.584164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.584164Z digest=sha256:ce5aef42386cb221f81c0b9b9c15531f381a963502d4bca1246ad66af04afe61

Observation 8744a1d0-620d-4afb-8e54-d2144bda7a69 · outbound

This paper cites Conceptual 12m: Pushing web- scale image-text pre-training to recognize long-tail visual concepts.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Conceptual 12m: Pushing web- scale image-text pre-training to recognize long-tail visual concepts

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.588996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.588996Z digest=sha256:822e8fb35df18e9fac152a44eacfa5ccc7e72fcaee10c3689831fe67031cffed

Observation 0d15c85a-743b-404b-a2bd-92fb6e79b5df · outbound

This paper cites Attend-and-excite: Attention- based semantic guidance for text-to-image diffusion models.ACM transactions on Graphics (TOG), 42 (4):1–10, 2023.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Attend-and-excite: Attention- based semantic guidance for text-to-image diffusion models.ACM transactions on Graphics (TOG), 42 (4):1–10, 2023

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.594301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.594301Z digest=sha256:394f3f7f53f75dc153aff561d4922e4d05b53340a0a1a163c94be0fa96e37bd1

Observation e03adb8d-a695-4f2b-a510-211206bdb279 · outbound

This paper cites Textdiffuser: Diffusion models as text painters.Advances in Neural Information Processing Systems, 36:9353–9387, 2023.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Textdiffuser: Diffusion models as text painters.Advances in Neural Information Processing Systems, 36:9353–9387, 2023

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.598554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.598554Z digest=sha256:8f264a2fce6b499627b301608a5a6ca303a0f468885b760551bb62317acd5afe

Observation d11eda94-b037-487d-b91b-fb9edf9641a0 · outbound

This paper cites BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.602753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.602753Z digest=sha256:5f6eeb1a125b23752468969be66c9e8eba6aa8f1e9a06021a8bbfaded7b55a45

Observation d52c7b55-7cc0-4e1a-b710-fc45ed3298d6 · outbound

This paper cites PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.608245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.608245Z digest=sha256:6eafcab1eaca0319a97c573bb79198653fcefb370be433a04913894f4c407b68

Observation 61fb0654-ee82-4366-9b49-ebdf97ab9ccb · outbound

This paper cites PosterCraft: Rethinking High-Quality Aesthetic Poster Generation in a Unified Framework.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark PosterCraft: Rethinking High-Quality Aesthetic Poster Generation in a Unified Framework

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.613130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.613130Z digest=sha256:fc64ec45da91e8a022f3bac9e9cafc0429446fb9349aab5c4e325d68b38ecc76

Observation 493d2111-f5c4-4ee3-847f-f6256d730c5a · outbound

This paper cites Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.618107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.618107Z digest=sha256:488ae3a99f68b330d07146608715930f16bd89a1e64b0851f77300bf7c02c8dd

Observation 1e079fd0-bf99-4dbb-9b5a-b28863b6d373 · outbound

This paper cites Davidsonian Scene Graph: Improving Reliability in Fine-grained Evaluation for Text-to-Image Generation.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Davidsonian Scene Graph: Improving Reliability in Fine-grained Evaluation for Text-to-Image Generation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.622457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.622457Z digest=sha256:295b73f1113856f2db2ac720853dc37fbd25dbcc9f2be7b2f6898bc6b783b375

Observation 102199fd-c156-441e-b7d0-d8e8ef9efe48 · outbound

This paper cites Emerging Properties in Unified Multimodal Pretraining.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Emerging Properties in Unified Multimodal Pretraining

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.626833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.626833Z digest=sha256:67de0922776add7ace79026c465477935e6a8e7d48bd8f5dcff57f11cee2e4b7

Observation 5aa4c4a5-d4bd-4dba-a56a-700e5605e989 · outbound

This paper cites GoT-R1: Unleashing Reasoning Capability of MLLM for Visual Generation with Reinforcement Learning.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark GoT-R1: Unleashing Reasoning Capability of MLLM for Visual Generation with Reinforcement Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.631140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.631140Z digest=sha256:0f8391845080d18c62ab3f2861fb628c36010985c2921df7c70f5e85d6be5ebb

Observation 9c2c9c9f-153d-4808-aad0-f3dc622ed302 · outbound

This paper cites Scaling rectified flow transformers for high-resolution image synthesis.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Scaling rectified flow transformers for high-resolution image synthesis

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.635989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.635989Z digest=sha256:32240a169e7b58caf85f351418279fbc56e69c24716c3c0bfed4f4792cf4d49c

Observation c3c5cd98-926e-4ae3-afbd-c637bb033a3c · outbound

This paper cites PUMA: Empowering Unified MLLM with Multi-granular Visual Generation.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark PUMA: Empowering Unified MLLM with Multi-granular Visual Generation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.640677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.640677Z digest=sha256:f66616e26bdcbeb4b59e625989f911b5965fbe219b3262939140d48d42454ed7

Observation 15feb0bb-9423-4260-a030-ccd50ba0b956 · outbound

This paper cites GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.645811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.645811Z digest=sha256:171d68cc12eeecbf670be5a69d010f7fdb26d118f7b352e86d83566380c8e5ae

Observation 0f6fadec-312e-44c2-b44d-623bec1bce2f · outbound

This paper cites Training-Free Structured Diffusion Guidance for Compositional Text-to-Image Synthesis.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Training-Free Structured Diffusion Guidance for Compositional Text-to-Image Synthesis

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.650851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.650851Z digest=sha256:bbe222f4c60e8f5a2da5c556feee7293926ae9ae19d34f4fd33167be10c9fbf7

Observation 66c32426-e674-488e-8e1d-aaad6cc0f225 · outbound

This paper cites Commonsense-T2I Challenge: Can Text-to-Image Generation Models Understand Commonsense?.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Commonsense-T2I Challenge: Can Text-to-Image Generation Models Understand Commonsense?

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.655845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.655845Z digest=sha256:58821cbddd1606869bfd7cff0f38c40ae32e75b64ea53c9f14bb24ee41a8d7d7

Observation 4a359ab1-f411-42fc-9429-91ab5da67e97 · outbound

This paper cites Datacomp: In search of the next generation of multimodal datasets.Advances in Neural Information Processing Systems, 36:27092–27112, 2023.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Datacomp: In search of the next generation of multimodal datasets.Advances in Neural Information Processing Systems, 36:27092–27112, 2023

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.661172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.661172Z digest=sha256:89b06961ad014b3916e55e9fb0d23cb35397878d28d3c3054c74833f434a2b78

Observation 22dd52d8-9ef1-4103-bfb2-6d994cfa619d · outbound

This paper cites Seedream 3.0 Technical Report.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Seedream 3.0 Technical Report

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.665647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.665647Z digest=sha256:61d50663f13ab58b8056bff250cc58251a69d0c4d39678f876ca2f6b23514776

Observation 1d147607-e9e8-4fc6-9e11-1252bb0c4383 · outbound

This paper cites Geneval: An object-focused framework for evaluating text-to-image alignment.Advances in Neural Information Processing Systems, 36:52132– 52152, 2023.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Geneval: An object-focused framework for evaluating text-to-image alignment.Advances in Neural Information Processing Systems, 36:52132– 52152, 2023

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.670397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.670397Z digest=sha256:8454cafb19f3c56fbb98b62f4cfef723b649f8c3d946c1e0a4939de15530af6f

Observation 20adb2b8-4782-4348-8e6b-348dc4ba1ab5 · outbound

This paper cites Seedream 2.0: A Native Chinese-English Bilingual Image Generation Foundation Model.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Seedream 2.0: A Native Chinese-English Bilingual Image Generation Foundation Model

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.674844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.674844Z digest=sha256:41ae14c1829b06993a3bab83492d602fbf3d6f1d40ec772291597f66cc8a7574

Observation a7103d71-2262-4a62-a77e-d9d3114b67ba · outbound

This paper cites Gemini2.5-pro, 2025.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Gemini2.5-pro, 2025

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.679408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.679408Z digest=sha256:c09cfa9b68426f2aa50c2f1c1314c1a1090c5c6b1d0d8a54205bd85caecc7215

Observation a4855d2b-abd7-4496-988e-e78d7d438bb5 · outbound

This paper cites Imagen4, 2025.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Imagen4, 2025

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.683844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.683844Z digest=sha256:c18e0472e051bb27bdb0bbf60af9f79b698adf3d4324acfa61ca127fe17c34ba

Observation 88f7b782-fd25-43b6-a00a-426d7ba02dfe · outbound

This paper cites Gemini2.5-flash-image, 2025.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Gemini2.5-flash-image, 2025

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.688106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.688106Z digest=sha256:68ad786490276c977b881b09121f1e0c0dee1b084376356fc929892dfe16a970

Observation f050f52b-e189-4d85-9826-0aa32acfc810 · outbound

This paper cites EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.692481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.692481Z digest=sha256:96ea2ec303008e397a487c2dadfb86730be73da0c515d14a5a46713fc4ae8d05

Observation 6f6cefc9-a127-45e4-9762-eedb1cfcf7f9 · outbound

This paper cites CLIPScore: A Reference-free Evaluation Metric for Image Captioning.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark CLIPScore: A Reference-free Evaluation Metric for Image Captioning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.697231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.697231Z digest=sha256:7abc976a58678a1b7bc85535fc113b53130eedad1d743495736d6109e7a80c4a

Observation d3674176-1a2c-4c13-a57e-88a56dbbaee2 · outbound

This paper cites Denoising diffusion probabilistic models.Advances in neural information processing systems, 33:6840–6851, 2020.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Denoising diffusion probabilistic models.Advances in neural information processing systems, 33:6840–6851, 2020

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.701850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.701850Z digest=sha256:4fd5d8f2fc27371b7f5c7c411a8a3c3adcc7ecf320a3d1d7500e02150dafbd0b

Observation b0fb7470-f858-49c1-a777-2b6f6d7da733 · outbound

This paper cites Scaling up vision-language pre-training for image captioning.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Scaling up vision-language pre-training for image captioning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.706247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.706247Z digest=sha256:518ff4402154663bd7c3fd509d37f7a6ce716d4bdadc5caf8817a4a7256c0bbc

Observation bb319ba8-5638-4b97-a767-a9b0cb625bb2 · outbound

This paper cites ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.710756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.710756Z digest=sha256:371ecd01323ead81681a6e881a4bac51320e96edc30e72a12718b08e71059c95

Observation d87c14be-6654-48a6-9516-ba005052566b · outbound

This paper cites Tifa: Accurate and interpretable text-to-image faithfulness evaluation with question answer- ing.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Tifa: Accurate and interpretable text-to-image faithfulness evaluation with question answer- ing

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.715302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.715302Z digest=sha256:2489b2fa9f51a64e75a85429fbb2f124718541437c205850ca060a2848ac361c

Observation c7fce712-9928-48a5-b394-133de9993800 · outbound

This paper cites T2i-compbench: A comprehensive benchmark for open-world compositional text-to-image generation.Advances in Neural Information Processing Systems, 36:78723–78747, 2023.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark T2i-compbench: A comprehensive benchmark for open-world compositional text-to-image generation.Advances in Neural Information Processing Systems, 36:78723–78747, 2023

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.719868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.719868Z digest=sha256:32d578d6f25ed4fa4b0439af94e81b12a16e3e443717f5e21c0209d78d2618bc

Observation 7944bef0-bd39-4828-a8ff-83ebf1886063 · outbound

This paper cites Scaling up visual and vision-language representation learning with noisy text supervision.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Scaling up visual and vision-language representation learning with noisy text supervision

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.724083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.724083Z digest=sha256:21463093cfefe1727f0248957c7fa2e03f4b873eef9a85b9ca639918ed42fc5f

Observation 19fe8c2c-f8bf-4602-a64b-228326cd7725 · outbound

This paper cites Genetic k-means algorithm.IEEE Transactions on Systems, Man, and Cybernetics, Part B (Cybernetics), 29(3):433–439, 1999.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Genetic k-means algorithm.IEEE Transactions on Systems, Man, and Cybernetics, Part B (Cybernetics), 29(3):433–439, 1999

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.728017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.728017Z digest=sha256:5fb7fbf5e3dd9e4dfeb5c9fcc1e9fe965a9931fcda2e08910b5d03e5ff93bc3a

Observation 817fe4d2-c078-4b26-b7d9-ae1f0aec9d92 · outbound

This paper cites GenAI-Bench: Evaluating and Improving Compositional Text-to-Visual Generation.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark GenAI-Bench: Evaluating and Improving Compositional Text-to-Visual Generation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.732027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.732027Z digest=sha256:e65f1f9effbdbc44bc07adde2f9a3b62d864c2af49a65953177ec45e1a1ad89e

Observation b6ce2eb6-4a67-476f-b610-2c536fc11828 · outbound

This paper cites Playground v2.5: Three Insights towards Enhancing Aesthetic Quality in Text-to-Image Generation.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Playground v2.5: Three Insights towards Enhancing Aesthetic Quality in Text-to-Image Generation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.736630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.736630Z digest=sha256:5e9db1784e45e26a82f69d5f618084746483ecb730e2f30b24ae239ae2eb6032

Observation b371269d-14db-4774-babd-8ae7f7a9625d · outbound

This paper cites Oscar: Object-semantics aligned pre-training for vision-language tasks.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Oscar: Object-semantics aligned pre-training for vision-language tasks

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.741275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.741275Z digest=sha256:e36301756ec18faa5f0ed1252b404cf868034df78d0bea560d77b5ac4d449ca3

Observation 0356551c-13d7-480c-8ac5-f129a797c09f · outbound

This paper cites Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.745641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.745641Z digest=sha256:9c6316f73cad08a5846618627b2a488226f8608a94acce3fba7cbc6298bb047c

Observation c24289ca-fdc0-419f-a700-62dc2927313b · outbound

This paper cites Nuwa-infinity: Autoregressive over autoregressive generation for infinite visual synthesis.Advances in Neural Information Processing Systems, 35:15420–15432, 2022.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Nuwa-infinity: Autoregressive over autoregressive generation for infinite visual synthesis.Advances in Neural Information Processing Systems, 35:15420–15432, 2022

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.749799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.749799Z digest=sha256:c30fe0293c70996cd23ee5f5164dda072e2c42158a22c702a20e2ba4572f2777

Observation 2157b36c-7885-484e-a4f1-3c621f24ba8b · outbound

This paper cites Evaluating text-to-visual generation with image-to-text generation.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Evaluating text-to-visual generation with image-to-text generation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.754532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.754532Z digest=sha256:360e731e44d5b0fe4ebf9c01dc3243b0571e46b55e7ee7f0ef462df74294cb49

Observation 8ca082a6-7773-467d-af2e-1bcf828cd6e2 · outbound

This paper cites Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.759187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.759187Z digest=sha256:b920c939d9908fa53c220a67d95e9cacab75be5ff86c411518098042d0e67591

Observation dda439af-19b4-429b-938c-c35248794e9a · outbound

This paper cites Summarizing emotions from text using plutchik’s wheel of emotions.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Summarizing emotions from text using plutchik’s wheel of emotions

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.764203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.764203Z digest=sha256:105ace93294250eca23826d8c037ef506b7ba4364ec110b284cd3d8419651613

Observation 3d5ed760-fcb6-457c-a127-e274c368c6a3 · outbound

This paper cites Gpt-4.1, 2025.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Gpt-4.1, 2025

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.768998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.768998Z digest=sha256:debb3faa5a683391ab78a45eba5fb01bab48ea9b16f92b5dc91c0a8119942f32

Observation 14637733-d07d-4f0d-a3a7-94d010d9e8fa · outbound

This paper cites Gpt-image-1, 2025.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Gpt-image-1, 2025

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.773709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.773709Z digest=sha256:527795c3fc9de43cdbc2a6ddbd171935fef3536482df1b7847e6088592be2bb1

Observation b8bc7723-71da-420b-afb6-57ca6ea7c95e · outbound

This paper cites Dall·e 3, September 2023.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Dall·e 3, September 2023

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.778454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.778454Z digest=sha256:eb7b0d8ecad8a322ac2637f15066bf11c5ebe332fdb5ef2694bb46d32343464e

Observation 20cb7434-a7c0-466f-a31e-d8119da7e8d8 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.782559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.782559Z digest=sha256:62f98535bd9b5e271f86a146cd12c27609277fac9fdf1c25bd241982249c56c9

Observation cbbeac40-3f01-459a-84ed-0493b3223119 · outbound

This paper cites Lumina-Image 2.0: A Unified and Efficient Image Generative Framework.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Lumina-Image 2.0: A Unified and Efficient Image Generative Framework

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.787032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.787032Z digest=sha256:4209f4dd26f4ad06b27b2c03d1796b37d0b8e01980b8ccd58f5b534bd301ea23

Observation e2299ed6-5290-44bf-9101-c6001ce104c8 · outbound

This paper cites High- resolution image synthesis with latent diffusion models.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark High- resolution image synthesis with latent diffusion models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.791588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.791588Z digest=sha256:a14bd4d704f67e0407829c935f0e1ec9b5872916d0c26b711a6cb0420ee19311

Observation 00b83a60-5eb1-4ff5-8eb9-0d3a3afa232f · outbound

This paper cites Laion-5b: An open large-scale dataset for training next generation image-text models.Advances in neural information processing systems, 35:25278–25294, 2022.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Laion-5b: An open large-scale dataset for training next generation image-text models.Advances in neural information processing systems, 35:25278–25294, 2022

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.795622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.795622Z digest=sha256:f68a962bed46e9a8b788803e944614f21d27cf485e0a5a6cbbf01b07b903efd6

Observation bb70cb03-f76e-4543-bde6-2c52f6442a69 · outbound

This paper cites Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.799595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.799595Z digest=sha256:087918122464cc46a1cba6b384b1e1a781ca657c79b10d10f12e36c9241b05b3

Observation a40919a6-6724-4fa7-b093-3f5f9d14a557 · outbound

This paper cites From Pixels to Prose: A Large Dataset of Dense Image Captions.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark From Pixels to Prose: A Large Dataset of Dense Image Captions

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.803853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.803853Z digest=sha256:b3ea6a8f50c4f8339b3ccdd80a56fbc39cb4c8ab84bb09d4834e8f4ba6b2a233

Observation 767ffcd6-8a63-472b-838c-284ae1be0a57 · outbound

This paper cites Stable diffusion 2.1, 2022.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Stable diffusion 2.1, 2022

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.808078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.808078Z digest=sha256:22a0856f84c3d99bd530a3e84d593f5ef3bd8c0fc952a8287bf33d9fa271efff

Observation c8ce336c-b8b7-4bfa-a370-63d1d229dd7f · outbound

This paper cites Stable diffusion 3, 2024.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Stable diffusion 3, 2024

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.812209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.812209Z digest=sha256:801a669b207eb6a7c5ed23092276a5682042467de9f7a71ae561bede6ad00240

Observation 6437c0c5-32ed-4980-8808-28a8ca0458c8 · outbound

This paper cites Stable diffusion 3.5, 2024.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Stable diffusion 3.5, 2024

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.816335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.816335Z digest=sha256:38470c90c0f6042090ab81af28ceb900aa172e7fdbc0f36a6241770ee7272093

Observation 05a59200-9ff6-43f0-b1f9-ba3cd7497707 · outbound

This paper cites T2I-ReasonBench: Benchmarking Reasoning-Informed Text-to-Image Generation.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark T2I-ReasonBench: Benchmarking Reasoning-Informed Text-to-Image Generation

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.820245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.820245Z digest=sha256:39dc18230755a747f92dc7d0ca9ed9b7578ae6c6db99d6134eac7b9e0d31e4a6

Observation 9d54b302-4927-4779-9b8d-1947f2a21fa7 · outbound

This paper cites Journeydb: A benchmark for generative image understanding.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Journeydb: A benchmark for generative image understanding

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.824740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.824740Z digest=sha256:55263d1b45d005f7444634c0f31f52052e21282bd67e7d67eb4028f22ea3b069

Observation d8dfc13f-ab11-4415-b4fb-afc6ddb6ecd0 · outbound

This paper cites AnyText: Multilingual Visual Text Generation And Editing.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark AnyText: Multilingual Visual Text Generation And Editing

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.828806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.828806Z digest=sha256:71285fe407061d4462bf3ce73dfcce3699368a68da294a0b55ec838e1c243c11

Observation 9314c1d1-ec1f-4db1-867a-30d8c7dfebc2 · outbound

This paper cites Textatlas5m: A large-scale dataset for dense text image generation.arXiv preprint arXiv:2502.07870, 2025.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Textatlas5m: A large-scale dataset for dense text image generation.arXiv preprint arXiv:2502.07870, 2025

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.832885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.832885Z digest=sha256:9ef8178b7f5532809da7b30519958d21e738a1bb762dc81142b6eccbf4bfa0d0

Observation 7a952e4b-07f4-4ffa-92fe-5179ad069cae · outbound

This paper cites Nuwa: Visual synthesis pre-training for neural visual world creation.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Nuwa: Visual synthesis pre-training for neural visual world creation

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.836817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.836817Z digest=sha256:1afe5fe4fa446b47fac64397512a55f944913af593aa02c234c6cbefa8752bc1

Observation 7d331a1d-46c9-4d4d-b0a4-e012fdae3b65 · outbound

This paper cites Qwen-Image Technical Report.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Qwen-Image Technical Report

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.840789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.840789Z digest=sha256:0cc4bed884a142bdd7369021f6b6b82bfee36a498fcc5d6850e1626ae1e33111

Observation 41db8f31-f4b4-4a70-b193-2734420b0857 · outbound

This paper cites OmniGen2: Towards Instruction-Aligned Multimodal Generation.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.845193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.845193Z digest=sha256:1fd92bb1643cfcca8b7fbc61fa54365f594d95eac71abd5bccc39f7a1855b46a

Observation 938a92ab-78fd-47e7-a096-b670c3016288 · outbound

This paper cites Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.849652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.849652Z digest=sha256:c52b2a84b07bf31a8d7aa935f4b86b19ed6b54d5148fca898537b33be20a776f

Observation d8d8bce3-6b48-4a3e-9891-050a87ba4cf2 · outbound

This paper cites Conceptmix: A com- positional image generation benchmark with controllable difficulty.Advances in Neural Information Processing Systems, 37:86004–86047, 2024.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Conceptmix: A com- positional image generation benchmark with controllable difficulty.Advances in Neural Information Processing Systems, 37:86004–86047, 2024

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.854018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.854018Z digest=sha256:7a50d7a5b526adfea7dfa8404ba39318341212a78b04c5f0d3da74f894456e16

Observation 29c414a0-efcf-470c-a956-7d6eb10b54d1 · outbound

This paper cites SANA 1.5: Efficient Scaling of Training-Time and Inference-Time Compute in Linear Diffusion Transformer.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark SANA 1.5: Efficient Scaling of Training-Time and Inference-Time Compute in Linear Diffusion Transformer

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.858574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.858574Z digest=sha256:40d90130e0ee0a3834f8ebef99d9ab8bbaab786353408c707375ff299b731a89

Observation e6ece735-e814-416a-ba27-7dc34e34b37f · outbound

This paper cites Show-o2: Improved Native Unified Multimodal Models.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Show-o2: Improved Native Unified Multimodal Models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.862906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.862906Z digest=sha256:91bc9e5da62051cfa1a706154d6b50e474f38d2662d983cb5d9dc3a7504d1669

Observation db94fd1b-4d1a-4e33-80da-4069954ad002 · outbound

This paper cites Qwen3 Technical Report.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Qwen3 Technical Report

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.867345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.867345Z digest=sha256:a3a58a0797cf85e506223ba08d757629fb450da4cd605a843a1f91151895ab21

Observation d6030455-2072-4d5c-8360-fa3b2aee2280 · outbound

This paper cites Scaling Autoregressive Models for Content-Rich Text-to-Image Generation.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.872017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.872017Z digest=sha256:0eeaea10135c8279680aad423a0e5a48e9416ce6d05b7656dbeac8c4f8dfdfaa

Observation 5216cd83-fb07-494a-a7c0-8828e1a8916c · outbound

This paper cites Capsfusion: Rethinking image-text data at scale.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Capsfusion: Rethinking image-text data at scale

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.876743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.876743Z digest=sha256:f2474b60f684a9225d56ca93105ec534ddf37ce9dbefd05ceb2c328af798459c

Observation 700b9ab6-fd88-4774-99c9-fa3b8739461a · outbound

This paper cites VinVL: Revisiting Visual Representations in Vision-Language Models.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark VinVL: Revisiting Visual Representations in Vision-Language Models

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.880972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.880972Z digest=sha256:af02d6c0d8d7bbb7fa2c364b6dbcace69d2d5b3ae60379f4efd92d312eb0e999

Observation 8da24999-f075-454c-88ba-faedb3fde794 · outbound

This paper cites Lumina-next: Making lumina-t2x stronger and faster with next-dit.Advances in Neural Information Processing Systems, 37:131278–131315, 2024.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Lumina-next: Making lumina-t2x stronger and faster with next-dit.Advances in Neural Information Processing Systems, 37:131278–131315, 2024

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.885237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.885237Z digest=sha256:0556bec745ad44e6f682581793744724ed0887515f34b218964b2e3c7676810a

Pith citing papers

Observation a2d368cc-6e46-4376-b939-0cb33154a5af · inbound

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer cites this paper.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:08:37.030219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-11T14:08:36.801359Z digest=sha256:07b9f704d92893e3a0a4fdeb9a87b86d0007a346b93bf402ed3f5b38d6996b5f

Observation 6b6f9400-4ee2-466a-9de1-97b015cb413c · inbound

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer cites this paper.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:31.327318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:31.327318Z digest=sha256:3253bd0ae05b3822afdec5badee7fd4ce6e7272519e58d50e5ac883723d02853

Observation fce86933-3f87-4417-89e9-eac9301d2ba8 · inbound

Guiding Token-Sparse Diffusion Models cites this paper.

Guiding Token-Sparse Diffusion Models FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T12:50:55.743597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:50:55.743597Z digest=sha256:7d3058eeef588f1129d1272f7ca52b6e607a5268d9c8fce35f3cf5260669d23a

Observation 3dccef82-3cf6-4e2a-b349-6088b0cc0d1d · inbound

Self-Adversarial One Step Generation via Condition Shifting cites this paper.

Self-Adversarial One Step Generation via Condition Shifting FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:11:01.970606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T15:37:36.994169Z digest=sha256:ef6a50d05ec9a2f1dabb82f1a03be9e0c10053c13ef9a21e98711a37ee725ec6

Observation d77b53bb-b502-4f90-a876-656d0aa4aa6a · inbound

LLaDA2.0-Uni: Unifying Multimodal Understanding and Generation with Diffusion Large Language Model cites this paper.

LLaDA2.0-Uni: Unifying Multimodal Understanding and Generation with Diffusion Large Language Model FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-10T00:49:48.344114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T00:49:38.156237Z digest=sha256:69b502f4d1678c7de198ced9b05e2a902a0736fbd5bdfb736612499a07fa7c81

Observation df3ef91b-bdce-4eab-b933-b60bd3102b55 · inbound

InsightTok: Improving Text and Face Fidelity in Discrete Tokenization for Autoregressive Image Generation cites this paper.

InsightTok: Improving Text and Face Fidelity in Discrete Tokenization for Autoregressive Image Generation FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:13:30.594327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T02:10:18.004724Z digest=sha256:64862dbc277af76b59b1e67e9584b4ca8172c7c7269acff2a633ec65009bd227

Observation fe9d9296-2407-4e88-90a3-9ca455a86712 · inbound

Unlocking Complex Visual Generation via Closed-Loop Verified Reasoning cites this paper.

Unlocking Complex Visual Generation via Closed-Loop Verified Reasoning FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-19T16:37:39.647701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T16:35:43.166697Z digest=sha256:cbd8fcebb3e2cde7fafc9fc0ad1ce33c24877bea4e6936ae188e940a25aec920

Observation 6850081c-98f7-4258-831d-423ea5956fb1 · inbound

SciIR: A Large-scale Training Dataset and Benchmark for Scientific Image Reasoning Generation cites this paper.

SciIR: A Large-scale Training Dataset and Benchmark for Scientific Image Reasoning Generation FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T06:54:20.704034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T06:48:22.594002Z digest=sha256:82434259dec7ee7e0dfe02656433ff036d935a89f766adf23a82a105e8df38bd

Observation 0b8bbb34-8a38-4ddd-9730-8cd48d5abea2 · inbound

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric cites this paper.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-01T15:42:58.828756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:42:58.828756Z digest=sha256:c28f58a6677568dcb2fab5a54bd74104103089d2b8a31a9aeaedcf90b925a63c

Observation a87f0171-fe34-42dc-86dc-ead1f635dfe0 · inbound

CAPEval: A Decoupled Caption Evaluation across Understanding and Generation cites this paper.

CAPEval: A Decoupled Caption Evaluation across Understanding and Generation FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T04:09:49.613004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:09:49.613004Z digest=sha256:3246d0936a70a7c32579c2c0679e49af8e99d07c5365d72629bf89762a5ed513