Pith. sign in

Paper Citation Record · LEDGER

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation

As of 20 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 10 inbound Pith citation observations for arXiv:2412.18150.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.18150 v2

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T05:02:21.265519Z

measured 62 of 62 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:40:26.460713Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-22T09:04:45.978885Z

Reference resolution

52 of 52 outbound references displayed

  • verified exact0
  • verified fuzzy31
  • unresolved21
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 307de585-4f03-4f32-96f1-07fbf6d6be76 · outbound

This paper cites GPT-4 Technical Report.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T05:02:20.851688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:02:20.851688Z digest=sha256:2041f0092248a75ae38be9d44da6a1a4074f78cab1177316bc86c44707a873c0

Observation ad95ae33-07e9-4e5f-90e2-3a866e39fc66 · outbound

This paper cites Kandinsky 3.0 Technical Report.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation Kandinsky 3.0 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T05:02:20.857101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:02:20.857101Z digest=sha256:fab607e8c61797a29e5d5edc8bc79aa95c40a5fc52b4ea7e432f41366f2b0b10

Observation fbda3c3d-9c86-458f-90b2-542372b9899f · outbound

This paper cites an unresolved cited work.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-11T05:02:22.745634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T05:02:20.862210Z digest=sha256:e78697ed5eb8b4f6f96ea08a81e557ba22d6a9bf2fc6d5709597d8b4f7a109c3

Observation 76909cde-a0cb-4a6f-ba78-e441b7868eff · outbound

This paper cites PixArt-\Sigma: Weak-to-Strong Training of Diffusion Transformer for 4K Text-to-Image Generation.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation PixArt-\Sigma: Weak-to-Strong Training of Diffusion Transformer for 4K Text-to-Image Generation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T05:02:20.871423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:02:20.871423Z digest=sha256:09ec0d0effd4e60108f72e6e81c3651909b0aede6b676fadb94ec6eb58b41191

Observation 7b580605-deab-43ad-8b41-8a96588eb7a5 · outbound

This paper cites PIXART-{\delta}: Fast and Controllable Image Generation with Latent Consistency Models.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation PIXART-{\delta}: Fast and Controllable Image Generation with Latent Consistency Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T05:02:20.879214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:02:20.879214Z digest=sha256:77fe852b574b6e2fc43fc9252bb430cc614afb1ab5fb0ce5f73ef94a65c91c21

Observation cb1c4a1b-9fee-49e9-9f1c-cffd3c41a872 · outbound

This paper cites Pixart-α: Fast training of diffusion trans- former for photorealistic text-to-image synthesis.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation Pixart-α: Fast training of diffusion trans- former for photorealistic text-to-image synthesis

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:02:22.718503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T05:02:20.886875Z digest=sha256:24024c210ccf1d5e3960f755a56e3e45296dd1cb702c4228e03d64fac072797f

Observation feea6f55-9308-4406-b65d-a3da77087c31 · outbound

This paper cites Dall-eval: Probing the reasoning skills and social biases of text-to- image generation models.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation Dall-eval: Probing the reasoning skills and social biases of text-to- image generation models

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:02:22.694638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T05:02:20.894200Z digest=sha256:057f1fa12a6dcfc8d3e09e8d690c298673a4f58ccdd48aa3efa4a11af0930d76

Observation c4652b9c-b24b-40c6-b5a5-2ff49340a79d · outbound

This paper cites If-i-xl-v1.0.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation If-i-xl-v1.0

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:02:22.675019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T05:02:20.904653Z digest=sha256:82b86961791b18426bfa221e8e26631c6973467f37078c1dd55182c5118661e4

Observation 8602313f-eb29-4ee1-9494-e29fc6325577 · outbound

This paper cites Dreamina.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation Dreamina

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:02:22.657134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T05:02:20.912278Z digest=sha256:906e929b0899e157f2220df5e04326976a506bcd8ec29c7a9c50118d31e4bccb

Observation 8c2fe1ef-eeb8-4120-a6e2-afbf475a170a · outbound

This paper cites Scaling recti- fied flow transformers for high-resolution image synthesis.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation Scaling recti- fied flow transformers for high-resolution image synthesis

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:02:22.637201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T05:02:20.918953Z digest=sha256:e0a63085310af118ca8d9cf9b6a3f83477a6707922afbf677001d70285d7f605

Observation 3da1b9f5-41a1-4ca2-9ce8-4e57d21880be · outbound

This paper cites Progressive Knowledge Distillation Of Stable Diffusion XL Using Layer Level Loss.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation Progressive Knowledge Distillation Of Stable Diffusion XL Using Layer Level Loss

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T05:02:20.925998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:02:20.925998Z digest=sha256:5ad29c534b474c7564fc20adde16352a6d2a30d142406be465b01d18a8070513

Observation 3e78fa1f-f29e-411d-8a9e-907b01b74f5f · outbound

This paper cites CLIPScore: A Reference-free Evaluation Metric for Image Captioning.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation CLIPScore: A Reference-free Evaluation Metric for Image Captioning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T05:02:20.937773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:02:20.937773Z digest=sha256:659ac8dae1dc9c5c0e5b4b0c692e1c3bf78e0fd767c65918a769b947e1763945

Observation 455001f5-1923-493b-adae-2ab67bd2dda2 · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilib- rium.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation Gans trained by a two time-scale update rule converge to a local nash equilib- rium

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:02:22.607738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T05:02:20.948276Z digest=sha256:dfb8c66589b576540f6842b2b35098244f10f86eb82edfa50b766f8e6ef004a3

Observation 55b022e9-6f76-4d76-a03b-436436feaf0e · outbound

This paper cites Denoising diffu- sion probabilistic models.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation Denoising diffu- sion probabilistic models

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:02:22.588710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T05:02:20.959363Z digest=sha256:6cac2bac540f1822ac7b0e924286ddfaff6392c464c0347983a4793a6c00dc20

Observation 2523502c-afa7-4a2d-91c2-7d540c0db447 · outbound

This paper cites Midjourney.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation Midjourney

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:02:22.564578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T05:02:20.965988Z digest=sha256:6084b353dcf6dc979018e27ef8fd810363211e6e1da0dd2da2e3bfb283769e1e

Observation 7037f01e-36fa-4e62-b6b6-918a8050cb37 · outbound

This paper cites Tifa: Accurate and interpretable text-to-image faithfulness evaluation with question answering.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation Tifa: Accurate and interpretable text-to-image faithfulness evaluation with question answering

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:02:22.538887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T05:02:20.972653Z digest=sha256:a2d8beea1ac448a92bb6cbc766506087b410e0bac5cf25db41c1a1fe71b756c5

Observation a98bdefe-d7df-4fb8-9920-c50cde40387a · outbound

This paper cites T2i-compbench: A comprehensive benchmark for open-world compositional text-to-image generation.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation T2i-compbench: A comprehensive benchmark for open-world compositional text-to-image generation

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:02:22.497299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T05:02:20.978469Z digest=sha256:6a5fa56debcac2b86d8042af54f1dba50f15ac444413fa483465a0afdfbb5993

Observation 2600e025-9e03-4bb4-bc86-c62ce442f7d9 · outbound

This paper cites Bert: Pre-training of deep bidirectional trans- formers for language understanding.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation Bert: Pre-training of deep bidirectional trans- formers for language understanding

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:02:22.463105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T05:02:20.986792Z digest=sha256:ab14ebdc1ee192a74a3f497681cab89690eeea98001df76e92f2259d6630bb99

Observation 23a18070-6271-4b03-87ec-cd4747603d55 · outbound

This paper cites Pick-a-pic: An open dataset of user preferences for text-to-image generation.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation Pick-a-pic: An open dataset of user preferences for text-to-image generation

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:02:22.437220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T05:02:20.992411Z digest=sha256:b872d7e14f7dbf7947b216bf05a23f662d711bf1acf2151a4a831c11014aa924

Observation cb6b0374-f776-4503-8711-0803822a4cb4 · outbound

This paper cites GenAI-Bench: Evaluating and Improving Compositional Text-to-Visual Generation.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation GenAI-Bench: Evaluating and Improving Compositional Text-to-Visual Generation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T05:02:20.998785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:02:20.998785Z digest=sha256:2fb4cf5b7b37814bac0072c18ea5c07629fd229f19a66acc33cd0955131e9fc4

Observation be980eac-6c0c-4333-92a6-fdd6b999f4bc · outbound

This paper cites Evaluating and improving composi- tional text-to-visual generation.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation Evaluating and improving composi- tional text-to-visual generation

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:02:22.401777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T05:02:21.005276Z digest=sha256:ec6ab44885fbc5ebef2d228b805d752fe2f17fff505fce4a0a1540aad8f47226

Observation f89b545d-0bd5-4bf8-ac60-28d90d10c2d7 · outbound

This paper cites Playground v2.5: Three Insights towards Enhancing Aesthetic Quality in Text-to-Image Generation.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation Playground v2.5: Three Insights towards Enhancing Aesthetic Quality in Text-to-Image Generation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T05:02:21.012013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:02:21.012013Z digest=sha256:d452d294a1505742b4adb1aa727609aa6710b917e3a46a351227c6c2b8d24eb3

Observation 4d610c48-2b54-43ab-8d34-229cab943032 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T05:02:21.021586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:02:21.021586Z digest=sha256:669abc64c4056a0b83f24082ac6b84931892e27063f40b742363d639410a304e

Observation 6decca13-abc0-427c-a26b-d644163b295b · outbound

This paper cites Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T05:02:21.029310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:02:21.029310Z digest=sha256:9cc3933d20ab22e4a5a37e800648125d03102a1c7ec95aebccb6c944e2e85d29

Observation 5cb528c2-36a5-4668-866e-1f1b969f270d · outbound

This paper cites Rich human feedback for text-to-image generation.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation Rich human feedback for text-to-image generation

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:02:22.356693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T05:02:21.036558Z digest=sha256:e053d71a1cf8f42e028d4e62590cded83ceb77f605c3f5053ee3e032bf53ef84

Observation 47dd1d66-3a5d-4564-83eb-8f86b908a486 · outbound

This paper cites SDXL-Lightning: Progressive Adversarial Diffusion Distillation.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation SDXL-Lightning: Progressive Adversarial Diffusion Distillation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T05:02:21.042923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:02:21.042923Z digest=sha256:e05a87028009c4160cc6891095dc672c038be770b56a4b9048087eefd1580b37

Observation 93f56d51-6c78-40f0-a08f-3f7866d4b12b · outbound

This paper cites Microsoft coco: Common objects in context.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation Microsoft coco: Common objects in context

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T05:02:21.051255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:02:21.051255Z digest=sha256:016c016676cee83fe75de55a6b27e5c3d752a9458de39b4220973e00f33ce209

Observation fe456f20-c1d2-4ce5-9daa-aacb5dbf0b22 · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:02:22.311721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T05:02:21.059585Z digest=sha256:8b0e2e82376aa749a436811f207bb444737d767050b008a2c8e17e845649ab3e

Observation bbf407be-a7e0-4ae2-a39d-3e819f06ae8e · outbound

This paper cites Latent Consistency Models: Synthesizing High-Resolution Images with Few-Step Inference.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation Latent Consistency Models: Synthesizing High-Resolution Images with Few-Step Inference

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T05:02:21.071349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:02:21.071349Z digest=sha256:b63c9adbe5a74e3e983b7a7f0ced8fab7ecec4da30232c1c3308da03d3cb4342

Observation cc432f8e-9bcb-4ad5-8dc5-06bbb44ffefd · outbound

This paper cites Scalable diffusion models with transformers.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation Scalable diffusion models with transformers

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:02:22.280148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T05:02:21.079970Z digest=sha256:a0ccd98ed867a94e6b0809ebb634b41b832f4a51465fc9f82dcc32b4eda9c970

Observation 594a4e38-b197-4553-95b2-511c39c83f54 · outbound

This paper cites Sdxl: Improving latent diffusion models for high-resolution image synthesis.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation Sdxl: Improving latent diffusion models for high-resolution image synthesis

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:02:22.249478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T05:02:21.089847Z digest=sha256:851bd01e5d41acc6ff4c26e0a76291f30e02b3362e530b86feadbeee525fe12e

Observation 8eedc475-0ccb-406e-8f84-98bbc525c32a · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T05:02:21.097690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:02:21.097690Z digest=sha256:42db1189f2b28daa63efa90db292b01e96dbe48c9513895c58b9fff4f2e394ba

Observation 1e5849c8-95e4-48d9-bdec-b36e2f2386d4 · outbound

This paper cites High-resolution image syn- thesis with latent diffusion models.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation High-resolution image syn- thesis with latent diffusion models

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:02:22.221393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T05:02:21.103787Z digest=sha256:21cf9a3104630c2d13a638555f15793e3ad2fd364e87d8eb7432d2f8ed28efb8

Observation 7f5035af-9229-42c9-b7b7-3e740d5f785c · outbound

This paper cites Photorealistic text-to-image diffusion models with deep language understanding.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation Photorealistic text-to-image diffusion models with deep language understanding

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:02:22.188309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T05:02:21.111622Z digest=sha256:ad12517f75a4050c47ed87f1cb4dbe942be49b4cb16cd0e84d5bdfc80c569183

Observation d49c2c9b-c03a-409a-910a-1e5532028457 · outbound

This paper cites Improved techniques for training gans.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation Improved techniques for training gans

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:02:22.157587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T05:02:21.117927Z digest=sha256:71cc72bde9e39a3a0ab710799a6eb9f73f56b53bb39718ecb0f04451c7eef950

Observation dcd8729b-d21b-4149-8598-4bab0acc0422 · outbound

This paper cites Adversarial diffusion distillation.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation Adversarial diffusion distillation

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:02:22.125556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T05:02:21.126537Z digest=sha256:60b116c1e77f0090e8b769d869dcb137c1dc9d79c5d35dbdd9b90b45e020cb72

Observation 2e0a4791-8562-48ad-8c09-f65c5e542251 · outbound

This paper cites EVALALIGN: Supervised Fine-Tuning Multimodal LLMs with Human-Aligned Data for Evaluating Text-to-Image Models.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation EVALALIGN: Supervised Fine-Tuning Multimodal LLMs with Human-Aligned Data for Evaluating Text-to-Image Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T05:02:21.134390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:02:21.134390Z digest=sha256:490a9199d6a42e86f33d0f780b61f6e97399a601c15f8323b89f3f44ca351a14

Observation ec67be34-20d5-4d73-8b90-7687ac6dd39c · outbound

This paper cites Kolors: Effective training of diffusion model for photorealistic text-to-image synthesis.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation Kolors: Effective training of diffusion model for photorealistic text-to-image synthesis

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T05:02:21.146494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:02:21.146494Z digest=sha256:bc960a88e83ff88f64085e8577aed322f4222e61ac89666da6cc4bc52227557b

Observation 5611f23b-ea4e-4db1-b371-7ab0b513843e · outbound

This paper cites Shaping datasets: Optimal data selection for spe- cific target distributions across dimensions.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation Shaping datasets: Optimal data selection for spe- cific target distributions across dimensions

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:02:22.090609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T05:02:21.155621Z digest=sha256:aca19b580179e46ba07395760aed60e0819f6364933e3d81bd8cfaa4d8b7070a

Observation d903efde-b1a8-4d39-83ca-20655a7eb158 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T05:02:21.167608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:02:21.167608Z digest=sha256:ba3cd451c4d7e123b28f57f9d86cf9940c43ebe60e9fce643284cc7534f5f5d1

Observation 824b497c-b825-43cc-8ecd-15fad71e2ded · outbound

This paper cites Diffu- siondb: A large-scale prompt gallery dataset for text-to- image generative models.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation Diffu- siondb: A large-scale prompt gallery dataset for text-to- image generative models

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:02:22.056488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T05:02:21.177037Z digest=sha256:e5f076e5c5eca819bcabb8d3e068fb15cf5a1c938be0d8a2f9d11aeef1409726

Observation 8b36e445-d6cc-49f1-9bc1-a2fdc6827740 · outbound

This paper cites Revisiting Text-to-Image Evaluation with Gecko: On Metrics, Prompts, and Human Ratings.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation Revisiting Text-to-Image Evaluation with Gecko: On Metrics, Prompts, and Human Ratings

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T05:02:21.183802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:02:21.183802Z digest=sha256:bfa9c52467a32143a40148fafb0e63a933e2a9014d5513c3230a294f974d13bd

Observation e73ef5b9-5d03-40f0-9ce5-f3d0c7663fa5 · outbound

This paper cites Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T05:02:21.191344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:02:21.191344Z digest=sha256:a2df4df54066c738a6f40a3966d981ba39cedd10c7aa6410977b24bec1bfe86f

Observation fce54b42-52b2-4ef2-a7f1-a8f8d5e8fc25 · outbound

This paper cites Imagere- ward: Learning and evaluating human preferences for text- to-image generation.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation Imagere- ward: Learning and evaluating human preferences for text- to-image generation

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:02:22.003846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T05:02:21.197982Z digest=sha256:9f1e738ebb7d6dd97095941de64ec8734a24fccd78d42f86eacf659a636eef28

Observation 0668094b-562a-4122-8821-04ba1d075eac · outbound

This paper cites What you see is what you read? improving text- image alignment evaluation.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation What you see is what you read? improving text- image alignment evaluation

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:02:21.970675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T05:02:21.206424Z digest=sha256:6bcebd9eeda29c14131433071ba5264e22c9301d7cd3c2cc02e4d469a24e71e1

Observation b767f916-d5ab-4753-b078-ea09397f85ca · outbound

This paper cites mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T05:02:21.213390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:02:21.213390Z digest=sha256:7cc8eb9a8a8941e7970a8ef14ea841d8507f1c2e4d114c46fd2e2a2990dca8de

Observation 15f99ce7-d24c-4a56-b0cb-d1f90b984ea8 · outbound

This paper cites Scaling autoregressive models for content-rich text-to-image generation.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation Scaling autoregressive models for content-rich text-to-image generation

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:02:21.947283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T05:02:21.230479Z digest=sha256:652242d3774d2ee955502576fc0b02696b055b46157d49d820b304141483807d

Observation 080476d3-f4b0-4183-bb0f-b3ee56730743 · outbound

This paper cites The unreasonable effectiveness of deep features as a perceptual metric.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation The unreasonable effectiveness of deep features as a perceptual metric

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:02:21.921018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T05:02:21.237859Z digest=sha256:d97f99dcc7262a7c173929ac0fc683c7b69fe166c82d9930f1732cefba8d0916

Observation d590b938-a58a-41b9-bf4c-abc96f605fdb · outbound

This paper cites In data collection details, we de- tail the classification and sampling of real user prompts in Sections 8.1 and 8.2, ensuring diversity and balance among the real prompts.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation In data collection details, we de- tail the classification and sampling of real user prompts in Sections 8.1 and 8.2, ensuring diversity and balance among the real prompts

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:02:21.894425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T05:02:21.243077Z digest=sha256:425977b3a04ad92d67933623937ba2c85d0ca9df90b57b59936a35724e85fd0e

Observation 5eb7cd80-4821-41c8-bbe2-19b2ca55d1fa · outbound

This paper cites 1 cat and some dogs.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation 1 cat and some dogs

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:02:21.865554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T05:02:21.251563Z digest=sha256:c9c7fd585f590d9e7a56d182fa394521500684f5f7a7bed190985251bfd36ab8

Observation 7a7ddfa5-7fda-4b6e-b168-15ffacdffd25 · outbound

This paper cites Yes.” and “No.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation Yes.” and “No

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:02:21.841468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T05:02:21.258732Z digest=sha256:ebd7578099d086d91f49c9127e9298d11b98b4aa1eef3b19d042fae99fb20f63

Observation f9334116-0c08-4333-94c8-1cfcc72905b6 · outbound

This paper cites Evaluation of image-text alignment across different T2I models.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation Evaluation of image-text alignment across different T2I models

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:02:21.821794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T05:02:21.265519Z digest=sha256:d35cea5b3789628d8c8724e9bb7a5f14041bcf096f47e15ed0b2cd7d6a150fdf

Pith citing papers

Observation 332d43b9-40a9-4b4d-9326-6a9fb54f5748 · inbound

Unified Reward Model for Multimodal Understanding and Generation cites this paper.

Unified Reward Model for Multimodal Understanding and Generation EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T00:44:30.874520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T00:44:30.558048Z digest=sha256:7f2c0be3eae013463ca04a516d94add8904c3f4788399a15680dcd64ba087ec7

Observation 875f6439-a1f1-4702-ac45-987eb212db60 · inbound

Seedream 2.0: A Native Chinese-English Bilingual Image Generation Foundation Model cites this paper.

Seedream 2.0: A Native Chinese-English Bilingual Image Generation Foundation Model EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T08:27:36.292569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-17T08:27:36.242416Z digest=sha256:c7dcf41955a1016758a13d54a09d5b8e86b2166dabe64cb12f04c459c44e49e7

Observation 24604afa-ef4f-44c1-9f38-99a20c6f70e7 · inbound

Seedream 3.0 Technical Report cites this paper.

Seedream 3.0 Technical Report EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T07:55:38.744353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-13T07:55:38.690569Z digest=sha256:e2636a2ec8363a1523bd1d0252f1dc21bf07cf9f218e0c690f18488133e5e236

Observation 73ddecb0-e4a0-4d62-b9b0-dec03dc27092 · inbound

Instruction-augmented Multimodal Alignment for Image-Text and Element Matching cites this paper.

Instruction-augmented Multimodal Alignment for Image-Text and Element Matching EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T12:40:26.460713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:40:26.460713Z digest=sha256:23c7341baea9771606e7bd17901a592d7328d87094042b5fa5515360a323443e

Observation 0717342c-df3f-486f-b6ef-76aa03d125a6 · inbound

LLM Code Customization with Visual Results: A Benchmark on TikZ cites this paper.

LLM Code Customization with Visual Results: A Benchmark on TikZ EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T23:37:38.074067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:37:38.074067Z digest=sha256:3b498358179d96ad03fc706309441a5e86f40e8ab90bb224b7a3ab0e883f5c83

Observation 6c5980de-0401-4a87-bc22-9002791225ac · inbound

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation cites this paper.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:46.138692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:46.138692Z digest=sha256:153b258f1f4d0f496fe4a0b78d506efcf49fccad5689a88f3e22784c66ee96a8

Observation 20e1c83d-9c5e-4fbe-8e3d-fccc0850c1db · inbound

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models cites this paper.

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:33.225548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:33.225548Z digest=sha256:6879413e3fe4c0fafb7a215373a29fcc5f9179e7d3e84d2ca4bcb54d93eb30a5

Observation 0a48ac7b-7546-4399-aed5-fbcb6dd31e9c · inbound

ViDA-UGC: Detailed Image Quality Analysis via Visual Distortion Assessment for UGC Images cites this paper.

ViDA-UGC: Detailed Image Quality Analysis via Visual Distortion Assessment for UGC Images EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T17:27:52.889724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:27:52.889724Z digest=sha256:111a2e05c37342a74a430d60d4c15e0f2644587c2542b2436648f835687d529a

Observation f050f52b-e189-4d85-9826-0aa32acfc810 · inbound

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark cites this paper.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.692481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.692481Z digest=sha256:151208469716ac81bac170b4844422664dc0b63a14e7b47440429bda998ba23f

Observation de3d8133-aed9-47b5-abd9-3b631ea87206 · inbound

BEiTScore: Reference-free Image Captioning Evaluation with an Efficient Cross-Encoder Model cites this paper.

BEiTScore: Reference-free Image Captioning Evaluation with an Efficient Cross-Encoder Model EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:04:45.982530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T09:01:24.453821Z digest=sha256:6e90ab3df407231e9d0c07f3805a38bd7975706df40192bfa21488d1387bbe26