EvalMuse-40K adds 40K image-text pairs and about 1M human annotations to T2I evaluation, and the FGA-BLIP2 and PN-VQA metrics show higher correlation with human alignment scores than prior zero-shot metrics.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.CV 1years
2024 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation
EvalMuse-40K adds 40K image-text pairs and about 1M human annotations to T2I evaluation, and the FGA-BLIP2 and PN-VQA metrics show higher correlation with human alignment scores than prior zero-shot metrics.