Pith. sign in

Paper Citation Record · LEDGER

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models

As of 23 August 2026, this Paper Citation Record lists 90 of 90 outbound references and 11 inbound Pith citation observations for arXiv:2505.19415.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.19415 v2

Coverage vector

measured 90 of 90 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:17:40.539908Z

measured 101 of 101 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T12:16:44.681599Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

90 of 90 outbound references displayed

  • verified exact4
  • verified fuzzy27
  • unresolved58
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation ebc1c556-8043-4c91-9050-e5814de20360 · outbound

This paper cites Qwen2.5-VL Technical Report.

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models Qwen2.5-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:32.198193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:32.198193Z digest=sha256:e24c03d298d02b707065a50344b1a2fdfaaaa31fc9e1fe2163cda2966a7360f1

Observation 1f03809d-4cb9-41d9-a07d-ce2184c25094 · outbound

This paper cites Pixart- σ: Weak-to-strong training of diffusion transformer for 4k text-to-image generation.

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models Pixart- σ: Weak-to-strong training of diffusion transformer for 4k text-to-image generation

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:32.253961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:32.253961Z digest=sha256:e39a9d562a5f046917e75565bae1d62e20f70ba329f9931142a959e003e2dca3

Observation 956a145d-9daa-47d9-af78-c6027cfcaf70 · outbound

This paper cites Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling.

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:32.455904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:32.455904Z digest=sha256:91cf66bce725f5b188dcf2693fce7e5f5e90955d4ed02c7b593b15dd6c0b0741

Observation 319cba81-a631-4770-95e8-e365b887cfe8 · outbound

This paper cites Davidsonian Scene Graph: Improving Reliability in Fine-grained Evaluation for Text-to-Image Generation.

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models Davidsonian Scene Graph: Improving Reliability in Fine-grained Evaluation for Text-to-Image Generation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:32.547410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:32.547410Z digest=sha256:386d73a5bbe2867cf21ace3c1eb030a7f06e35416a5571c860320cf928fdc9ba

Observation 6278fd2d-b2e0-4c2b-8581-0dbde1d2c4c3 · outbound

This paper cites Dreambench, 2022.

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models Dreambench, 2022

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:32.623206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:32.623206Z digest=sha256:a532c9ce6edf0be57ff143b40fb2f6c333342a0890dea88875fa2dca6a6f2a6a

Observation 8ea65807-b9b2-4183-ae09-469ba6671eac · outbound

This paper cites Scaling Rectified Flow Transformers for High-Resolution Image Synthesis.

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models Scaling Rectified Flow Transformers for High-Resolution Image Synthesis

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:32.692807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:32.692807Z digest=sha256:9e63de6dc777633e8db0869bf3104eb46ce1fe9a1308f5f530d3408a11a4be00

Observation 1fc9f34c-9f78-4159-8240-ee816df5ed74 · outbound

This paper cites Personalize Anything for Free with Diffusion Transformer.

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models Personalize Anything for Free with Diffusion Transformer

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:32.762978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:32.762978Z digest=sha256:05d8025c89673af9d98343ea4a12935219069ee1f13c0126b3a8878508331db0

Observation 9b784e3f-2441-487a-a237-cf01ea15bdca · outbound

This paper cites Commonsense-T2I Challenge: Can Text-to-Image Generation Models Understand Commonsense?.

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models Commonsense-T2I Challenge: Can Text-to-Image Generation Models Understand Commonsense?

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:32.854514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:32.854514Z digest=sha256:7376965c3aa8271a555812b29d986c9cc55a5abb7de8effb42687f7aef2e571c

Observation 53f7d114-3ea6-4ae5-91c3-d463759d2edc · outbound

This paper cites An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion.

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:32.973359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:32.973359Z digest=sha256:9ed5bb6ab28fc84bf3a350b5fe0af91b436f02216c4527f2d3df03ecb1be6c25

Observation 826c7544-3676-4543-8fe7-c1950397d96c · outbound

This paper cites Geneval: An object-focused framework for evaluating text-to-image alignment.

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models Geneval: An object-focused framework for evaluating text-to-image alignment

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:33.046998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:33.046998Z digest=sha256:aab17cd952d50aacb09b7eb1318fb266b1ce7a905db2864f5ddc2b9e062c8390

Observation 8710b60f-8f99-42ca-9c35-fd726c61decc · outbound

This paper cites Gemini 2.0 flash, 2025.

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models Gemini 2.0 flash, 2025

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:33.128761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:33.128761Z digest=sha256:26b079619a54787c0d6a1c17eebe0a602d087cff7619164423607bd70a785d78

Observation 20e1c83d-9c5e-4fbe-8e3d-fccc0850c1db · outbound

This paper cites EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation.

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:33.225548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:33.225548Z digest=sha256:6879413e3fe4c0fafb7a215373a29fcc5f9179e7d3e84d2ca4bcb54d93eb30a5

Observation f6659ad7-bccc-4c04-be4e-03e0bfc88c44 · outbound

This paper cites ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment.

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:33.298586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:33.298586Z digest=sha256:29d4f1560f59669551208325737606c53c9089d58f38afa6e3ee8cb5dfe2aef9

Observation d55fa72d-bf2f-4a12-959e-5e0f50f827f7 · outbound

This paper cites PromptCap: Prompt-Guided Task-Aware Image Captioning.

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models PromptCap: Prompt-Guided Task-Aware Image Captioning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:33.369845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:33.369845Z digest=sha256:314db7038b4e7bf6de7b725abc44ab9ffbd3190b7606bd24e7680b894979c302

Observation 1df7e5ed-c53e-4a36-a83d-fcd0d8c2804e · outbound

This paper cites an unresolved cited work.

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:33.476722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:33.476722Z digest=sha256:7fdd2dbcacde0e1c26d0f57ca08b1268dd7d8fc8fc13ae5397ad93fd93b747b0

Observation c75cbd43-9ac1-4832-8609-29b20633391c · outbound

This paper cites Tifa: Accurate and interpretable text-to-image faithfulness evaluation with question answering.

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models Tifa: Accurate and interpretable text-to-image faithfulness evaluation with question answering

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:33.566324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:33.566324Z digest=sha256:af2cf8383faec93c3c62f043b5ff9bd90bd51dddf08ec2570940d1df8ce8c933

Observation 102a8bb1-d20b-4a86-bbb4-b1cecbb5e0f6 · outbound

This paper cites Finematch: Aspect-based fine-grained image and text mismatch detection and correction.

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models Finematch: Aspect-based fine-grained image and text mismatch detection and correction

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:33.636999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:33.636999Z digest=sha256:705257769effb13f154672f83c0d9afe0f1d09ec8b11c0f56dbb74efb9bce3c9

Observation a27e9937-2230-4ba7-bf44-1d7987b2a080 · outbound

This paper cites MMCOMPOSITION: Revisiting the Compositionality of Pre-trained Vision-Language Models.

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models MMCOMPOSITION: Revisiting the Compositionality of Pre-trained Vision-Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:33.709860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:33.709860Z digest=sha256:8776876a409b0434f35e92a7d078fefd2d5d02f0a7f81548e90b36c74a1abdfb

Observation 5f8d031a-0586-43c7-9b1c-814c32fe0561 · outbound

This paper cites T2i- compbench++: An enhanced and comprehensive benchmark for compositional text-to-image generation.

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models T2i- compbench++: An enhanced and comprehensive benchmark for compositional text-to-image generation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:33.783607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:33.783607Z digest=sha256:138f0edfdbe4035a060563d2346e72747411bf6bce97124ca409a46c8f9a959a

Observation 8f14445d-b9a3-4d2a-88dd-81f43faf8828 · outbound

This paper cites T2I-FactualBench: Benchmarking the Factuality of Text-to-Image Models with Knowledge-Intensive Concepts.

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models T2I-FactualBench: Benchmarking the Factuality of Text-to-Image Models with Knowledge-Intensive Concepts

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:33.979030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:33.979030Z digest=sha256:22cb3ab8a7acb7a5320a41f8b1efba0a18a2b73301fb54085e260bf46e16f97e

Observation facc9f8d-5068-4fa6-a19c-23d5ed5e58a7 · outbound

This paper cites Imagic: Text-based real image editing with diffusion models.

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models Imagic: Text-based real image editing with diffusion models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:17:48.094741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:17:34.067778Z digest=sha256:91c6e3ba360f317656726df900e955b128d1497d6b21aa079acf41979a6b58d6

Observation 1be814f8-41b5-4078-a06f-8ecc2696b562 · outbound

This paper cites Profashion: Prototype-guided fashion video generation with multiple reference images.

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models Profashion: Prototype-guided fashion video generation with multiple reference images

Reference 22

Resolution
verified exact
raw_fallback, observed 2026-08-07T14:17:42.018740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:17:34.166191Z digest=sha256:eaa5d73d3b142f8a97f90db5838e63cabba91771e68fb05f95b1577ebab75282

Observation f793b838-9fb8-4b72-99f4-f778726fc54a · outbound

This paper cites Are These the Same Apple? Comparing Images Based on Object Intrinsics.

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models Are These the Same Apple? Comparing Images Based on Object Intrinsics

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:17:41.674371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:17:34.260701Z digest=sha256:112965af1bf0748878b016788c9b1f02f461dfa337a3204b6194dc9cf06192bf

Observation c53286f2-f14c-4b16-946f-37d1001fc762 · outbound

This paper cites Multi- concept customization of text-to-image diffusion.

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models Multi- concept customization of text-to-image diffusion

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:17:47.918627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:17:34.334393Z digest=sha256:132d2d45fa8c9830870eb5af607fb55e4f66b4ebf4b5bec930fbe0eeca7f70ef

Observation 4a731ae5-bc4e-427a-8a05-85dbd9c687fb · outbound

This paper cites Flux.1, 2024.

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models Flux.1, 2024

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:17:47.702937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:17:34.415113Z digest=sha256:33d7328905778425857e997cb272cccf16b492dcd0168c5d78e4c00aa61c9f9e

Observation 18bb1ba9-1589-4f9b-bc54-7448ac5064d0 · outbound

This paper cites an unresolved cited work.

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:34.494533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:34.494533Z digest=sha256:2802871a73b21a5f525736277b757cd47315d17d9855c46e17d7b099894b06cf

Observation 9453b890-c1fc-457f-aff5-3b4d4ecc2bb0 · outbound

This paper cites Holistic evaluation of text-to-image models.

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models Holistic evaluation of text-to-image models

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:17:47.511934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:17:34.588651Z digest=sha256:7d658752af80249f15c16dbf3b7140c62fcb23a642821243955cbc9b3e1fafa1

Observation a6d82efc-93f1-4e0a-9be3-468e66cde8ea · outbound

This paper cites GenAI-Bench: Evaluating and Improving Compositional Text-to-Visual Generation.

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models GenAI-Bench: Evaluating and Improving Compositional Text-to-Visual Generation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:34.670156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:34.670156Z digest=sha256:89eeb1dce489510db2915895b8f563821902cf8622d169432f5d416d429a6bc4

Observation 23844377-395b-457c-90d4-a8abc9cd86da · outbound

This paper cites Blip-diffusion: Pre-trained subject representation for controllable text-to-image generation and editing.

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models Blip-diffusion: Pre-trained subject representation for controllable text-to-image generation and editing

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:34.761980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:34.761980Z digest=sha256:584d1ece4085b50f163f5658762e937e4a42f493a305c910c39d3b173b5b0108

Observation cd72b125-1a48-4f03-a3b9-fd888d5f8ff6 · outbound

This paper cites UNIMO-G: Unified Image Generation through Multimodal Conditional Diffusion.

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models UNIMO-G: Unified Image Generation through Multimodal Conditional Diffusion

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:34.850238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:34.850238Z digest=sha256:3299a6bc5899936b0bb55432009c51951fa43db60c0cd95a25b48e1d7c47b711

Observation 7464fcef-16b5-4120-bebc-5a2d6fff5501 · outbound

This paper cites Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding.

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:34.926538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:34.926538Z digest=sha256:fb5ca85f77688c75276d98042574f7f8dfebca0862b2d092feb53c96aeb1ba91

Observation bf9c1b92-59a5-4e20-a3a0-2fb338cd9c6a · outbound

This paper cites Evaluating text-to-visual generation with image-to-text generation.

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models Evaluating text-to-visual generation with image-to-text generation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:35.004844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:35.004844Z digest=sha256:b6ceb758c721d69ab86599b93962cde2e9bee33ba0dcefc4586ecf897733191f

Observation 4b4d0343-7423-42f1-a3a3-b60338de0636 · outbound

This paper cites JanusFlow: Harmonizing Autoregression and Rectified Flow for Unified Multimodal Understanding and Generation.

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models JanusFlow: Harmonizing Autoregression and Rectified Flow for Unified Multimodal Understanding and Generation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:35.084460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:35.084460Z digest=sha256:f9f976a930315c8d37b8608b1b8bdb99eb3e298e6cb34bf452f07fd88e9211a1

Observation 5a465f76-35ec-4672-9868-ce9356b68a4d · outbound

This paper cites Dreamo: A unified framework for image customization.

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models Dreamo: A unified framework for image customization

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:35.188045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:35.188045Z digest=sha256:74b613651619c62c9db368e2eb247cd8b72bbc0d8b90af9fd0586505466a1b71

Observation 167f5396-337a-49bb-8910-86bba60b3567 · outbound

This paper cites GPT-4 Technical Report.

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models GPT-4 Technical Report

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:35.272677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:35.272677Z digest=sha256:673897e8c6399bb5f698fc4cbd212dda426dce739b6d2b0d750bf20087b78742

Observation e6805a9a-285d-490d-aeb1-ddcb2fc189cb · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models DINOv2: Learning Robust Visual Features without Supervision

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:35.348998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:35.348998Z digest=sha256:fbfc3c2a1693c25b15eb5a99a4499def5acb97394b23fba45bac52763c449bfb

Observation f90bcb09-fe16-46c2-b5b9-42cba39f935c · outbound

This paper cites Kosmos-G: Generating Images in Context with Multimodal Large Language Models.

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models Kosmos-G: Generating Images in Context with Multimodal Large Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:35.431664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:35.431664Z digest=sha256:91cf24791f32a5028b9033e4f0325b7e813dabde8272a30766720fbe651fe629

Observation b107bfca-6bd5-4d03-bca3-0983d68e75d2 · outbound

This paper cites DreamBench++: A Human-Aligned Benchmark for Personalized Image Generation.

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models DreamBench++: A Human-Aligned Benchmark for Personalized Image Generation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:35.573979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:35.573979Z digest=sha256:b69992992710c5376cdf1a41f87dc9deb8daa4ee0330a1fe37f5c7788ba59c13

Observation 69fc506b-bb5b-4d22-a343-267aaf41f310 · outbound

This paper cites Pexels, 2014.

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models Pexels, 2014

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:17:47.157579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:17:35.642706Z digest=sha256:4fd8445b42e80d2c72fc0a5f9cf6c702e911f88f045df229317dc66dbfc42287

Observation 8217eeda-c442-454c-be1e-0e4a8dd56b7e · outbound

This paper cites an unresolved cited work.

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:17:47.356538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:17:35.505841Z digest=sha256:d30c14b23cd512cbe8c67bf23b1cb5ab29453f15a274b29dc8e9cd9b9a7136de

Observation f83cb8f0-3318-47aa-9d71-55c18f2ecb40 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:35.788931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:35.788931Z digest=sha256:abb201ed7c3e3aca30b593b8e72c618a40675630732a35bf2f67b78a982f1b67

Observation c3b6b41e-ba0c-419d-a92b-f756d57ba2b1 · outbound

This paper cites Lumina-Image 2.0: A Unified and Efficient Image Generative Framework.

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models Lumina-Image 2.0: A Unified and Efficient Image Generative Framework

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:35.875375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:35.875375Z digest=sha256:bf3177b5b0f15fdeb3c5332e6b8c7196cd990171ba5735c880aac80ef25ad008

Observation 0ff6849a-201c-4bbc-b620-e72f5cf7cd6c · outbound

This paper cites Photon-v1.

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models Photon-v1

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:17:46.988130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:17:35.710792Z digest=sha256:d5d5efb91e131e3aee8f91a17e3a5245d07ac15b454d1fe974935c357087a0f5

Observation 468fbcd0-f144-4b3b-8730-38fc84468f23 · outbound

This paper cites Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation.

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:36.061706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:36.061706Z digest=sha256:bdaa9477788a8501616d384d55e1a5117aed229ba0e51189271e4313c9c55289

Observation dc3dfb22-7bf2-43c9-8ec9-8455237536e5 · outbound

This paper cites Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation.

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:36.141458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:36.141458Z digest=sha256:064330bf9cc78065c9455178200bfefc0055ffc3171e24ca14efc42efd016413

Observation 6c3cbda5-e4ef-4687-b65c-36fb8587355a · outbound

This paper cites Learning transferable visual models from natural language supervision.

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models Learning transferable visual models from natural language supervision

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:17:46.853853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:17:35.997395Z digest=sha256:e52a9ca9172e0346e46fb29a768b072af56d61e42eff3c9fd85cb377f1ed7a24

Observation 05756949-a24e-49d3-88a4-2009cb79e1fd · outbound

This paper cites Instantbooth: Personalized text-to- image generation without test-time finetuning.

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models Instantbooth: Personalized text-to- image generation without test-time finetuning

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:17:46.719733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:17:36.285110Z digest=sha256:6ceb25274cf867d3c4ae7ca18542e5b2cdb810343093ba505098da9f617115bb

Observation 8b1da1ca-b31c-4b14-9176-dcd357f87f5f · outbound

This paper cites Generative multimodal models are in-context learners.

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models Generative multimodal models are in-context learners

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:36.405666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:36.405666Z digest=sha256:aa13d28eb717ed901154c4f30287241c5085011b0717e65c84be4d99b62bebea

Observation 69cf6613-43d6-4a08-9c2e-3fa75499fcf5 · outbound

This paper cites Hyperdreambooth: Hypernetworks for fast personalization of text-to-image models, 2023.

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models Hyperdreambooth: Hypernetworks for fast personalization of text-to-image models, 2023

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:36.211649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:36.211649Z digest=sha256:bf8306fb103bad4db5cc1f132d73f6c09b1ce6e778f1bcaff68aabb4b8d05616

Observation 7ec74167-9983-4116-852e-cf64f108f876 · outbound

This paper cites Hidream-i1: A 17b parameter open chinese text-to-image generation model.

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models Hidream-i1: A 17b parameter open chinese text-to-image generation model

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:17:46.602631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:17:36.556940Z digest=sha256:80e293ced9a3a11c219761f7221bfa8785edafbc0d5b7ebca7eb00526fb084c6

Observation 7a151e95-6de1-45ee-be06-a50e0ac452d1 · outbound

This paper cites Lmm4lmm: Benchmarking and evaluating large-multimodal image generation with lmms.

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models Lmm4lmm: Benchmarking and evaluating large-multimodal image generation with lmms

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:17:46.491274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:17:36.625308Z digest=sha256:4cc77040a3be00381ffe295b06ad28a231f955522c19db97f242542e4bdf16b0

Observation 0502b1cb-59d9-4cd7-9c72-05f7b681da84 · outbound

This paper cites Vidcomposition: Can mllms analyze compositions in compiled videos? arXiv preprint arXiv:2411.10979, 2024.

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models Vidcomposition: Can mllms analyze compositions in compiled videos? arXiv preprint arXiv:2411.10979, 2024

Reference 52

Resolution
verified exact
raw_fallback, observed 2026-08-07T14:17:41.178957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:17:36.490359Z digest=sha256:25fc12c2b3e4615a074ac931d233e675d65eb027f242b895f0a06bedbde3918d

Observation 5402eb20-67f8-40bb-8f0f-1af7c66d91a3 · outbound

This paper cites InstantID: Zero-shot Identity-Preserving Generation in Seconds.

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models InstantID: Zero-shot Identity-Preserving Generation in Seconds

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:36.766473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:36.766473Z digest=sha256:b79641965198e5316322e7c69c12d557744e0eb88eb7ae6482094f4c91d7b179

Observation 3d81d521-1b15-49d3-a362-31d62bcf6695 · outbound

This paper cites MS-Diffusion: Multi-subject Zero-shot Image Personalization with Layout Guidance.

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models MS-Diffusion: Multi-subject Zero-shot Image Personalization with Layout Guidance

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:36.838755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:36.838755Z digest=sha256:6c085015e512ff5aa8302882616b99671d40281ba9a7dacf2ede86fe3d3ac26a

Observation 71ddb135-76bf-4ce2-a798-d99731f022ae · outbound

This paper cites PF-LRM: Pose-Free Large Reconstruction Model for Joint Pose and Shape Prediction.

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models PF-LRM: Pose-Free Large Reconstruction Model for Joint Pose and Shape Prediction

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:36.698633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:36.698633Z digest=sha256:0257a3d54341a7340c98e475f4c4c384d11ad1f43cc727529a17527f1d3382c9

Observation 20a05d4f-385e-4ad3-841c-da2f9ff8e0e6 · outbound

This paper cites Personalized Image Generation with Deep Generative Models: A Decade Survey.

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models Personalized Image Generation with Deep Generative Models: A Decade Survey

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:37.020233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:37.020233Z digest=sha256:d31ae48d2d5263983b9b4072b4146f5f465c91a3f82202882521f04864fba155

Observation 58ad5f0f-98bf-4095-99ed-04eb8fb06ffb · outbound

This paper cites Revisiting Text-to-Image Evaluation with Gecko: On Metrics, Prompts, and Human Ratings.

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models Revisiting Text-to-Image Evaluation with Gecko: On Metrics, Prompts, and Human Ratings

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:37.111721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:37.111721Z digest=sha256:ee936bea6a3a7e6d5b76525faa53f410880faff42a58ca15fa33d1ae7e4cb330

Observation 641edb97-7833-4ee0-8720-1cf388511536 · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models Emu3: Next-Token Prediction is All You Need

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:36.950135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:36.950135Z digest=sha256:8c4189e79c7df03124150478d09268e0ba92e93e62dd00d5920286bfa0e114ec

Observation 84b0be7e-5a20-4e90-bef7-5d873ddbddf3 · outbound

This paper cites What You See is What You Read? Improving Text-Image Alignment Evaluation.

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models What You See is What You Read? Improving Text-Image Alignment Evaluation

Reference 59

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:17:40.822374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:17:37.320626Z digest=sha256:156ff0b2a17ca2926e84948106953a3eb1d43507e5a46428070804a634a0d410

Observation 7560ce16-f122-4c6f-a068-4fd984d2c536 · outbound

This paper cites IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models.

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:37.464479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:37.464479Z digest=sha256:8ae96d71e2a2dde4379f9927664a91ddc04cae3d2dfdb690bee15b62ea0a4297

Observation 98efdf5c-162b-436f-a7ae-e9d13498de4a · outbound

This paper cites GroundingBooth: Grounding Text-to-Image Customization.

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models GroundingBooth: Grounding Text-to-Image Customization

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:37.221759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:37.221759Z digest=sha256:07bd422d79e243474156183867885ddcf81be3526ed0df454a80fb5d6f05cc32

Observation 6d14cf9a-04ef-479f-978f-2ca49404fe8e · outbound

This paper cites Perceptual artifacts localization for image synthesis tasks.

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models Perceptual artifacts localization for image synthesis tasks

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:17:46.295760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:17:37.679832Z digest=sha256:4f7309c0cc1036dfbead2ed5478845a70652234444a86ada592f144a6ac6bbc0

Observation b0149de3-cf12-4115-9dc5-5f713a901145 · outbound

This paper cites Q-Eval-100K: Evaluating Visual Quality and Alignment Level for Text-to-Vision Content.

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models Q-Eval-100K: Evaluating Visual Quality and Alignment Level for Text-to-Vision Content

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:37.730771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:37.730771Z digest=sha256:c8a5b3649fef651ff86170902ff16a2aae36e9ec880d8a7d12a9a3b4b0345ddc

Observation 00ddbd14-d2c0-43e8-82e2-bc26f12bed67 · outbound

This paper cites Scaling autoregressive models for content-rich text-to-image generation.

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models Scaling autoregressive models for content-rich text-to-image generation

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:17:46.364786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:17:37.566774Z digest=sha256:ec3ee10edad2977bd7e43afa02da1f57c1f6c7394e305453cdb3c6567090e94e

Observation d96c673d-0991-499b-b77a-10501cf48a12 · outbound

This paper cites EasyRef: Omni-Generalized Group Image Reference for Diffusion Models via Multimodal LLM.

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models EasyRef: Omni-Generalized Group Image Reference for Diffusion Models via Multimodal LLM

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:38.004401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:38.004401Z digest=sha256:6b604043d6b984dca41c9d52c1e9040b03d9a632e095fd431faf9617f307309f

Observation 1dd032ed-7d5e-43ad-9f67-9266eec42b7f · outbound

This paper cites CogView3: Finer and Faster Text-to-Image Generation via Relay Diffusion.

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models CogView3: Finer and Faster Text-to-Image Generation via Relay Diffusion

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:37.891666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:37.891666Z digest=sha256:9a2b5962620646ea5329700cbd37bb8031322aa448cd6b6da634e8455fa63695

Observation a6bc888e-597b-4db3-887d-187f3a74ca97 · outbound

This paper cites entity" should be common objects; e.g., chair, dog, car, lamp, etc.

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models entity" should be common objects; e.g., chair, dog, car, lamp, etc

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:17:45.929187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:17:38.111737Z digest=sha256:28f5f71182b5d391185656325829762247dc758618319abb50c9e00923a14b83

Observation 0fcde27d-f7e5-404b-9e57-2b7ad9ba2367 · outbound

This paper cites interaction.

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models interaction

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:17:45.586397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:17:38.349730Z digest=sha256:76c5a085d68329e20cef2c64e07b87fcac9b52f54df85640ee438ef3e9a4ee55

Observation e8a75886-d99c-4e0f-8e5a-05466eb33169 · outbound

This paper cites scene description.

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models scene description

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:17:45.181082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:17:38.630873Z digest=sha256:ba67158f5bd97ff7fda88873953643c37203cbb7bd5d07ed3f0ad7043def403c

Observation 5be5e40a-857a-46a3-a811-208eacc51a2b · outbound

This paper cites A robot and a dolphin dancing under the ocean, surrounded by swirling schools of fish.

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models A robot and a dolphin dancing under the ocean, surrounded by swirling schools of fish

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:17:44.886420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:17:38.875063Z digest=sha256:6e6807dbdaf1d0f84fc0b8b38672f3fa218146ed33a03ab497292c99791aa809

Observation 735d8e0d-beb4-449d-a6b1-f560dbc19cb0 · outbound

This paper cites minimalism meets hygge vibes / editorial photoshoot style / baroque detail / etc.

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models minimalism meets hygge vibes / editorial photoshoot style / baroque detail / etc

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:17:44.743428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:17:38.960242Z digest=sha256:f5a31117f2fbe090cbcd667622e7dada0132bf6fd793e4ea783705c3d457914b

Observation 5cc877b0-2182-4f37-8d98-e93c45ffa3c1 · outbound

This paper cites negation.

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models negation

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:17:44.580482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:17:39.173467Z digest=sha256:22098206fb6e8b0df194f708df5f36d04c779e3cb99c05428622d015259d1cf9

Observation 66a896e8-1995-400f-b294-dcd635eff30e · outbound

This paper cites an unresolved cited work.

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models Unresolved cited work

Reference 81

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:17:45.033190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:17:39.228478Z digest=sha256:c98b8f80b4fa81a2db0d86ecf039d923a1d2ea9cb942f1f55db8d4a25f990f4a

Observation f5cff7ec-054c-4c11-8c73-0cf056664659 · outbound

This paper cites comparison.

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models comparison

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:17:44.446713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:17:39.326589Z digest=sha256:dc77a286ba84b54c49b2e69b1bf6b842e2dce3ab7e28a9ec44ea05144ce26713

Observation 9737e77f-fa97-42e5-a886-d1312f516c8c · outbound

This paper cites an unresolved cited work.

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models Unresolved cited work

Reference 83

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:17:44.340233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:17:39.437725Z digest=sha256:a8c02afa0b3df1039b765035704abc4e637b92efdcd1e14ac6f2d8cdb41fe356

Observation 6655b434-904b-4727-83dd-986640f3082a · outbound

This paper cites an unresolved cited work.

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models Unresolved cited work

Reference 84

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:17:44.286726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:17:39.529673Z digest=sha256:aaaa6f0e711a4c2aae8e8ac4648ec0aee3b1ea5ab20ff2a41a52d431c65455d2

Observation 5cb1be73-2647-4c9c-a686-99b7f8ed8684 · outbound

This paper cites an unresolved cited work.

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models Unresolved cited work

Reference 85

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:17:44.159110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:17:39.623980Z digest=sha256:fa96d0350e735753bd9c436517d0a23ea229bcf4cf7c98b49294631f1435976a

Observation 877d5ceb-73bf-4f1e-8be8-e009a76cccc9 · outbound

This paper cites an unresolved cited work.

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models Unresolved cited work

Reference 86

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:17:44.006604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:17:39.748781Z digest=sha256:e7459f3cfe518d392416129529c86d25b82e95e0dc977a394d47735d02d71420

Observation f600f510-0200-4e99-bcbc-6637f29b999a · outbound

This paper cites an unresolved cited work.

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models Unresolved cited work

Reference 87

Resolution
parse uncertain
raw_fallback, observed 2026-08-07T14:17:43.852269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:17:39.870642Z digest=sha256:4e1b85c4eed5a825d0671264d0026c7bb19e44ef67a929f2df6c92bd8f947c66

Observation da5dab7a-73a5-42ad-b19a-2c7d40487f86 · outbound

This paper cites [scene description (op- tional)] + [number][attribute][entity1] + [interaction (spatial or action)] + [number (optional)][attribute][entity2].

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models [scene description (op- tional)] + [number][attribute][entity1] + [interaction (spatial or action)] + [number (optional)][attribute][entity2]

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:17:43.723163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:17:39.933217Z digest=sha256:e75d78257b62734b3950731e6bae8a1ab7a91c93dc82db8146c766b4bfc10ba6

Observation 0d566e46-0c93-4de3-a12b-717ca7d71231 · outbound

This paper cites entity" should be common objects; e.g., chair, dog, car, lamp, etc.

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models entity" should be common objects; e.g., chair, dog, car, lamp, etc

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:17:43.559229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:17:40.060937Z digest=sha256:9957a116b4233be6546e9b8b6459597360bc20ac5ae7c99994896642b586eb14

Observation 565409cb-f28d-4d98-95c6-948cf8dc5c63 · outbound

This paper cites attribute.

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models attribute

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:17:45.747117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:17:40.151845Z digest=sha256:985426d47d6120173306730e05bad48a394b85a7df205e094fc4f9031476b48a

Observation b7602f8b-d29c-4bc9-8ccf-fcc9fbe2f955 · outbound

This paper cites number" should be.

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models number" should be

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:17:43.136198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:17:40.209199Z digest=sha256:76c21cb3d3e96dd32552ea01250e5e563213de0d8a6f5912b621f565ee95bc53

Observation ade96861-c5d7-4e22-9424-e63d40f00876 · outbound

This paper cites interaction.

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models interaction

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:17:42.711995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:17:40.307178Z digest=sha256:517855ddb056b3528807982ba38f3efbcdd72c51c6d2f35f6ac8d3d837e3f189

Observation 647baadd-296f-4805-976d-e3cbbc00122c · outbound

This paper cites scene description.

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models scene description

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:17:45.450909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:17:40.375452Z digest=sha256:1d3b38132e7724fe586b752485f9967a7b8ecfac4807ec381bcd9933854f00b9

Observation 9c65d3ed-347e-481f-a3dc-469ea52434c1 · outbound

This paper cites interaction action.

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models interaction action

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:17:45.307905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:17:40.477293Z digest=sha256:a2da94a70f36f7045a8636b9713e75f745d726a7d3e0b2ca7b9b889b73b06362

Observation dc85f316-5abb-46c4-8812-1e264cfd02a0 · outbound

This paper cites scene description.

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models scene description

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:17:42.341356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:17:40.539908Z digest=sha256:cf02229ff5bfb876841f036d208fed83eedc05767f8d329689b83cb20afc8ff2

Observation efc1e1c9-8c17-4742-b9b0-0139c9dba46c · outbound

This paper cites an unresolved cited work.

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models Unresolved cited work

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:33.880000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:33.880000Z digest=sha256:f903873c2588876099a927e7b682fd70431ba6c9e6b073c433d711d7dba9d296

Observation d18b1dc0-f4fa-498c-909f-2b88b1c80adb · outbound

This paper cites an unresolved cited work.

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models Unresolved cited work

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:32.351287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:32.351287Z digest=sha256:b6aadc496ee7d50f9a14f875679f5e98e2e59000358305d209b264df4c47971f

Observation 4cbbba06-4115-462f-96b0-f121258372c7 · outbound

This paper cites an unresolved cited work.

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models Unresolved cited work

Reference 2025

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:17:46.110682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:17:37.799160Z digest=sha256:0b4807d14c464ad87d9621da08b07fb953cc726d067ff7e3a9df8d76ffa86efb

Pith citing papers

Observation 5d2f1016-d7b9-4996-bbf9-ccf4cb5a9a2d · inbound

ChartNet: A Million-Scale, High-Quality Multimodal Dataset for Robust Chart Understanding cites this paper.

ChartNet: A Million-Scale, High-Quality Multimodal Dataset for Robust Chart Understanding MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T22:39:32.913108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-14T22:39:06.113655Z digest=sha256:69599b3b64565ddee3f2e5d4adf268e33c77f3f170dd1ba51138f0a5d310a07a

Observation d0fbf8be-567f-4290-890c-45d1d41a5c17 · inbound

Skill-Aligned Annotation for Reliable Evaluation in Text-to-Image Generation cites this paper.

Skill-Aligned Annotation for Reliable Evaluation in Text-to-Image Generation MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:19:28.535511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-14T20:13:21.750639Z digest=sha256:b357b453f1ff8aec7f2e5bc78a0c462945e588a696bd70ca6c7ebace6f1a90da

Observation 79e960fa-1a9a-4a6d-ae8d-e8d73862762e · inbound

MementoGUI: Learning Agentic Multimodal Memory Control for Long-Horizon GUI Agents cites this paper.

MementoGUI: Learning Agentic Multimodal Memory Control for Long-Horizon GUI Agents MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:58:15.046298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-20T11:55:36.734758Z digest=sha256:bb86f43820abc6dd5db0a40d0039929d6733020a961cbe7ab7a6e7986cb51f1a

Observation dfe5db5b-ba7d-4510-b033-d47557dce56a · inbound

Agent Skills Should Go Beyond Text: The Case for Visual Skills cites this paper.

Agent Skills Should Go Beyond Text: The Case for Visual Skills MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-06-28T17:12:23.921117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T17:11:37.457810Z digest=sha256:c79350ed8f29cf0c3e4c78da47032b970b2c1f29369eb34d63d9f9e7c8338ba1

Observation 6702f4e2-7b6d-4114-840b-f92783a07124 · inbound

MemoBench: Benchmarking World Modeling in Dynamically Changing Environments cites this paper.

MemoBench: Benchmarking World Modeling in Dynamically Changing Environments MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T18:25:57.796713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T02:03:45.564122Z digest=sha256:399ca4520e6cdbcc0711818d15bb1d4a99b1087c628ade01bf8e0c464e3d5ccc

Observation 540ad50b-fbea-492e-9c25-5def5ee16eb6 · inbound

MemoBench: Benchmarking World Modeling in Dynamically Changing Environments cites this paper.

MemoBench: Benchmarking World Modeling in Dynamically Changing Environments MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T09:35:40.637749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T06:25:58.872140Z digest=sha256:59ce5a0ba0cb396bffccb73a254d0ce6c1eec4522f39e5f76da16f1dbeb32279

Observation 9da87cec-30ae-4453-8d54-b06e830de4b7 · inbound

MemoBench: Benchmarking World Modeling in Dynamically Changing Environments cites this paper.

MemoBench: Benchmarking World Modeling in Dynamically Changing Environments MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T20:57:22.796523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-02T20:52:28.444524Z digest=sha256:651ee43f4f67814bd97356326f00dce838fe5dc18534088e4c1f8141879ae4b5

Observation cb39151b-0dc9-4e56-a9ef-4390812c75ec · inbound

MemoBench: Benchmarking World Modeling in Dynamically Changing Environments cites this paper.

MemoBench: Benchmarking World Modeling in Dynamically Changing Environments MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T22:49:00.966399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-03T22:44:16.272541Z digest=sha256:3eddf532f7b84a87a8517e248338a6c1752ed9ab036bf0824def84754ed901a6

Observation a2758f89-69d3-4344-b147-f6a0645cd615 · inbound

MemoBench: Benchmarking World Modeling in Dynamically Changing Environments cites this paper.

MemoBench: Benchmarking World Modeling in Dynamically Changing Environments MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-14T17:14:19.770867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T17:14:19.770867Z digest=sha256:faea68fda7a443ca9b67d3eb1f528b05edb622778aa03e24a767fb4b5c9561c3

Observation a397dfef-885e-4bdc-ab6d-621ecb59b936 · inbound

MemoBench: Benchmarking World Modeling in Dynamically Changing Environments cites this paper.

MemoBench: Benchmarking World Modeling in Dynamically Changing Environments MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T10:00:10.910569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:00:10.910569Z digest=sha256:5ac2ed35d380c4e453607bed5d8dc6d5d7d6bc8986f766ccac8be9ce45c0d812

Observation f88a1b98-a342-4fd0-93d6-3d990ade54d9 · inbound

VideoArgus: Agentic Rubric-Grounded Unified Evaluation for Video Generation and Editing cites this paper.

VideoArgus: Agentic Rubric-Grounded Unified Evaluation for Video Generation and Editing MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T12:16:44.681599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:16:44.681599Z digest=sha256:52f215c2ce6ee6ac50b9ce6ea48e423a1136eca7a6e8c42b025c8e80562e9416