Pith. sign in

Paper Citation Record · LEDGER

Visual question answering based evaluation metrics for text-to-image generation

As of 16 August 2026, this Paper Citation Record lists 16 of 16 outbound references and 0 inbound Pith citation observations for arXiv:2411.10183.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.10183 v1

Coverage vector

measured 16 of 16 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T19:54:42.529392Z

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

16 of 16 outbound references displayed

  • verified exact0
  • verified fuzzy2
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 59e36181-0e9f-4932-9507-d1e8149697b1 · outbound

This paper cites Training language models to follow instructions with human feedback,.

Visual question answering based evaluation metrics for text-to-image generation Training language models to follow instructions with human feedback,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T19:54:42.465801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:54:42.465801Z digest=sha256:d0c24adce63fe5586cb4d386f2ea4df4a431eb78a6bb100963cabe5e52de8c48

Observation a0d5c683-8111-43e9-93a9-130cd7213da4 · outbound

This paper cites Manigan: Text-guided image manipulation,.

Visual question answering based evaluation metrics for text-to-image generation Manigan: Text-guided image manipulation,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:54:42.682442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T19:54:42.470176Z digest=sha256:4e008b19d57c440c2e3bb23550e6c4c62eef4f23decbb2f222c2e1915dc0cf27

Observation 27e1e406-4bc2-4ce7-b72c-32ea8e07ed2a · outbound

This paper cites Unpaired image-to-image translation using cycle-consistent adversarial networks,.

Visual question answering based evaluation metrics for text-to-image generation Unpaired image-to-image translation using cycle-consistent adversarial networks,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T19:54:42.474766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:54:42.474766Z digest=sha256:b8862b120b5e31ac15932afcf2143b06a6ac956a48bc481d9ff075fc87869162

Observation 9e669340-ef06-44cc-b87f-bd161c177e01 · outbound

This paper cites High- resolution image synthesis with latent diffusion models,.

Visual question answering based evaluation metrics for text-to-image generation High- resolution image synthesis with latent diffusion models,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T19:54:42.478533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:54:42.478533Z digest=sha256:7bf946c564c0000ae1f75711c6fea32f73042e61e5bbc84a6b59b7043d3e4681

Observation 9b6c6e34-fcb1-450b-af34-91336205b16c · outbound

This paper cites Photorealistic text-to-image diffusion models with deep language understanding,.

Visual question answering based evaluation metrics for text-to-image generation Photorealistic text-to-image diffusion models with deep language understanding,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T19:54:42.483989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:54:42.483989Z digest=sha256:5c54fcdaa2a64062760801e78b088faa28390e10932bea8d6fdd3a17bd423425

Observation 76eb0e26-c0e6-4b4a-831a-58bf917d342a · outbound

This paper cites Microsoft coco: Common objects in context,.

Visual question answering based evaluation metrics for text-to-image generation Microsoft coco: Common objects in context,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T19:54:42.488583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:54:42.488583Z digest=sha256:dd190159de42e9d92aceb7dfcd229a16b7c5d82578eb9d4bbe510d471bfb2b6d

Observation 9945f1e4-f41c-42b2-ba2c-a52286b0d781 · outbound

This paper cites The caltech-ucsd birds-200-2011 dataset,.

Visual question answering based evaluation metrics for text-to-image generation The caltech-ucsd birds-200-2011 dataset,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T19:54:42.493087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:54:42.493087Z digest=sha256:fa5de9b4bd4dd07646bed974ab70a0336bb100a6665ddd478d627044899db1d7

Observation dfb46f02-7a86-42d8-b605-bfc4e4db71e3 · outbound

This paper cites DiffusionDB: A Large-scale Prompt Gallery Dataset for Text-to-Image Generative Models.

Visual question answering based evaluation metrics for text-to-image generation DiffusionDB: A Large-scale Prompt Gallery Dataset for Text-to-Image Generative Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T19:54:42.497045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:54:42.497045Z digest=sha256:7f55be52cca529986f2bf4fab526736cc65b40829348d583275261a54c393cc9

Observation ca985a11-068e-45e9-a019-3c82f2c0b846 · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilibrium,.

Visual question answering based evaluation metrics for text-to-image generation Gans trained by a two time-scale update rule converge to a local nash equilibrium,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T19:54:42.501323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:54:42.501323Z digest=sha256:ba8f855002a905765c595dc180c05fd9b83afc96921d3b0820e1c3120ecd425e

Observation 52aa3298-4076-4ac3-8e9f-b977db37caee · outbound

This paper cites CLIPScore: A Reference-free Evaluation Metric for Image Captioning.

Visual question answering based evaluation metrics for text-to-image generation CLIPScore: A Reference-free Evaluation Metric for Image Captioning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T19:54:42.505048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:54:42.505048Z digest=sha256:fbb25ee97e0bd66fd38043a2e6e8c4c5bd4b299471484f8cd3e7b4feb7acd96b

Observation 326edf4f-de19-4c23-8210-299b58136bd6 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Visual question answering based evaluation metrics for text-to-image generation Learning transferable visual models from natural language supervision,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T19:54:42.509243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:54:42.509243Z digest=sha256:e9e8514a477b9f5911c1d204b5491104cc689c89db255fdd8b8af6e8723c78de

Observation bb05ecb7-969c-4455-b32c-5e5e1f6be4ef · outbound

This paper cites Maniqa: Multi-dimension attention network for no-reference image quality assessment,.

Visual question answering based evaluation metrics for text-to-image generation Maniqa: Multi-dimension attention network for no-reference image quality assessment,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:54:42.625826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T19:54:42.512992Z digest=sha256:aef587f55802b477cb38018ffd720e56efdfb1e78c7433ddf8883f95eeef8233

Observation 6a16cdf1-f185-4c77-9d59-042e2ea74df1 · outbound

This paper cites ImageReward: Learning and Evaluating Human Preferences for Text-to-Image Generation.

Visual question answering based evaluation metrics for text-to-image generation ImageReward: Learning and Evaluating Human Preferences for Text-to-Image Generation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T19:54:42.516786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:54:42.516786Z digest=sha256:341ab43199c3a5d1f9837f0f96056791c331cb5abd2153414328ff1aa07c78e9

Observation 7e13e8d9-2fa7-4c44-8b52-07db5136489c · outbound

This paper cites Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks.

Visual question answering based evaluation metrics for text-to-image generation Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T19:54:42.521365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:54:42.521365Z digest=sha256:0fafa9f1be4c965cd2214b6983bdacd8dda6062e4b739274870090b70a6f9023

Observation 096d89bc-7cd3-4476-811a-827ad5502dab · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Visual question answering based evaluation metrics for text-to-image generation An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T19:54:42.525433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:54:42.525433Z digest=sha256:5d0689589ff110de69b0f11b60ae00abfeb0f886d3df1ef837cc4495605b650e

Observation 7c3d6c14-fe95-45b3-87cb-44f768a1a5f6 · outbound

This paper cites Attention is all you need,.

Visual question answering based evaluation metrics for text-to-image generation Attention is all you need,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T19:54:42.529392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:54:42.529392Z digest=sha256:b473273e0465aac2722f222ad18c8626e99fa2ff91dff402c0e9631fe8dd37ab

Pith citing papers

No inbound Pith citation observations are available.