Pith. sign in

Paper Citation Record · LEDGER

TIFA: Accurate and Interpretable Text-to-Image Faithfulness Evaluation with Question Answering

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 26 inbound Pith citation observations for arXiv:2303.11897.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2303.11897 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 26 of 26 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T04:43:11.260035Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T14:57:03.744199Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f3b4525a-3f97-4bed-9eb8-9dd8ca0791b8 · inbound

GAIA: a benchmark for General AI Assistants cites this paper.

GAIA: a benchmark for General AI Assistants TIFA: Accurate and Interpretable Text-to-Image Faithfulness Evaluation with Question Answering

Reference 168

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T15:46:03.602428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-12T15:46:03.247029Z digest=sha256:6d10ed38645a23f1cff8fd2a058d8f790e361909df9e309764ee2a998f52f55f

Observation 532dd131-098b-434e-a28d-6f80c7074153 · inbound

BLINK: Multimodal Large Language Models Can See but Not Perceive cites this paper.

BLINK: Multimodal Large Language Models Can See but Not Perceive TIFA: Accurate and Interpretable Text-to-Image Faithfulness Evaluation with Question Answering

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T20:18:15.689615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:3e2344baf382ad54708fae9f2435c3830e3f4fc7d72bcc02216709f656438048

Observation 7f9a7d9d-96ab-4de3-aaec-cd9e8c6a90f8 · inbound

Evaluating Hallucination in Text-to-Image Diffusion Models with Scene-Graph based Question-Answering Agent cites this paper.

Evaluating Hallucination in Text-to-Image Diffusion Models with Scene-Graph based Question-Answering Agent TIFA: Accurate and Interpretable Text-to-Image Faithfulness Evaluation with Question Answering

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T20:28:49.170451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:28:49.170451Z digest=sha256:5cefd2128c5c563d343a13558b8e1340a119e62f22bc4856a74b85834bb1bc22

Observation 30cca4b0-6871-4574-a296-a74f49fb3437 · inbound

Efficient Scaling of Diffusion Transformers for Text-to-Image Generation cites this paper.

Efficient Scaling of Diffusion Transformers for Text-to-Image Generation TIFA: Accurate and Interpretable Text-to-Image Faithfulness Evaluation with Question Answering

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-11T14:11:00.550971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:11:00.550971Z digest=sha256:3aa448dc502d4cd5f1dcbaa080e79b5c2f16a5b64a574d4487e2d034111cdbba

Observation 1a453dab-995e-4016-99c2-20ff644478da · inbound

What makes a good metric? Evaluating automatic metrics for text-to-image consistency cites this paper.

What makes a good metric? Evaluating automatic metrics for text-to-image consistency TIFA: Accurate and Interpretable Text-to-Image Faithfulness Evaluation with Question Answering

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T12:38:06.589512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:38:06.589512Z digest=sha256:337e5bd80586af2cca6c405f9baa8ba4f0409f4e79a0b44d3f53220bea06f9b2

Observation 17686a8f-78c5-4de7-a3ab-0112db891e5d · inbound

Multi-Modal Language Models as Text-to-Image Model Evaluators cites this paper.

Multi-Modal Language Models as Text-to-Image Model Evaluators TIFA: Accurate and Interpretable Text-to-Image Faithfulness Evaluation with Question Answering

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:11.260035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:11.260035Z digest=sha256:6642c1e8fc55d10627c03f7f9f0e9365d170d06bc2987cba4cb431045dfd0cc7

Observation 338c2a94-95c9-43ff-94cb-622dd6a851e1 · inbound

TNG-CLIP:Training-Time Negation Data Generation for Negation Awareness of CLIP cites this paper.

TNG-CLIP:Training-Time Negation Data Generation for Negation Awareness of CLIP TIFA: Accurate and Interpretable Text-to-Image Faithfulness Evaluation with Question Answering

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:35:00.351522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:35:00.351522Z digest=sha256:d9bc5eac7ee47c3ac91426088132204fe2c76004ec5923e0b471d81d16474127

Observation ac972a64-6cde-41a2-b6f0-db927337c602 · inbound

CoCA: Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning cites this paper.

CoCA: Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning TIFA: Accurate and Interpretable Text-to-Image Faithfulness Evaluation with Question Answering

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:26.101775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:26.101775Z digest=sha256:6ead479ca5774361d3f0e66eba9259d3e907fdac98c888c4c2c59a33bf568bbf

Observation ca6781f0-1d25-4fe7-b785-4916d580ee93 · inbound

DIMCIM: A Quantitative Evaluation Framework for Default-mode Diversity and Generalization in Text-to-Image Generative Models cites this paper.

DIMCIM: A Quantitative Evaluation Framework for Default-mode Diversity and Generalization in Text-to-Image Generative Models TIFA: Accurate and Interpretable Text-to-Image Faithfulness Evaluation with Question Answering

Reference 2020

Resolution
malformed identifier
no resolver link, observed 2026-08-07T10:30:27.996361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:30:27.996361Z digest=sha256:94a628d6ab81b20bf258d5670d6b7fbee082a017cae49ec8bb8f92c8f8350478

Observation f0bc7fa2-ab92-42b3-8734-66d2cc1803e0 · inbound

Audit & Repair: An Agentic Framework for Consistent Story Visualization in Text-to-Image Diffusion Models cites this paper.

Audit & Repair: An Agentic Framework for Consistent Story Visualization in Text-to-Image Diffusion Models TIFA: Accurate and Interpretable Text-to-Image Faithfulness Evaluation with Question Answering

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T18:47:32.947904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:47:32.947904Z digest=sha256:68e1d284fac97f8aa9797ecab404bb876ad2c43fb3fbe9c3fd04e08becec4d0a

Observation ea392ca7-4baf-476d-9aa9-5cd84cd6d310 · inbound

AIGVE-MACS: Unified Multi-Aspect Commenting and Scoring Model for AI-Generated Video Evaluation cites this paper.

AIGVE-MACS: Unified Multi-Aspect Commenting and Scoring Model for AI-Generated Video Evaluation TIFA: Accurate and Interpretable Text-to-Image Faithfulness Evaluation with Question Answering

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T21:03:00.879353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:03:00.879353Z digest=sha256:48dc747a5ba197c644c1b8b98c55953ecf26ae1bd0746bb43d65de0ba61251b0

Observation 676ddf00-af40-400e-bdb1-fad7f263a898 · inbound

Towards Evaluating Robustness of Prompt Adherence in Text to Image Models cites this paper.

Towards Evaluating Robustness of Prompt Adherence in Text to Image Models TIFA: Accurate and Interpretable Text-to-Image Faithfulness Evaluation with Question Answering

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T18:50:42.666253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:50:42.666253Z digest=sha256:b2a2e8c8ce27e2ed013b1312e65476d09a4fd0f0e3e596984807f85f67c48c18

Observation 67850597-ba20-4148-bcdb-b7fd79c80a67 · inbound

Disentanglement in T-space for Faster and Distributed Training of Diffusion Models with Fewer Latent-states cites this paper.

Disentanglement in T-space for Faster and Distributed Training of Diffusion Models with Fewer Latent-states TIFA: Accurate and Interpretable Text-to-Image Faithfulness Evaluation with Question Answering

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T18:41:32.084239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T18:41:32.084239Z digest=sha256:c55c5fa0ab00d06fa84aa80615658be11ff2d1b162d53359b5d08b526b5849a2

Observation 513f1f98-cbeb-4fb5-ba78-095efd307efc · inbound

The Mind's Eye: A Multi-Faceted Reward Framework for Guiding Visual Metaphor Generation cites this paper.

The Mind's Eye: A Multi-Faceted Reward Framework for Guiding Visual Metaphor Generation TIFA: Accurate and Interpretable Text-to-Image Faithfulness Evaluation with Question Answering

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T17:00:55.999832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:00:55.999832Z digest=sha256:c2f6765afcf7f6ebe27d33d84f5be4cdd85b58c3363691be5bffcc4a93ec05ca

Observation 5fa3a9d5-e248-4f60-a87a-216b5424476f · inbound

Understanding and evaluating computer vision models through the lens of counterfactuals cites this paper.

Understanding and evaluating computer vision models through the lens of counterfactuals TIFA: Accurate and Interpretable Text-to-Image Faithfulness Evaluation with Question Answering

Reference 208

Resolution
unresolved
no resolver link, observed 2026-08-05T14:49:31.846334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:49:31.846334Z digest=sha256:265c994c2cdd77f9895382a916a8f65ce98a2a29575e83012eaedf631fcafe1d

Observation e75287e3-5e9d-47c4-a4c7-88e450328cb7 · inbound

Discovering Divergent Representations between Text-to-Image Models cites this paper.

Discovering Divergent Representations between Text-to-Image Models TIFA: Accurate and Interpretable Text-to-Image Faithfulness Evaluation with Question Answering

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T19:59:32.044233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:59:32.044233Z digest=sha256:7536929cdbf10039558b86bc29cbe0efc6b7112412136dd655d3fae2c832fa52

Observation 58469a3f-45db-4005-a7a3-ef05c61946c9 · inbound

Faithful-First Reasoning, Planning, and Acting for Multimodal LLMs cites this paper.

Faithful-First Reasoning, Planning, and Acting for Multimodal LLMs TIFA: Accurate and Interpretable Text-to-Image Faithfulness Evaluation with Question Answering

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-17T23:45:30.714518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-17T23:43:43.201888Z digest=sha256:b7276ba2d7fec69687b9cc3533296ea56f12450e042c46840713a3feb3ea81a4

Observation 63356f5d-cb41-4586-97de-a4ed4b9b54bd · inbound

Evaluating AI-Generated Images of Cultural Artifacts with Community-Informed Rubrics cites this paper.

Evaluating AI-Generated Images of Cultural Artifacts with Community-Informed Rubrics TIFA: Accurate and Interpretable Text-to-Image Faithfulness Evaluation with Question Answering

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-13T21:03:19.982943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-13T20:59:52.448832Z digest=sha256:338aab57797c40856d105f7664261c2a6b6a17c77f1359162f7af483624455eb

Observation 737195c3-2cc8-435c-bc32-5f5e19954836 · inbound

Evaluating AI-Generated Images of Cultural Artifacts with Community-Informed Rubrics cites this paper.

Evaluating AI-Generated Images of Cultural Artifacts with Community-Informed Rubrics TIFA: Accurate and Interpretable Text-to-Image Faithfulness Evaluation with Question Answering

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-21T10:40:00.755494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-21T10:35:39.269869Z digest=sha256:a990711833b30ede85982837220cdc9f48d435278a5644aec954bec38d3ec349

Observation 71710eff-027c-49da-8616-035590152c32 · inbound

QEVA: A Reference-Free Evaluation Metric for Narrative Video Summarization with Multimodal Question Answering cites this paper.

QEVA: A Reference-Free Evaluation Metric for Narrative Video Summarization with Multimodal Question Answering TIFA: Accurate and Interpretable Text-to-Image Faithfulness Evaluation with Question Answering

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T21:36:16.404118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-08T04:51:18.203826Z digest=sha256:a55b1624d09b7d81114e4c2e71c075940b687a64ed76957803d0c7f8864b88d8

Observation 89b589b9-67e5-432f-b3c4-ee3559a019dc · inbound

AutoRubric-T2I: Robust Rule-Based Reward Model for Text-to-Image Alignment cites this paper.

AutoRubric-T2I: Robust Rule-Based Reward Model for Text-to-Image Alignment TIFA: Accurate and Interpretable Text-to-Image Faithfulness Evaluation with Question Answering

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:13:16.092852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T12:11:23.775843Z digest=sha256:74801680635697e8282ea35d1cde04acf9bbc86a33a7b80e9cb4d5b0e900b9ae

Observation 7b2dff6e-02cf-4253-90af-4f3cb4d3f9a3 · inbound

AutoRubric-T2I: Robust Rule-Based Reward Model for Text-to-Image Alignment cites this paper.

AutoRubric-T2I: Robust Rule-Based Reward Model for Text-to-Image Alignment TIFA: Accurate and Interpretable Text-to-Image Faithfulness Evaluation with Question Answering

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:21:21.500312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-22T09:19:39.848194Z digest=sha256:e125f88f8a36740514b08ec28ed4527bf597e19b415fc69c573f9f596c0f1496

Observation fd596e6c-e198-4a01-b9fb-ad3df36462e8 · inbound

The Illusion of High Utility in Safety Alignment of Text-to-Image Diffusion Models cites this paper.

The Illusion of High Utility in Safety Alignment of Text-to-Image Diffusion Models TIFA: Accurate and Interpretable Text-to-Image Faithfulness Evaluation with Question Answering

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T14:57:03.745541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-02T14:56:40.860766Z digest=sha256:9ddf21efed81e4280f1775176a8daaeec409a06a60ad550c9da4bca6884660ec

Observation b557d234-f48b-4829-8f35-10d73131098d · inbound

A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions cites this paper.

A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions TIFA: Accurate and Interpretable Text-to-Image Faithfulness Evaluation with Question Answering

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T00:03:54.867139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:03:54.867139Z digest=sha256:5359a5060968bb32afb37476af1ca60a30aebc6d7b378dc30b06bd9bb4e4d837

Observation 6dee82f8-7c2c-4f64-8a3c-7490ed143fdd · inbound

Poplar: A Scalable Pipeline for Human-Centric Image Dataset Synthesis cites this paper.

Poplar: A Scalable Pipeline for Human-Centric Image Dataset Synthesis TIFA: Accurate and Interpretable Text-to-Image Faithfulness Evaluation with Question Answering

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T01:00:04.666098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T01:00:04.666098Z digest=sha256:806355d163eb521605911aa2939436801ebe94bf67dc787cf95887a4e92b9fff

Observation b279b81b-e9aa-4d95-8231-9e132b100059 · inbound

Can Text-to-Image Models Draw from the Right Frame of Reference? cites this paper.

Can Text-to-Image Models Draw from the Right Frame of Reference? TIFA: Accurate and Interpretable Text-to-Image Faithfulness Evaluation with Question Answering

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T14:59:46.915311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:59:46.915311Z digest=sha256:0578374075979f68c6bdad51ca9ab92b9e1b80b02495cce2a4055a37e5bffe96