Pith. sign in

REVIEW 4 cited by

Perspective (In)consistency of Paint by Text

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2206.14617 v1 pith:LKYIU33J submitted 2022-06-27 cs.GR cs.AIcs.CVcs.CY

classification cs.GRcs.AIcs.CVcs.CY
keywords imagesconsistencydall-e-2geometricpaintperspectivetextwill
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Type "a sea otter with a pearl earring by Johannes Vermeer" or "a photo of a teddy bear on a skateboard in Times Square" into OpenAI's DALL-E-2 paint-by-text synthesis engine and you will not be disappointed by the delightful and eerily pertinent results. The ability to synthesize highly realistic images -- with seemingly no limitation other than our imagination -- is sure to yield many exciting and creative applications. These images are also likely to pose new challenges to the photo-forensic community. Motivated by the fact that paint by text is not based on explicit geometric modeling, and the human visual system's often obliviousness to even glaring geometric inconsistencies, we provide an initial exploration of the perspective consistency of DALL-E-2 synthesized images to determine if geometric-based forensic analyses will prove fruitful in detecting this new breed of synthetic media.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Any-Resolution AI-Generated Image Detection by Spectral Learning

    cs.CV 2024-11 conditional novelty 6.0 of 10

    SPAI uses spectral reconstruction similarity from a frozen masked-frequency ViT plus attention pooling to reach 91.0 average AUC for AI-generated image detection across 13 unseen generators.

  2. Relations, Negations, and Numbers: Looking for Logic in Generative Text-to-Image Models

    cs.CV 2024-11 conditional novelty 4.0 of 10

    DALL-E 3 fails to reliably produce images matching prompts about physical relations, negations, and exact numbers beyond three, and a grounded diffusion pipeline does worse on relations.

  3. Survey on AI-Generated Media Detection: From Non-MLLM to MLLM

    cs.CV 2025-02 unverdicted novelty 3.0 of 10

    A survey organizing AI-generated media detection into Non-MLLM and MLLM based methods, with task and benchmark taxonomies.

  4. Artificial Intelligence for Geometry-Based Feature Extraction, Analysis and Synthesis in Artistic Images: A Survey

    cs.AI 2024-12 conditional novelty 2.0 of 10

    A survey reviewing how geometric features (bounding boxes, keypoints, poses, 3D representations) are used in AI for extracting, analyzing, and synthesizing artistic images, concluding that geometry improves performanc...

Pith tools