Pith. sign in

REVIEW 4 cited by

TextCaps: a Dataset for Image Captioning with Reading Comprehension

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2003.12462 v2 pith:IE4QB2EO submitted 2020-03-24 cs.CV cs.CL

classification cs.CVcs.CL
keywords textimagedatasettextcapsvisualapproachescaptioningchallenges
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Image descriptions can help visually impaired people to quickly understand the image content. While we made significant progress in automatically describing images and optical character recognition, current approaches are unable to include written text in their descriptions, although text is omnipresent in human environments and frequently critical to understand our surroundings. To study how to comprehend text in the context of an image we collect a novel dataset, TextCaps, with 145k captions for 28k images. Our dataset challenges a model to recognize text, relate it to its visual context, and decide what part of the text to copy or paraphrase, requiring spatial, semantic, and visual reasoning between multiple text tokens and visual entities, such as objects. We study baselines and adapt existing approaches to this new task, which we refer to as image captioning with reading comprehension. Our analysis with automatic and human studies shows that our new TextCaps dataset provides many new technical challenges over previous datasets.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DODO: Discrete OCR Diffusion Models

    cs.CV 2026-02 conditional novelty 6.0 of 10

    Block-based discrete diffusion can transcribe documents in parallel, roughly matching autoregressive OCR accuracy while cutting inference time by up to about 3x in a lower-accuracy fast variant.

  2. Pixel-Level Reasoning Segmentation via Multi-turn Conversations

    cs.CV 2025-02 conditional novelty 6.0 of 10

    PRIST, a benchmark for pixel-level segmentation through multi-turn conversations, and the MIRAS model achieve the best reported scores on this new task.

  3. Sealing The Backdoor: Unlearning Adversarial Text Triggers In Diffusion Models Using Knowledge Distillation

    cs.CV 2025-08 unverdicted novelty 5.0 of 10

    SKD-CAG erases adversarial text triggers from diffusion models by distilling the model's own clean outputs through cross-attention guidance, claiming 100% and 93% removal for pixel and style backdoors.

  4. TextPixs: Glyph-Conditioned Diffusion with Character-Aware Attention and OCR-Guided Supervision

    cs.CV 2025-07 reject novelty 5.0 of 10

    The GCDA framework claims state-of-the-art text rendering in diffusion images via dual-stream encoding, attention segregation, and OCR supervision, but the paper lacks verifiable artifacts and contains internal incons...

Pith tools