Pith. sign in

REVIEW 1 cited by

A Generative Approach for Wikipedia-Scale Visual Entity Recognition

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.02041 v2 pith:HI2M62Y2 submitted 2024-03-04 cs.CV

classification cs.CV
keywords entityrecognitiongivenimagevisualcaptioningdual-encodergenerative
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this paper, we address web-scale visual entity recognition, specifically the task of mapping a given query image to one of the 6 million existing entities in Wikipedia. One way of approaching a problem of such scale is using dual-encoder models (eg CLIP), where all the entity names and query images are embedded into a unified space, paving the way for an approximate k-NN search. Alternatively, it is also possible to re-purpose a captioning model to directly generate the entity names for a given image. In contrast, we introduce a novel Generative Entity Recognition (GER) framework, which given an input image learns to auto-regressively decode a semantic and discriminative ``code'' identifying the target entity. Our experiments demonstrate the efficacy of this GER paradigm, showcasing state-of-the-art performance on the challenging OVEN benchmark. GER surpasses strong captioning, dual-encoder, visual matching and hierarchical classification baselines, affirming its advantage in tackling the complexities of web-scale recognition.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Reverse Region-to-Entity Annotation for Pixel-Level Visual Entity Linking

    cs.CV 2024-12 conditional novelty 7.0 of 10

    Introduces PL-VEL, a pixel-mask-based visual entity linking task, and MaskOVEN-Wiki, a 5.2M-annotation dataset built via reverse annotation, plus a semantic tokenization method that yields a 5-point accuracy gain.

Pith tools