Pith. sign in

REVIEW 2 cited by

IMProv: Inpainting-based Multimodal Prompting for Computer Vision Tasks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.01771 v1 pith:EFOEWMDN submitted 2023-12-04 cs.CV

classification cs.CV
keywords modelin-contexttasksvisioncomputerdatasetimprovlearning
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In-context learning allows adapting a model to new tasks given a task description at test time. In this paper, we present IMProv - a generative model that is able to in-context learn visual tasks from multimodal prompts. Given a textual description of a visual task (e.g. "Left: input image, Right: foreground segmentation"), a few input-output visual examples, or both, the model in-context learns to solve it for a new test input. We train a masked generative transformer on a new dataset of figures from computer vision papers and their associated captions, together with a captioned large-scale image-text dataset. During inference time, we prompt the model with text and/or image task example(s) and have the model inpaint the corresponding output. We show that training our model with text conditioning and scaling the dataset size improves in-context learning for computer vision tasks by over +10\% AP for Foreground Segmentation, over +5\% gains in AP for Single Object Detection, and almost 20\% lower LPIPS in Colorization. Our empirical results suggest that vision and language prompts are complementary and it is advantageous to use both to achieve better in-context learning performance. Project page is available at https://jerryxu.net/IMProv .

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Show Me Examples: Inferring Visual Concepts from Image Sets

    cs.CV 2026-07 unverdicted novelty 7.0 of 10

    Introduces VICIS task and training framework for inferring visual concepts from image sets, with experiments showing better accuracy, diversity, and generalization than standard VLMs on synthetic and ImageNet data.

  2. Stable Diffusion Models are Secretly Good at Visual In-Context Learning

    cs.CV 2025-08 conditional novelty 6.0 of 10

    A training-free attention recomputation inside Stable Diffusion self-attention enables visual in-context learning across six vision tasks.

Pith tools