Pith. sign in

REVIEW 1 cited by

InstructDiffusion: A Generalist Modeling Interface for Vision Tasks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.03895 v1 pith:BQXYX4JB submitted 2023-09-07 cs.CV

InstructDiffusion: A Generalist Modeling Interface for Vision Tasks

classification cs.CV
keywords tasksvisioninstructdiffusionspacecomputergeneralisthandleinstructions
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

We present InstructDiffusion, a unifying and generic framework for aligning computer vision tasks with human instructions. Unlike existing approaches that integrate prior knowledge and pre-define the output space (e.g., categories and coordinates) for each vision task, we cast diverse vision tasks into a human-intuitive image-manipulating process whose output space is a flexible and interactive pixel space. Concretely, the model is built upon the diffusion process and is trained to predict pixels according to user instructions, such as encircling the man's left shoulder in red or applying a blue mask to the left car. InstructDiffusion could handle a variety of vision tasks, including understanding tasks (such as segmentation and keypoint detection) and generative tasks (such as editing and enhancement). It even exhibits the ability to handle unseen tasks and outperforms prior methods on novel datasets. This represents a significant step towards a generalist modeling interface for vision tasks, advancing artificial general intelligence in the field of computer vision.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Hidden-Shot: Towards One-Shot Task Generalization for Low-Level Vision Generalist Models

    cs.CV 2026-07 unverdicted novelty 5.0

    Hidden-Shot adds an implicit visual-task prompt and selective merging step to existing low-level vision generalist models, paired with a 3C4U/3C7U evaluation framework that reports outperformance on seven and ten data...