REVIEW 4 cited by
InstructIR: High-Quality Image Restoration Following Human Instructions
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Image restoration is a fundamental problem that involves recovering a high-quality clean image from its degraded observation. All-In-One image restoration models can effectively restore images from various types and levels of degradation using degradation-specific information as prompts to guide the restoration model. In this work, we present the first approach that uses human-written instructions to guide the image restoration model. Given natural language prompts, our model can recover high-quality images from their degraded counterparts, considering multiple degradation types. Our method, InstructIR, achieves state-of-the-art results on several restoration tasks including image denoising, deraining, deblurring, dehazing, and (low-light) image enhancement. InstructIR improves +1dB over previous all-in-one restoration methods. Moreover, our dataset and results represent a novel benchmark for new research on text-guided image restoration and enhancement. Our code, datasets and models are available at: https://github.com/mv-lab/InstructIR
Forward citations
Cited by 4 Pith papers
-
Uni-DocDiff: A Unified Document Restoration Model Based on Diffusion
A dual-stream diffusion model with a handcrafted prior pool and a prior fusion module unifies six document restoration tasks and matches task-specific specialists.
-
Grounding Degradations in Natural Language for All-In-One Video Restoration
RONIN distills per-frame language descriptions of video degradations into lightweight input-conditioned prompts, achieving all-in-one video restoration without any text encoder or MLLM at inference and outperforming p...
-
Olympus: A Universal Task Router for Computer Vision Tasks
Olympus is a trained MLLM router that delegates 20 vision tasks to specialist models and supports chain-of-action execution of up to five tasks per instruction.
-
LossAgent: Towards Any Optimization Objectives for Image Processing with LLM Agents
An LLM agent reads past loss weights and quality scores, then writes new loss weights, letting image processing models be trained toward non-differentiable objectives like IQA scores and text feedback.
Discussion (0). Continue with ORCID to comment.