Pith. sign in

REVIEW 17 cited by

Controlling Vision-Language Models for Multi-Task Image Restoration

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.01018 v2 pith:LPZK7BAY submitted 2023-10-02 cs.CV

Controlling Vision-Language Models for Multi-Task Image Restoration

classification cs.CV
keywords imagerestorationvision-languagemodelsda-clipdegradationtasksclip
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Vision-language models such as CLIP have shown great impact on diverse downstream tasks for zero-shot or label-free predictions. However, when it comes to low-level vision such as image restoration their performance deteriorates dramatically due to corrupted inputs. In this paper, we present a degradation-aware vision-language model (DA-CLIP) to better transfer pretrained vision-language models to low-level vision tasks as a multi-task framework for image restoration. More specifically, DA-CLIP trains an additional controller that adapts the fixed CLIP image encoder to predict high-quality feature embeddings. By integrating the embedding into an image restoration network via cross-attention, we are able to pilot the model to learn a high-fidelity image reconstruction. The controller itself will also output a degradation feature that matches the real corruptions of the input, yielding a natural classifier for different degradation types. In addition, we construct a mixed degradation dataset with synthetic captions for DA-CLIP training. Our approach advances state-of-the-art performance on both \emph{degradation-specific} and \emph{unified} image restoration tasks, showing a promising direction of prompting image restoration with large-scale pretrained vision-language models. Our code is available at https://github.com/Algolzw/daclip-uir.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 17 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. LL-Bench: Rethinking Low-Level Vision Evaluation in the Era of Large-Scale Generative Models

    cs.CV 2026-06 unverdicted novelty 7.0

    LL-Bench supplies a human-annotated dataset exposing generative model weaknesses in low-level restoration and introduces LL-Score as an MLLM evaluator that outperforms existing quality metrics and can serve as a train...

  2. Degradation-Aware Adaptive Context Gating for Unified Image Restoration

    cs.CV 2026-05 unverdicted novelty 7.0

    DACG-IR adds a lightweight degradation-aware module that generates prompts to adaptively gate attention temperature, output features, and spatial-channel fusion in an encoder-decoder network for unified image restoration.

  3. PhySe-RPO: Physics and Semantics Guided Relative Policy Optimization for Diffusion-Based Surgical Smoke Removal

    cs.AI 2026-03 unverdicted novelty 7.0

    PhySe-RPO enables diffusion-based surgical smoke removal by converting restoration into a stochastic policy optimized with physics consistency and CLIP semantic rewards under limited supervision.

  4. Toward Generalizable Forgery Detection and Reasoning

    cs.CV 2025-03 unverdicted novelty 7.0

    FakeReasoning is an MLLM-based framework for unified forgery detection and reasoning on AI-generated images, supported by the new MMFR-Dataset of 120K images and 378K annotations across 10 generators.

  5. SpikeRestormer: Towards Energy-Efficient All-in-One Image Restoration via Unified Event Reasoning

    cs.CV 2026-08 conditional novelty 6.0

    A spiking neural network with subtractive and additive attention performs all-in-one image restoration in one time step, matching older ANN baselines with much lower estimated energy.

  6. SCRIPT: Scalable Diffusion Policy with Multi-stage Training for Language-driven Physics-based Humanoid Control

    cs.GR 2026-05 unverdicted novelty 6.0

    A new diffusion transformer policy with joint attention over actions, states, and text plus RL post-training outperforms prior methods on language alignment and motion quality for humanoid control.

  7. SCRIPT: Scalable Diffusion Policy with Multi-stage Training for Language-driven Physics-based Humanoid Control

    cs.GR 2026-05 unverdicted novelty 6.0

    SCRIPT presents a scalable diffusion policy with JAST-DiT architecture, nonlinear history conditioning, and RLHR post-training that claims to outperform prior methods on text alignment, motion quality, and physical re...

  8. EvoIR-Agent: Self-Evolving Image Restoration Agentic System via Experience-Driven Learning

    cs.CV 2026-05 unverdicted novelty 6.0

    EvoIR-Agent formulates experience components into a hierarchical pool with a self-evolving update mechanism to improve performance and efficiency of training-free MLLM image restoration agents over prior paradigms.

  9. Degradation Frequency Curve: An Explicit Frequency-Quantified Representation for All-in-One Image Restoration

    cs.CV 2026-05 unverdicted novelty 6.0

    The paper proposes the Degradation Frequency Curve (DFC) as an explicit spectral representation for quantifying degradations and develops a DFC-guided multi-scale restorer that achieves state-of-the-art performance on...

  10. TIR-Agent: Training an Explorative and Efficient Agent for Image Restoration

    cs.CV 2026-03 conditional novelty 6.0

    A vision-language agent trained with SFT plus RL, exploration-driven trajectory perturbation, and adaptive multi-metric rewards learns direct tool selection for composite image restoration, beating training-free agent...

  11. Adapting Large VLMs with Iterative and Manual Instructions for Generative Low-light Enhancement

    cs.CV 2025-07 conditional novelty 6.0

    VLM-IMI adapts VLMs with iterative and manual instructions plus a learnable fusion module to guide diffusion-based generative low-light image enhancement, outperforming prior methods in perceptual quality.

  12. DVANet: Degradation-aware Visual-prior Alignment Network for Image Restoration

    cs.CV 2026-06 unverdicted novelty 5.0

    DVANet proposes a deep unfolding network combining degradation representation with DINOv3 visual priors for unified restoration under complex degradations.

  13. EvoIR-Agent: Self-Evolving Image Restoration Agentic System via Experience-Driven Learning

    cs.CV 2026-05 unverdicted novelty 5.0

    EvoIR-Agent introduces a hierarchical experience pool and self-evolving mechanism to improve training-free image restoration agents, claiming significant metric leads and better performance-efficiency balance.

  14. AccelAes: Accelerating Diffusion Transformers for Training-Free Aesthetic-Enhanced Image Generation

    cs.CV 2026-03 unverdicted novelty 5.0

    Training-free AccelAes accelerates DiTs with aesthetic focus masks and step caches, reporting 2.11× speedup and +11.9% ImageReward on Lumina-Next.

  15. TPGDiff: Hierarchical Triple-Prior Guided Diffusion for Image Restoration

    cs.CV 2026-01 unverdicted novelty 5.0

    TPGDiff introduces hierarchical triple-prior guidance in a diffusion network, placing degradation priors throughout, structural priors in shallow layers, and semantic priors in deep layers for improved all-in-one imag...

  16. Bridging Information Asymmetry: A Hierarchical Framework for Deterministic Blind Face Restoration

    cs.CV 2026-01 conditional novelty 5.0

    Pref-Restore combines AR semantic tokens, a diffusion generator, and DiffusionNFT-style RL to make blind face restoration more consistent, but its deterministic-identity claim is weakened by self-referential rewards a...

  17. Q-Agent: Quality-Driven Chain-of-Thought Image Restoration Agent through Robust Multimodal Large Language Model

    eess.IV 2025-04 unverdicted novelty 5.0

    Q-Agent uses CoT decomposition on a fine-tuned MLLM for multi-degradation perception plus IQA-driven greedy selection of restoration algorithms to claim better performance than All-in-One IR models.