HiTokSR uses a coarse-to-fine hierarchical tokenizer with frequency-aware sub-codebooks, vision foundation model priors, and index perturbation to achieve state-of-the-art perceptual quality and fidelity in real-world image super-resolution.
Diff- bir: Towards blind image restoration with generative diffu- sion prior
6 Pith papers cite this work. Polarity classification is still indexing.
fields
cs.CV 6verdicts
UNVERDICTED 6representative citing papers
GGT-100K is a 103k-pair LQ-HQ dataset generated via MFMs to enhance real-world generalization of image restoration models.
VARestorer converts a text-to-image VAR model into a fast one-step real-world image super-resolution model via distribution matching distillation and pyramid image conditioning.
EAM is a DiT-based blind super-resolution model that uses a triple-flow Ψ-DiT block, progressive masked image modeling, and in-context subject-aware prompting to reach state-of-the-art quantitative and visual results on standard datasets.
GleSAM integrates latent diffusion into SAM and SAM2 to boost segmentation robustness on low-quality images using minimal extra parameters and a new LQSeg dataset.
GleSAM++ improves SAM robustness on degraded images by using generative enhancement, feature alignment, and adaptive degradation prediction while adding few parameters.
citing papers explorer
-
HiTokSR: A Coarse-to-Fine Tokenizer with Hierarchical Codebooks for High-Fidelity Real-World Image Super-Resolution
HiTokSR uses a coarse-to-fine hierarchical tokenizer with frequency-aware sub-codebooks, vision foundation model priors, and index perturbation to achieve state-of-the-art perceptual quality and fidelity in real-world image super-resolution.
-
GGT-100K: Generative Ground Truth for Generalizable Real-World Image Restoration
GGT-100K is a 103k-pair LQ-HQ dataset generated via MFMs to enhance real-world generalization of image restoration models.
-
VARestorer: One-Step VAR Distillation for Real-World Image Super-Resolution
VARestorer converts a text-to-image VAR model into a fast one-step real-world image super-resolution model via distribution matching distillation and pyramid image conditioning.
-
EAM: Enhancing Anything with Diffusion Transformers for Blind Super-Resolution
EAM is a DiT-based blind super-resolution model that uses a triple-flow Ψ-DiT block, progressive masked image modeling, and in-context subject-aware prompting to reach state-of-the-art quantitative and visual results on standard datasets.
-
Segment Any-Quality Images with Generative Latent Space Enhancement
GleSAM integrates latent diffusion into SAM and SAM2 to boost segmentation robustness on low-quality images using minimal extra parameters and a new LQSeg dataset.
-
Towards Any-Quality Image Segmentation via Generative and Adaptive Latent Space Enhancement
GleSAM++ improves SAM robustness on degraded images by using generative enhancement, feature alignment, and adaptive degradation prediction while adding few parameters.