Pith. sign in

REVIEW 8 cited by

Image Quality Assessment: Unifying Structure and Texture Similarity

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2004.07728 v3 pith:2KHTCDMI submitted 2020-04-16 cs.CV

classification cs.CV
keywords textureimagequalityhumanmethodsimilarityaveragescorrelations
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Objective measures of image quality generally operate by comparing pixels of a "degraded" image to those of the original. Relative to human observers, these measures are overly sensitive to resampling of texture regions (e.g., replacing one patch of grass with another). Here, we develop the first full-reference image quality model with explicit tolerance to texture resampling. Using a convolutional neural network, we construct an injective and differentiable function that transforms images to multi-scale overcomplete representations. We demonstrate empirically that the spatial averages of the feature maps in this representation capture texture appearance, in that they provide a set of sufficient statistical constraints to synthesize a wide variety of texture patterns. We then describe an image quality method that combines correlations of these spatial averages ("texture similarity") with correlations of the feature maps ("structure similarity"). The parameters of the proposed measure are jointly optimized to match human ratings of image quality, while minimizing the reported distances between subimages cropped from the same texture images. Experiments show that the optimized method explains human perceptual scores, both on conventional image quality databases, as well as on texture databases. The measure also offers competitive performance on related tasks such as texture classification and retrieval. Finally, we show that our method is relatively insensitive to geometric transformations (e.g., translation and dilation), without use of any specialized training or data augmentation. Code is available at https://github.com/dingkeyan93/DISTS.

Discussion (0). Sign in to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Your Demands Deserve More Bits: Referring Semantic Image Compression at Ultra-low Bitrate

    eess.IV 2025-05 conditional novelty 7.0 of 10

    RSIC allocates bits to user-specified image regions via a grounding model and guides a pretrained diffusion decoder with the compressed latent, boosting local fidelity at ultra-low rates.

  2. Online Neural Space Time Memory for Dynamic Novel View Synthesis

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Neural Space-Time Memory (NSTM) decouples low-frequency memory updates from per-frame synthesis with cross-view attention, enabling real-time minute-long dynamic novel view synthesis from multi-view streams.

  3. Style4D-Bench: A Benchmark Suite for 4D Stylization

    cs.CV 2025-08 conditional novelty 6.0 of 10

    Style4D-Bench introduces a 12-metric evaluation protocol and a 4DGS-based baseline, Style4D, claimed to achieve state-of-the-art 4D stylization.

  4. MicroZoom: Structure-Preserving Detail Synthesis at Extreme Scale

    cs.CV 2026-07 conditional novelty 5.0 of 10

    A cascaded, segmentation-conditioned, per-instance diffusion method synthesizes globally coherent gigapixel microscopic detail from a phone photo and sparse microscope references at up to 350×.

  5. Multi-Modal Learning meets Genetic Programming: Analyzing Alignment in Latent Space Optimization

    cs.NE 2026-04 unverdicted novelty 5.0 of 10

    SNIP's symbolic-numeric alignment stays coarse and does not improve during optimization, so multi-modal LSO does not yet deliver effective bi-modal search for symbolic regression.

  6. FasterVAR: Plug-and-Play Acceleration for Visual Autoregressive Models

    cs.CV 2025-12 conditional novelty 5.0 of 10

    Stage-aware pruning of late generation steps, using random projection and cached-feature restoration, speeds up VAR text-to-image models by up to 3.4x with minimal quality loss.

  7. Reverse Browser: Vector-Image-to-Code Generator

    cs.SE 2025-09 conditional novelty 5.0 of 10

    An open-weights system that turns vector images of web designs into HTML/CSS, with new datasets and a multi-scale pixel metric, though accuracy remains below production quality.

  8. Robustness as Architecture: Designing IQA Models to Withstand Adversarial Perturbations

    cs.CV 2025-06 conditional novelty 5.0 of 10

    An NR-IQA defense built from an FFT-domain orthogonal block, 10% pruning, and fine-tuning lowers adversarial AbsGain on some models with a modest SROCC decline, but the reported gains are mixed across architectures.

Pith tools