Pith. sign in

REVIEW 1 cited by

Foundation Models Boost Low-Level Perceptual Similarity Metrics

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.07650 v2 pith:5W7CMZG7 submitted 2024-09-11 cs.CV

classification cs.CV
keywords featuresmetricsmodelssimilarityfoundationimageintermediateperceptual
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

For full-reference image quality assessment (FR-IQA) using deep-learning approaches, the perceptual similarity score between a distorted image and a reference image is typically computed as a distance measure between features extracted from a pretrained CNN or more recently, a Transformer network. Often, these intermediate features require further fine-tuning or processing with additional neural network layers to align the final similarity scores with human judgments. So far, most IQA models based on foundation models have primarily relied on the final layer or the embedding for the quality score estimation. In contrast, this work explores the potential of utilizing the intermediate features of these foundation models, which have largely been unexplored so far in the design of low-level perceptual similarity metrics. We demonstrate that the intermediate features are comparatively more effective. Moreover, without requiring any training, these metrics can outperform both traditional and state-of-the-art learned metrics by utilizing distance measures between the features.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Latent Guidance in Diffusion Models for Perceptual Evaluations

    cs.CV 2025-05 conditional novelty 6.0 of 10

    LGDM uses perceptual guidance during Stable Diffusion sampling and aggregates multi-scale, multi-timestep U-Net features to predict human-rated image quality, achieving state-of-the-art correlations on ten NR-IQA datasets.

Pith tools