REVIEW 11 cited by
Quality-Aware Image-Text Alignment for Opinion-Unaware Image Quality Assessment
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Quality-Aware Image-Text Alignment for Opinion-Unaware Image Quality Assessment
read the original abstract
No-Reference Image Quality Assessment (NR-IQA) focuses on designing methods to measure image quality in alignment with human perception when a high-quality reference image is unavailable. Most state-of-the-art NR-IQA approaches are opinion-aware, i.e. they require human annotations for training. This dependency limits their scalability and broad applicability. To overcome this limitation, we propose QualiCLIP (Quality-aware CLIP), a CLIP-based self-supervised opinion-unaware approach that does not require human opinions. In particular, we introduce a quality-aware image-text alignment strategy to make CLIP generate quality-aware image representations. Starting from pristine images, we synthetically degrade them with increasing levels of intensity. Then, we train CLIP to rank these degraded images based on their similarity to quality-related antonym text prompts. At the same time, we force CLIP to generate consistent representations for images with similar content and the same level of degradation. Our experiments show that the proposed method improves over existing opinion-unaware approaches across multiple datasets with diverse distortion types. Moreover, despite not requiring human annotations, QualiCLIP achieves excellent performance against supervised opinion-aware methods in cross-dataset experiments, thus demonstrating remarkable generalization capabilities. The code and the model are publicly available at https://github.com/miccunifi/QualiCLIP.
Forward citations
Cited by 11 Pith papers
-
Banana100: Breaking NR-IQA Metrics by 100 Iterative Image Replications with Nano Banana Pro
Banana100 dataset shows that none of 21 popular NR-IQA metrics consistently rate images degraded by 100 iterative edits lower than clean originals.
-
Impulse-to-Peak-Output Norm Optimal State-Feedback Control of Linear PDEs
The authors use partial integral equations and Lyapunov techniques to cast impulse-to-peak norm analysis as convex optimization and derive optimal state-feedback controllers for linear PDEs via strong duality.
-
A Distributional View for Visual Mechanistic Interpretability: KL-Minimal Soft-Constraint Principle
The work introduces a distributional view of visual mechanistic interpretability that casts the task as KL-minimal optimization and realizes it through a soft-constraint principle implemented with energy-guided diffus...
-
Viewport-Unaware Blind Omnidirectional Image Quality Assessment: A Unified and Generalized Approach
Blind omnidirectional image quality assessment reduces to standard 2D blind IQA by skipping viewport generation, yielding a unified model that accepts equirectangular inputs directly.
-
Impulse-to-Peak-Output Norm Optimal State-Feedback Control of Linear PDEs
Using PIE representations and Lyapunov LMIs with strong duality, the authors give provable I2P-norm bounds and constructive optimal state-feedback for linear PDEs.
-
Spectral and Trajectory Regularization for Diffusion Transformer Super-Resolution
Asymmetric adversarial distillation plus frequency distribution matching lets DiT models perform one-step Real-ISR without the grid-like periodic artifacts that plague prior one-step DiT distillations.
-
From Global to Granular: Revealing IQA Model Performance via Correlation Surface
GMC maps an IQA model's agreement with human scores across the quality-level and quality-difference landscape, exposing local strengths that global PLCC/SRCC hide.
-
Latent Wavelet Diffusion For Ultra-High-Resolution Image Synthesis
Latent Wavelet Diffusion uses wavelet energy map masking and a scale-consistent VAE to improve detail fidelity in 2K-4K image generation without extra inference overhead.
-
Generalizable Video Quality Assessment via Weak-to-Strong Learning
Self-supervised ranking-based training on a 10x larger unlabeled video dataset enables a VQA model to match supervised zero-shot performance, show strong OOD generalization, and set new SOTA when fine-tuned.
-
HiRQA: Hierarchical Ranking and Quality Alignment for Opinion-Unaware Image Quality Assessment
HiRQA is a self-supervised NR-IQA framework trained on synthetic distortions, using a higher-order ranking loss, embedding distance loss, and text-guided contrastive alignment, claimed to generalize to authentic distortions.
-
A Lightweight Ensemble-Based Face Image Quality Assessment Method with Correlation-Aware Loss
An ensemble of MobileNetV3-Small and ShuffleNetV2 with a correlation-aware loss and test-time augmentation reaches SRCC 0.9829 and PLCC 0.9894 on the VQualA FIQA validation set.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.