Pith. sign in

REVIEW 15 cited by

Q-Insight: Understanding Image Quality via Visual Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.22679 v2 pith:GHEZNCL4 submitted 2025-03-28 cs.CV

Q-Insight: Understanding Image Quality via Visual Reinforcement Learning

classification cs.CV
keywords imagequalitydegradationq-insighttasksperceptionreasoningunderstanding
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Image quality assessment (IQA) focuses on the perceptual visual quality of images, playing a crucial role in downstream tasks such as image reconstruction, compression, and generation. The rapid advancement of multi-modal large language models (MLLMs) has significantly broadened the scope of IQA, moving toward comprehensive image quality understanding that incorporates content analysis, degradation perception, and comparison reasoning beyond mere numerical scoring. Previous MLLM-based methods typically either generate numerical scores lacking interpretability or heavily rely on supervised fine-tuning (SFT) using large-scale annotated datasets to provide descriptive assessments, limiting their flexibility and applicability. In this paper, we propose Q-Insight, a reinforcement learning-based model built upon group relative policy optimization (GRPO), which demonstrates strong visual reasoning capability for image quality understanding while requiring only a limited amount of rating scores and degradation labels. By jointly optimizing score regression and degradation perception tasks with carefully designed reward functions, our approach effectively exploits their mutual benefits for enhanced performance. Extensive experiments demonstrate that Q-Insight substantially outperforms existing state-of-the-art methods in both score regression and degradation perception tasks, while exhibiting impressive zero-shot generalization to comparison reasoning tasks. Code will be available at https://github.com/lwq20020127/Q-Insight.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 15 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Beyond Absolute Scores: Relative Edit-induced Difference for Generalizable Image Aesthetic Assessment

    cs.CV 2026-06 unverdicted novelty 7.0

    RED-Aes learns aesthetic changes from edit-induced image pairs and a new RED-20k dataset via three-stage relative ranking training, claiming SOTA generalization over absolute MOS regression.

  2. HiTokSR: A Coarse-to-Fine Tokenizer with Hierarchical Codebooks for High-Fidelity Real-World Image Super-Resolution

    cs.CV 2026-05 unverdicted novelty 7.0

    HiTokSR uses a coarse-to-fine hierarchical tokenizer with frequency-aware sub-codebooks, vision foundation model priors, and index perturbation to achieve state-of-the-art perceptual quality and fidelity in real-world...

  3. DRM: Diffusion-based Reward Model With Step-wise Guidance

    cs.CV 2026-05 unverdicted novelty 7.0

    DRM turns a pre-trained diffusion model into a step-wise reward model and uses it for dense RL training (Step-wise GRPO) and guided sampling to improve final image quality.

  4. RevealLayer: Disentangling Hidden and Visible Layers via Occlusion-Aware Image Decomposition

    cs.CV 2026-05 unverdicted novelty 7.0

    RevealLayer decomposes natural images into multiple RGBA layers using diffusion models with region-aware attention, occlusion-guided adaptation, and a composite loss, outperforming prior methods on a new benchmark dataset.

  5. Q-Probe: Scaling Image Quality Assessment to High Resolution via Context-Aware Agentic Probing

    eess.IV 2026-01 unverdicted novelty 7.0

    Q-Probe introduces the first agentic IQA framework that scales to high resolutions using context-aware probing, a new Vista-Bench benchmark, and three-stage training to reach state-of-the-art performance across scales.

  6. Zoom-IQA: Image Quality Assessment with Reliable Region-Aware Reasoning

    cs.CV 2026-01 conditional novelty 7.0

    Zoom-IQA lets a vision-language model iteratively crop and zoom into image regions before giving a quality score, improving reasoning and restoration guidance over single-pass IQA models.

  7. Peak-End-Net: A Peak-End Rule Inspired Framework for Generalizable Video Aesthetic Assessment

    cs.CV 2026-07 conditional novelty 6.0

    Peak-End-Net uses image-aesthetic priors and peak-end-rule frame weighting to set a new state of the art on VADB and zero-shot DIVIDE-3K video aesthetic assessment.

  8. Beyond Absolute Scores: Relative Edit-induced Difference for Generalizable Image Aesthetic Assessment

    cs.CV 2026-06 unverdicted novelty 6.0

    RED-Aes uses controllable image editing to generate relative aesthetic difference pairs and a ranking consistency reward for training, achieving SOTA generalization on IAA benchmarks via the new RED-20k dataset.

  9. Ultra-High-Definition Image Quality Assessment via Graph Representation Learning

    cs.CV 2026-05 unverdicted novelty 6.0

    UHD-GCN-BIQA models structural dependencies among sampled patches via a hybrid kNN graph and residual graph convolutions to achieve competitive PLCC and SRCC with the lowest RMSE on the UHD-IQA benchmark for blind ult...

  10. Injecting Distributional Awareness into MLLMs via Reinforcement Learning for Deep Imbalanced Regression

    cs.CL 2026-05 unverdicted novelty 6.0

    A plug-and-play RL method adds batch-level distributional supervision via CCC rewards to reduce regression-to-the-mean in MLLMs on imbalanced regression benchmarks.

  11. Injecting Distributional Awareness into MLLMs via Reinforcement Learning for Deep Imbalanced Regression

    cs.CL 2026-05 unverdicted novelty 6.0

    A Group Relative Policy Optimization framework with concordance correlation coefficient rewards improves MLLM regression accuracy on long-tailed distributions, especially in medium- and few-shot regimes, without model...

  12. Panoptic Pairwise Distortion Graph

    cs.CV 2026-04 unverdicted novelty 6.0

    A new Distortion Graph task with PandaSet dataset and PandaBench benchmark allows region-wise distortion analysis in image pairs, where current MLLMs fail even with region cues.

  13. ME-IQA: Memory-Enhanced Image Quality Assessment via Re-Ranking

    cs.CV 2026-03 conditional novelty 6.0

    A test-time re-ranking framework that uses a memory of similar images and pairwise VLM comparisons to densify and improve reasoning-based image quality scores.

  14. EmoFeedback$^2$: Reinforcement of Continuous Emotional Image Generation via LVLM-based Reward and Textual Feedback

    cs.CV 2025-11 reject novelty 6.0

    A closed-loop emotional image generation system uses a fine-tuned vision-language model both as a reinforcement-learning reward and as an iterative prompt refiner, claiming improved valence-arousal fidelity.

  15. Q-DeepSight: Incentivizing Thinking with Images for Image Quality Assessment and Refinement

    cs.CV 2026-04 unverdicted novelty 5.0

    Q-DeepSight proposes a think-with-image multimodal CoT framework trained via RL with perceptual curriculum rewards and evidence gradient filtering to achieve SOTA IQA performance and enable training-free perceptual re...