A challenge report showing multi-modal models reach SROCC 0.710 in predicting short-video engagement continuation rate, beating a 0.660 baseline.
InternVQA: Advancing Compressed Video Quality Assessment with Distilling Large Foundation Model
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
Video quality assessment tasks rely heavily on the rich features required for video understanding, such as semantic information, texture, and temporal motion. The existing video foundational model, InternVideo2, has demonstrated strong potential in video understanding tasks due to its large parameter size and large-scale multimodal data pertaining. Building on this, we explored the transferability of InternVideo2 to video quality assessment under compression scenarios. To design a lightweight model suitable for this task, we proposed a distillation method to equip the smaller model with rich compression quality priors. Additionally, we examined the performance of different backbones during the distillation process. The results showed that, compared to other methods, our lightweight model distilled from InternVideo2 achieved excellent performance in compression video quality assessment.
citation-role summary
citation-polarity summary
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
VQualA 2025 Challenge on Engagement Prediction for Short Videos: Methods and Results
A challenge report showing multi-modal models reach SROCC 0.710 in predicting short-video engagement continuation rate, beating a 0.660 baseline.