REVIEW 4 major objections 6 minor 63 references
Multi-Branch Collaborative Learning Network for Video Quality Assessment in Industrial Video Search
T0 review · 4 major / 6 minor · reviewed 2026-08-08 · deepseek-v4-flash
Pith's one-line read The paper claims that a four-branch quality scorer with learned aggregation identifies low-quality videos more accurately than the deployed monolithic CLIP model and thereby improves industrial video search ranking.
desk verdict A credible industrial VQA paper with a sensible four-branch decomposition and consistent offline evidence, but the online A/B testing is too thinly described to verify the headline gain. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the Multi-Branch Collaborative learning Network (MBCN), a four-branch quality scorer built on a shared multimodal encoder consisting of a Chinese BERT text encoder, a CLIP-initialized ViT frame encoder, and a two-layer temporal transformer. Each branch produces an independent score: the Video-Text Matching Assessment Branch uses multi-grained image-text similarity; the Frame Coherence Assessment Branch uses inter-frame differences and frame-versus-theme differences; the Frame Quality Assessment Branch uses an MLP on the video representation to catch mosaics, black boxes, and watermarks; and the Text Quality Assessment Branch uses an MLP on the text representation for title and OCR problems. The scores are then aggregated by a squeeze-and-excitation module that learns content-dependent weights, and training combines a pointwise squared loss with a pairwise margin loss to keep scores stable and discriminative.
What would settle it
Run a controlled A/B test with traffic randomly split and the rest of the ranking pipeline frozen: if replacing the baseline scorer with MBCN yields a GSB difference statistically indistinguishable from zero, or if the offline PNR/AUC gains disappear on a fresh held-out sample collected after deployment, the central claim that MBCN causes the ranking improvement would be falsified.
Extended reading notes
Core claim
The paper's central claim is that MBCN identifies low-quality videos better than the currently deployed single-model baseline, and that this identification directly improves the search engine's ranking. The model formalizes industrial low-quality video characteristics into four categories: visual low-level problems such as mosaics, black boxes, and watermarks; textual problems in titles and OCR content; semantic frame incoherence; and frame-text mismatch. Each category gets one assessment branch, and the four branch scores are aggregated by a squeeze-and-excitation mechanism that learns content-dependent weights. Training combines pointwise and pairwise losses so that scores are stable, calibrated, and discriminative. The reported evidence is offline PNR rising from 5.380 to 5.710 and AUC from 0.829 to 0.855; online gains of +7.00% in GSB and +3.55%/+4.18% in DCG2/DCG4; ablations showing every branch contributes; and on 200 manually selected low-quality AI-generated videos, accuracy rising from 84.0% to 93.5% at threshold 0.5 while the mean predicted score falls from 0.401 to 0.226.
Load-bearing premise
The online experiments attribute the +7.00% GSB and DCG gains entirely to replacing the baseline quality model with MBCN, but the deployment section does not describe how traffic was split, whether other ranking modules were held fixed, or how the two systems were matched, so uncontrolled deployment differences could explain part of the improvement.
Editorial extensions
If this is right
- Deployed video search quality scoring can be decomposed by failure mode: four specialized branches each contribute positively, and removing any one of them measurably lowers offline and online performance.
- The quality score is more discriminative than the baseline: the mean-score gap between 'good' and 'excellent' videos widens from 0.09 to 0.26, so high-quality videos are separated more cleanly from mediocre ones.
- Low-quality AI-generated videos are downgraded more aggressively: on the paper's 200-video subset, the mean predicted score drops from 0.401 to 0.226, and recognition accuracy rises by 9.5 to 24 percentage points depending on threshold.
- Because the final score is trained with pointwise plus pairwise objectives, it is designed to coexist with relevance, freshness, and authority signals in a real ranking pipeline.
- The method's offline gains translate to online ranking improvements under human evaluation, reported as a statistically significant +7.00% GSB shift over the deployed baseline.
Reading between the lines
- We infer that the four-category taxonomy likely transfers to other video surfaces such as short-video feeds, video ads, and user-generated-content platforms, with the squeeze-and-excitation weights re-learned for each domain.
- A testable extension would be to expose the four branch logits as explanations for demotion decisions, since each branch isolates a distinct failure mode and the ablation results show the branches are not redundant.
- The paper's AI-generated-video result rests on 200 manually selected clips; we infer that a larger, randomized sample of AI-generated videos would be needed to know how far the 93.5% accuracy generalizes.
- The authors do not report a sweep over the loss weight $\alpha$ or the ranking margin $\tau$; we infer that the stability of the offline and online gains across those hyperparameters remains an open quantitative question.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper addresses no-reference video quality assessment in an industrial video search engine. It taxonomizes low-quality videos into four categories (visual low-level problems, textual low-level problems, frame incoherence, and frame-text mismatch) and proposes MBCN, which encodes text and frames, computes four branch scores, aggregates them with a squeeze-and-excitation module, and is trained with a weighted pointwise and pairwise loss. Offline experiments on a proprietary training set of 110,502 videos and a test set of 15,068 videos report PNR and AUC gains over a CLIP-based baseline and over branch-removal variants. Online deployment results report gains in GSB and DCG, and a manually selected set of 200 AI-generated videos is used to evaluate performance on that emerging category.
Significance. The paper's four-way taxonomy of industrial video quality issues is a useful practical contribution that is largely absent from academic VQA benchmarks. The branch-level analyses in Tables 3 and 4, together with the ablation results, give credible evidence that each branch contributes something on this proprietary distribution. If the deployment effects are causal, the online results would be valuable for the applied community. However, the paper's central comparison confounds architecture with loss function, the online experiment lacks a described control protocol, and no variance information is provided, so the strength of the claims is currently conditional on fixing these issues.
major comments (4)
- [3.3 and 3.7.2, Table 1] The main MBCN-versus-Base comparison is not controlled: Section 3.7.2 states that the baseline uses only pairwise loss while MBCN uses a mixture of pointwise and pairwise losses (Eq. 11). The reported offline gains in Table 1 (+6.12% PNR, +3.13% AUC) could therefore be due to the loss change rather than to the four-branch architecture. The branch ablations share MBCN's loss and do support a branch contribution, but the central claim that a multi-branch network is better than a monolithic CLIP-based scorer needs a baseline trained with the same loss, or an MBCN variant trained with pairwise loss only.
- [3.6, Table 2] The online experiment is not interpretable as causal evidence for MBCN. Section 3.6 says only that MBCN was deployed and compared with the baseline in the real production environment; it does not specify the traffic split or randomization unit, the duration and query population, or whether non-quality ranking modules such as relevance, freshness, and authority were held fixed. Because the GSB annotation instructions ask judges to compare both relevance and quality, a concurrent change in any ranking signal or a shift in the query mix across the compared periods would move GSB and DCG in the reported direction. The t-test footnote establishes only that a difference is unlikely to be zero, not that MBCN caused it. The paper should provide the experiment protocol or explicitly downgrade the online claim to a deployment observation.
- [Table 1] No variance information is reported for the offline metrics. The test set is large, but a single run gives no indication of run-to-run stability; the paper should report means and standard deviations across at least three seeds, or a paired significance test, so that the headline PNR and AUC gains can be assessed.
- [3.7.3, Table 5] The AI-generated video evaluation uses a manually selected set of 200 low-quality videos and thresholds t=0.5 and t=0.4 chosen by the authors. Accuracy at a fixed threshold is not a threshold-free measure, and the sampling procedure is not described. Please report how the videos were selected, the label distribution, precision/recall or ROC over a range of thresholds, and confidence intervals; as written, the claim that MBCN significantly improves recognition accuracy on AI-generated videos is not adequately supported.
minor comments (6)
- [Eq. (8)] The squeeze operation [˜v_i; ˜t_i] W_S is dimensionally unclear as written: if [˜v_i; ˜t_i] is a vertical concatenation of two d-dimensional vectors, W_S should be in R^{2d×1} rather than R^{d×1}; if a horizontal concatenation is intended, the notation and dimensions should be corrected.
- [3.3, Table 1] The asterisk denotes fine-tuning on the authors' training set, but no hyperparameters, optimization steps, or stopping criteria are given for the COVER, DOVER, and X-CLIP baselines, so a reader cannot tell whether the comparisons are at comparable optimization strength.
- [2.5] No sensitivity analysis is reported for the hyper-parameters alpha, tau, and the soft-label mapping {0, 0.3, 0.6, 1}; since the score distribution in Figure 4 is used as evidence, a robustness check for these choices would strengthen the claims.
- [3.2, Eq. (12)] The PNR formula does not state how tied labels or tied predictions are handled, and the denominator can be zero in small lists; please specify the tie-handling convention used in the offline computation.
- [1] There is a typo in the first paragraph of Section 1: 'Visual-related low-leval' should be 'low-level'.
- [References] The reference list contains a typo in the title of [18] ('Dateset' for 'Dataset'), and the names 'DOVER', 'Dover', and 'COVER' are used inconsistently across the text and reference list.
Circularity Check
No significant circularity: the model is trained against human quality labels and evaluated on held-out data; no branch definition or aggregation reduces to the target metric.
full rationale
The paper's derivation chain is supervised and empirically evaluated rather than definitionally self-referential. MBCN's four branches produce intermediate scores (VTMAB, FCAB, FQAB, TQAB) from learned representations, and the final score is a learned squeeze-and-excitation aggregation of those branch scores. The labels {bad, fair, good, excellent} come from external human annotation, and the offline PNR/AUC metrics are computed on a held-out test set (110,502 training videos vs. 15,068 test videos). No branch score or aggregation weight is defined in terms of the final prediction, and no prediction metric is used as a fitting target. The use of CLIP, BERT, and SENet is standard external prior work, not a self-citation chain, and no uniqueness claim or ansatz is imported from the authors' own prior papers. The 'predicted scores are more discriminative' observation is consistent with the pairwise loss objective, but this is an empirical consequence of the chosen loss, not a fitted parameter renamed as a prediction. The main weakness—online GSB/DCG gains without reported traffic splitting or isolation of the quality module—is a threat to experimental validity and causal interpretation, not a circularity of the derivation. Hence the paper has no circular step under the specified criteria.
Assumptions & free parameters
free parameters (4)
- alpha (loss weight) =
0.5
- Soft label mapping =
{bad:0, fair:0.3, good:0.6, excellent:1}
- AI-generated evaluation thresholds (t) =
0.5 and 0.4
- Margin tau in pairwise loss =
not reported
assumptions (4)
- domain assumption Crowd-sourced human labels on Baidu's platform are reliable and consistent enough to train and evaluate a quality model.
- domain assumption The four quality categories (visual low-level, text quality, frame coherence, video-text match) are sufficient and roughly independent for industrial video quality.
- domain assumption Pretrained CLIP, ViT, and BERT representations transfer to quality assessment with encoders frozen.
- domain assumption Inter-frame differences in embedding space (Eq. 4) capture semantic frame coherence.
Cite this review
Pith. "Pith review of Multi-Branch Collaborative Learning Network for Video Quality Assessment in Industrial Video Search." pith.science (2026). https://pith.science/paper/PFNNOEJL
@misc{pith2026250205924,
author = {Pith},
title = {Pith review of: Multi-Branch Collaborative Learning Network for Video Quality Assessment in Industrial Video Search},
year = {2026},
howpublished = {\url{https://pith.science/paper/PFNNOEJL}},
note = {Machine review of arXiv:2502.05924}
}
read the original abstract
Video Quality Assessment (VQA) is vital for large-scale video retrieval systems, aimed at identifying quality issues to prioritize high-quality videos. In industrial systems, low-quality video characteristics fall into four categories: visual-related issues like mosaics and black boxes, textual issues from video titles and OCR content, and semantic issues like frame incoherence and frame-text mismatch from AI-generated videos. Despite their prevalence in industrial settings, these low-quality videos have been largely overlooked in academic research, posing a challenge for accurate identification. To address this, we introduce the Multi-Branch Collaborative Network (MBCN) tailored for industrial video retrieval systems. MBCN features four branches, each designed to tackle one of the aforementioned quality issues. After each branch independently scores videos, we aggregate these scores using a weighted approach and a squeeze-and-excitation mechanism to dynamically address quality issues across different scenarios. We implement point-wise and pair-wise optimization objectives to ensure score stability and reasonableness. Extensive offline and online experiments on a world-level video search engine demonstrate MBCN's effectiveness in identifying video quality issues, significantly enhancing the retrieval system's ranking performance. Detailed experimental analyses confirm the positive contribution of all four evaluation branches. Furthermore, MBCN significantly improves recognition accuracy for low-quality AI-generated videos compared to the baseline.
Figures
Reference graph
Works this paper leans on
-
[1]
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Floren- cia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774 (2023)
arXiv 2023
-
[2]
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh. 2015. Vqa: Visual question answering. In Proceedings of the IEEE international conference on computer vision . 2425–2433
2015
-
[3]
Tim Brooks, Bill Peebles, Connor Holmes, Will DePue, Yufei Guo, Li Jing, David Schnurr, Joe Taylor, Troy Luhman, Eric Luhman, Clarence Ng, Ricky Wang, and Aditya Ramesh. 2024. Video generation models as world simulators. (2024). https://openai.com/research/video-generation-models-as-world-simulators
2024
-
[4]
Joao Carreira and Andrew Zisserman. 2017. Quo vadis, action recognition? a new model and the kinetics dataset. In proceedings of the IEEE Conference on Computer Vision and Pattern Recognition . 6299–6308
2017
-
[5]
Minsuk Chang, Anh Truong, Oliver Wang, Maneesh Agrawala, and Juho Kim
-
[6]
Shyamprasad Chikkerur, Vijay Sundaram, Martin Reisslein, and Lina J Karam
-
[7]
Ohad Fried, Ayush Tewari, Michael Zollhöfer, Adam Finkelstein, Eli Shecht- man, Dan B Goldman, Kyle Genova, Zeyu Jin, Christian Theobalt, and Maneesh Agrawala. 2019. Text-based editing of talking-head video. ACM Transactions on Graphics (TOG) 38, 4 (2019), 1–14
2019
-
[8]
Chenlong He, Qi Zheng, Ruoxi Zhu, Xiaoyang Zeng, Yibo Fan, and Zhengzhong Tu. 2024. COVER: A comprehensive video quality evaluator. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 5799–5809
work page 2024
Show all 63 references
-
[9]
Jack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras, and Yejin Choi. 2021. Clipscore: A reference-free evaluation metric for image captioning.arXiv preprint arXiv:2104.08718 (2021)
2021 arXiv
-
[10]
Jonathan Ho, Ajay Jain, and Pieter Abbeel. 2020. Denoising diffusion probabilistic models. Advances in neural information processing systems 33 (2020), 6840–6851
2020
-
[11]
Vlad Hosu, Franz Hahn, Mohsen Jenadeleh, Hanhe Lin, Hui Men, Tamás Szirányi, Shujun Li, and Dietmar Saupe. 2017. The Konstanz natural video database (KoNViD-1k). In 2017 Ninth international conference on quality of multimedia experience (QoMEX). IEEE, 1–6
2017
-
[12]
Jie Hu, Li Shen, and Gang Sun. 2018. Squeeze-and-excitation networks. InProceed- ings of the IEEE conference on computer vision and pattern recognition . 7132–7141
2018
-
[13]
Ziqi Huang, Yinan He, Jiashuo Yu, Fan Zhang, Chenyang Si, Yuming Jiang, Yuan- han Zhang, Tianxing Wu, Qingyang Jin, Nattapol Chanpaisit, et al. 2024. Vbench: Comprehensive benchmark suite for video generative models. In Proceedings of the IEEE/CVF Conference on Computer Vision...
2024
-
[14]
Bernd Huber, Hijung Valentina Shin, Bryan Russell, Oliver Wang, and Gautham J Mysore. 2019. B-script: Transcript-based b-roll video editing with recommenda- tions. In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems. 1–11
2019
-
[15]
Mina Huh, Saelyne Yang, Yi-Hao Peng, Xiang’Anthony’ Chen, Young-Ho Kim, and Amy Pavel. 2023. AVscript: Accessible Video Editing with Audio-Visual Scripts. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems. 1–17
2023
-
[16]
Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig. 2021. Scaling up visual and vision- language representation learning with noisy text supervision. In International conference on machine learning . PMLR, 4904–4916
2021
-
[17]
Jari Korhonen. 2019. Two-Level Approach for No-Reference Consumer Video Quality Assessment. IEEE Transactions on Image Processing (Dec 2019), 5923–5938. https://doi.org/10.1109/tip.2019.2923051
2019
-
[18]
Tengchuan Kou, Xiaohong Liu, Zicheng Zhang, Chunyi Li, Haoning Wu, Xiongkuo Min, Guangtao Zhai, and Ning Liu. 2024. Subjective-Aligned Dateset KDD ’25, August 3–7, 2025, Toronto, ON, Canada Hengzhu Tang et al. and Metric for Text-to-Video Quality Assessment.arXiv preprint arXi...
2024 arXiv
-
[19]
Gierad P Laput, Mira Dontcheva, Gregg Wilensky, Walter Chang, Aseem Agar- wala, Jason Linder, and Eytan Adar. 2013. Pixeltone: A multimodal interface for image editing. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems. 2185–2194
2013
-
[20]
Dingquan Li, Tingting Jiang, and Ming Jiang. 2019. Quality Assessment of In- the-Wild Videos. In Proceedings of the 27th ACM International Conference on Multimedia. https://doi.org/10.1145/3343031.3351028
2019
-
[21]
Dingquan Li, Tingting Jiang, and Ming Jiang. 2021. Unified Quality Assessment of in-the-Wild Videos with Mixed Datasets Training. International Journal of Computer Vision (Apr 2021), 1238–1257. https://doi.org/10.1007/s11263-020- 01408-w
2021 doi
-
[22]
Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi. 2022. Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation. In International conference on machine learning . PMLR, 12888–12900
2022
-
[23]
Junnan Li, Ramprasaath Selvaraju, Akhilesh Gotmare, Shafiq Joty, Caiming Xiong, and Steven Chu Hong Hoi. 2021. Align before fuse: Vision and language repre- sentation learning with momentum distillation. Advances in neural information processing systems 34 (2021), 9694–9705
2021
-
[24]
Liunian Harold Li, Mark Yatskar, Da Yin, Cho-Jui Hsieh, and Kai-Wei Chang
-
[25]
Georgianna Lin, Jin Yi Li, Afsaneh Fazly, Vladimir Pavlovic, and Khai Truong
-
[26]
Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. 2019. Vilbert: Pretrain- ing task-agnostic visiolinguistic representations for vision-and-language tasks. Advances in neural information processing systems 32 (2019)
2019
-
[27]
arXiv preprint arXiv:1908.03557 (2019)
Visualbert: A simple and performant baseline for vision and language. arXiv preprint arXiv:1908.03557 (2019)
2019 arXiv
-
[28]
Saad, and Alan C
Anish Mittal, Michele A. Saad, and Alan C. Bovik. 2016. A Completely Blind Video Integrity Oracle. IEEE Transactions on Image Processing (Jan 2016), 289–300. https://doi.org/10.1109/tip.2015.2502725
2016
-
[29]
Ron Mokady, Amir Hertz, and Amit H Bermano. 2021. Clipcap: Clip prefix for image captioning. arXiv preprint arXiv:2111.09734 (2021)
2021 arXiv
-
[30]
Bolin Ni, Houwen Peng, Minghao Chen, Songyang Zhang, Gaofeng Meng, Jian- long Fu, Shiming Xiang, and Haibin Ling. 2022. Expanding language-image pretrained models for general video recognition. In European Conference on Computer Vision. Springer, 1–18
2022
-
[31]
Yiting Lu, Xin Li, Bingchen Li, Zihao Yu, Fengbin Guan, Xinrui Wang, Ruling Liao, Yan Ye, and Zhibo Chen. 2024. AIGC-VQA: A Holistic Perception Metric for AIGC Video Quality Assessment. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 6384–6394
2024
-
[32]
Amy Pavel, Gabriel Reyes, and Jeffrey P Bigham. 2020. Rescribe: Authoring and automatically editing audio descriptions. In Proceedings of the 33rd Annual ACM Symposium on User Interface Software and Technology . 747–759
2020
-
[33]
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. In International conference on machine learni...
2021
-
[34]
Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever. 2021. Zero-shot text-to-image generation. In International conference on machine learning . Pmlr, 8821–8831
2021
-
[35]
Junting Pan, Ziyi Lin, Yuying Ge, Xiatian Zhu, Renrui Zhang, Yi Wang, Yu Qiao, and Hongsheng Li. 2023. Retrieving-to-answer: Zero-shot video question answering with frozen large language models. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 272–283
2023
-
[36]
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2015. Neural machine translation of rare words with subword units. arXiv preprint arXiv:1508.07909 (2015)
2015 arXiv
-
[37]
Zeina Sinno and Alan Conrad Bovik. 2018. Large-scale study of perceptual video quality. IEEE Transactions on Image Processing 28, 2 (2018), 612–627
2018
-
[38]
Haoyu Song, Li Dong, Wei-Nan Zhang, Ting Liu, and Furu Wei. 2022. Clip models are few-shot learners: Empirical studies on vqa and visual entailment. arXiv preprint arXiv:2203.07190 (2022)
2022 arXiv
-
[39]
Saad, Alan C
Michele A. Saad, Alan C. Bovik, and Christophe Charrier. 2014. Blind Prediction of Natural Video Quality. IEEE Transactions on Image Processing (Mar 2014), 1352–1365. https://doi.org/10.1109/tip.2014.2299154
2014
-
[40]
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971 (2023)
2023 arXiv
-
[41]
Zhengzhong Tu, Yilin Wang, Neil Birkbeck, Balu Adsumilli, and Alan C. Bovik
-
[42]
Thomas Unterthiner, Sjoerd van Steenkiste, Karol Kurach, Raphaël Marinier, Marcin Michalski, and Sylvain Gelly. 2019. FVD: A new metric for video genera- tion. (2019)
2019
-
[43]
Matteo Stefanini, Marcella Cornia, Lorenzo Baraldi, Silvia Cascianelli, Giuseppe Fiameni, and Rita Cucchiara. 2022. From show to tell: A survey on deep learning- based image captioning. IEEE transactions on pattern analysis and machine intelligence 45, 1 (2022), 539–559
2022
-
[44]
Bryan Wang, Yuliang Li, Zhaoyang Lv, Haijun Xia, Yan Xu, and Raj Sodhi. 2024. LAVE: LLM-Powered Agent Assistance and Language Augmentation for Video Editing. In Proceedings of the 29th International Conference on Intelligent User Interfaces. 699–714
2024
-
[45]
Yilin Wang, Sasi Inguva, and Balu Adsumilli. 2019. YouTube UGC dataset for video compression research. In 2019 IEEE 21st International Workshop on Multimedia Signal Processing (MMSP). IEEE, 1–5
2019
-
[46]
Yilin Wang, Sasi Inguva, and Balu Adsumilli. 2019. YouTube UGC Dataset for Video Compression Research. In 2019 IEEE 21st International Workshop on Multi- media Signal Processing (MMSP) . https://doi.org/10.1109/mmsp.2019.8901772
2019
-
[47]
Haoning Wu, Chaofeng Chen, Liang Liao, Jingwen Hou, Wenxiu Sun, Qiong Yan, and Weisi Lin. 2022. DisCoVQA: Temporal Distortion-Content Transformers for Video Quality Assessment. (Jun 2022)
2022
-
[48]
Oriol Vinyals, Alexander Toshev, Samy Bengio, and Dumitru Erhan. 2015. Show and tell: A neural image caption generator. In Proceedings of the IEEE conference on computer vision and pattern recognition . 3156–3164
2015
-
[49]
Haoning Wu, Erli Zhang, Liang Liao, Chaofeng Chen, Jingwen Hou, Annan Wang, Wenxiu Sun, Qiong Yan, and Weisi Lin. 2023. Exploring video quality assessment on user generated contents from aesthetic and technical perspectives. In Proceedings of the IEEE/CVF International Confere...
2023
-
[50]
Haoning Wu, Erli Zhang, Liang Liao, Chaofeng Chen, Jingwen Hou, Annan Wang, Wenxiu Sun, Qiong Yan, and Weisi Lin. 2023. Towards explainable in-the-wild video quality assessment: a database and a language-prompted approach. In Proceedings of the 31st ACM International Conferenc...
2023
-
[51]
Xinyu Xia, Guohua Dong, Fengling Li, Lei Zhu, and Xiaomin Ying. 2023. When CLIP meets cross-modal hashing retrieval: A new strong baseline. Information Fusion 100 (2023), 101968
2023
-
[52]
Fengchuang Xing, Mingjie Li, Yuan-Gen Wang, Guopu Zhu, and Xiaochun Cao. 2024. CLIPVQA: Video Quality Assessment via CLIP. arXiv preprint arXiv:2407.04928 (2024)
2024 arXiv
-
[53]
Haoning Wu, Liang Liao, Chaofeng Chen, Jingwen Hou, Annan Wang, Wenxiu Sun, Qiong Yan, and Weisi Lin. 2022. Disentangling aesthetic and technical effects for video quality assessment of user generated content. arXiv preprint arXiv:2211.04894 2, 5 (2022), 6
2022 arXiv
-
[54]
Zhenqiang Ying, Mandal Maniratnam, Deepti Ghadiyaram, and AlanC. Bovik
-
[55]
Zicheng Zhang, Wei Wu, Wei Sun, Dangyang Tu, Wei Lu, Xiongkuo Min, Ying Chen, and Guangtao Zhai. 2023. MD-VQA: Multi-Dimensional Quality Assess- ment for UGC Live Videos. (Mar 2023)
2023
-
[56]
Kai Zhao, Kun Yuan, Ming Sun, and Xing Wen. 2023. Zoom-vqa: Patches, frames and clips integration for video quality assessment. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 1302–1310
2023
-
[57]
Liangli Zhen, Peng Hu, Xu Wang, and Dezhong Peng. 2019. Deep supervised cross-modal retrieval. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 10394–10403
2019
-
[58]
Zhenqiang Ying, Maniratnam Mandal, Deepti Ghadiyaram, and Alan Bovik. 2021. Patch-vq:’patching up’the video quality problem. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 14019–14029
2021
-
[2011]
IEEE transactions on broadcasting 57, 2 (2011), 165–182
Objective video quality assessment methods: A classification, review, and performance comparison. IEEE transactions on broadcasting 57, 2 (2011), 165–182
2011
-
[2019]
In Proceedings of the 2019 CHI conference on human factors in computing systems
How to design voice based navigation for how-to videos. In Proceedings of the 2019 CHI conference on human factors in computing systems . 1–11
2019
-
[2020]
Patching Up
Patch-VQ: “Patching Up” the Video Quality Problem. Cornell University - arXiv,Cornell University - arXiv (Nov 2020)
2020
-
[2021]
IEEE Transactions on Image Processing (Jan 2021), 4449–4464
UGC-VQA: Benchmarking Blind Video Quality Assessment for User Gen- erated Content. IEEE Transactions on Image Processing (Jan 2021), 4449–4464. https://doi.org/10.1109/tip.2021.3072221
2021
-
[2023]
InProceedings of the 2023 CHI Conference on Human Factors in Computing Systems
Identifying Multimodal Context Awareness Requirements for Supporting User Interaction with Procedural Videos. InProceedings of the 2023 CHI Conference on Human Factors in Computing Systems . 1–17
2023
Reviewed August 8, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.