REVIEW 4 major objections 6 minor 59 references
Video Quality Assessment for Online Processing: From Spatial to Temporal Sampling
T0 review · 4 major / 6 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read Heavily squeezed video—a few patches from a few frames—still predicts video quality on common in-the-wild databases.
desk verdict A useful empirical grid on joint spatial/temporal sampling for VQA, but the 99.83% cost-reduction headline is an input-pixel ratio, not measured compute. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The mechanism is a two-stage spatio-temporal sampling pipeline. Temporally, the video is divided into segments and keyframes are drawn by one of three strategies (TSN, TSM, or ECO); spatially, each keyframe is divided into uniform grids and small patches (e.g., $32\times32$ or $64\times64$) are randomly sampled from every grid and stitched into a fragment, following the GMS method. This produces stacked spatio-temporal blocks that retain enough local distortion evidence for quality regression while discarding most pixels. The MGQA model then pairs MobileNet as a lightweight spatial distortion capturer with a graph-based temporal fusion module and fully connected global regression.
What would settle it
Run VSFA and MGQA on the same hardware with the exact input sizes of Table IX and measure end-to-end latency and FLOPs; if the runtime reduction is far smaller than the $99.83\%$ data reduction, the paper's computational-cost claim fails as stated. Alternatively, evaluate the trained MGQA on a high-motion video database with rapidly moving objects or frequent scene changes; if PLCC/SRCC drop below acceptable levels, the claim that heavily squeezed video predicts original quality does not generalize to dynamic content.
Extended reading notes
Core claim
The central claim is that the heavily squeezed video can be used to predict the quality of the original video: after aggressive temporal sampling (TSN, TSM, or ECO) and spatial grid-patch sampling (GMS), a handful of small fragments—for example $10 \times 160 \times 160$ pixel patches on KoNViD-1k instead of $208 \times 540 \times 960$—suffices to keep correlation with human scores close to the full-input model. The paper shows this for VSFA and a transformer variant across six databases, and then instantiates an online model, MGQA, whose MobileNet spatial extractor and graph-network temporal fusion achieve PLCC/SRCC values such as $0.87/0.89$ on CVD2014 and $0.83/0.86$ on LIVE-Qualcomm while processing about $0.17\%$ of the original pixel data on average. The authors read this as demonstrating the feasibility of online VQA through joint sampling and a deliberately simple architecture.
Load-bearing premise
The load-bearing premise is that reducing the amount of input data by $99.83\%$ reduces computational cost by roughly the same proportion, since no wall-clock time, FLOPs, or energy is measured and the architecture changes from ResNet50 plus GRU to MobileNet plus Graph.
Editorial extensions
If this is right
- Video quality monitoring can be shifted to tiny inputs: on the six tested databases, roughly $0.17\%$ of the pixel data suffices to keep correlation within a small margin of the full-input VSFA baseline.
- Aggressive joint sampling is safest where spatial distortion dominates (CVD2014, KoNViD-1k, LSVQ); on temporal-distortion-heavy content (LIVE-VQC, LIVE-Qualcomm) the paper reports larger drops, so sampling density should be content-adaptive.
- The order of sampled frames barely matters, so online systems can process short, shuffled fragments rather than long ordered sequences.
- A graph-based temporal fusion regressor is the key to MGQA's accuracy: replacing it with GRU, LSTM, transformer, or plain FC degrades LIVE-VQC performance substantially in the paper's ablation.
- Lightweight spatial backbones are not interchangeable: among MobileNet, FasterNet, EfficientNet, ShuffleNet, and MobileOne, only MobileNet keeps PLCC above 0.65 on LIVE-VQC, which points to backbone choice as a major efficiency-accuracy lever.
Reading between the lines
- The paper's 'computational cost' is measured as processed data volume, not runtime; an equal-hardware latency and energy benchmark is the natural next experiment and would either support or qualify the $99.83\%$ claim.
- The orderless robustness suggests that much of VQA's difficulty is spatial; an implied testable extension is to combine aggressive sampling with a motion-aware light module for high-motion and HDR content, which the paper explicitly flags as an open problem.
- The large backbone gap implies that data reduction alone does not determine efficiency; a future study could treat sampling density and backbone width as a joint trade-off curve rather than fixing one.
- The graph regressor's advantage over recurrent and transformer fusion on tiny inputs hints that relational pooling over patch positions is well matched to fragment-based input, a hypothesis the paper does not directly test.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper investigates how much spatial and temporal information can be removed from videos before a no-reference VQA model degrades unacceptably. It first measures VSFA and a Transformer-based variant under combinations of four spatial sampling schemes (grid-based patch extraction) and three temporal sampling strategies (TSN, TSM, ECO) on six public databases. It then proposes MGQA, a lightweight model using MobileNet and a graph network, and reports that its processed data is about 99.8% smaller than VSFA's, which is claimed as a computational cost reduction. The paper concludes that heavily squeezed video can predict original video quality and that an online VQA model is feasible.
Significance. If the efficiency claim is validated with actual compute measurements, the MGQA design would be a practical contribution to low-resource VQA. The systematic sampling study itself is a useful empirical reference for understanding the trade-off between input size and quality prediction accuracy, and the paper's explicit limitation statement about high-motion videos is honest. However, the headline 99.83% figure currently rests on an unvalidated proxy, and the absence of comparisons with other efficient VQA models limits the assessment of the method's relative value.
major comments (4)
- [Section IV-C, Table IX, and Introduction contribution 3] The '99.83% computational cost reduction' claim is computed as the ratio of input tensor element counts (VSFA input dimensions versus MGQA input dimensions), not as any measured computational cost. The paper reports no wall-clock time, FLOPs, or energy measurements anywhere, even though MGQA replaces ResNet50+GRU with MobileNet plus a graph module and adds the GMS patch-extraction pipeline. As written, the headline efficiency claim is an assertion about processed data volume, not about computational cost, and the paper provides no evidence that the two are interchangeable. Please provide direct efficiency measurements (latency, FLOPs, or energy) under the same hardware and protocol, or revise the claim to state precisely that what is reduced is input pixel count.
- [Section IV-A3 and Table IX] The MGQA default temporal sampler is stated to be ECO, which the text describes as extracting M=20 segments, with a possible extra keyframe when N mod M is nonzero (i.e., 20 or 21 keyframes). However, Table IX lists MGQA input frame counts of 10, 16, 22, 15, 12, and 14 for the six databases. These numbers are inconsistent with the described ECO default, making the reported processed-data ratios non-reproducible as written. Please clarify the actual frame counts and explain how they were derived.
- [Section IV-C (page 7)] The sentence 'The model consistently performs well, demonstrating state-of-the-art results on the CVD2014 and LIVE-Qualcomm datasets' is not supported by any comparison with existing BVQA methods under the same training/test protocol; Tables V-VIII only compare MGQA variants against each other. The 'state-of-the-art' claim should be substantiated with a table of recent lightweight or efficient VQA methods evaluated on the same splits, or removed.
- [Section IV-B, Tables II and III] The claim that joint spatial and temporal sampling 'does not lead to a significant drop in performance across the datasets' is not uniformly supported by the tables. For example, on LIVE-Qualcomm, VSFA with TSM and S1 drops from PLCC 0.73 to 0.57; on LIVE-VQC, the TSM/S1 setting drops from PLCC 0.72 to 0.63. The paper's own text concedes 'a significant performance drop on the temporal distortion dominated databases.' The conclusion should be scoped accordingly, e.g., to spatial-distortion-dominated databases, and the abstract and conclusion should be amended to match.
minor comments (6)
- [Section IV-A3] The phrase 'we set the sampling stepstep to' contains a typo; it should be 'step'.
- [Section II-B] 'sptaiotemporal sampling' appears to be a typo for 'spatiotemporal sampling'.
- [Table VII] The row labeled 'Transofrmer' should be 'Transformer'.
- [Section IV-C] The text 'the Grap and fully connected layers' has an incomplete word; it should likely be 'the Graph and fully connected layers'.
- [Figure 3 caption] The caption states that 'TSM extracts four frames ... while TSM only samples one frame from each segment,' which confuses TSN and TSM; the second occurrence should presumably refer to TSN.
- [Section IV-D] 'Followed by our previous study [13]' should be 'Following our previous study [13]'.
Circularity Check
No significant circularity: the sampling study and MGQA evaluation are empirical and self-contained.
full rationale
This paper is an empirical study of joint spatial and temporal sampling for VQA. The central claims—(i) aggressively squeezed video inputs yield acceptable quality prediction on six public databases (Tables II–V), and (ii) the MGQA model built on MobileNet plus a graph regressor achieves competitive accuracy—are established by training on human-annotated quality scores and testing on held-out splits. No prediction is defined in terms of a fitted constant, and no equation equates the output with an input by construction. The redundancy premise is motivated by prior self-citations ([12], [13], [38]) but is re-tested in the present settings: Tables II–V and Table X directly measure the effect of sampling, so the self-citations are not load-bearing. The 99.83% 'computational cost reduction' (Table IX) is explicitly a ratio of processed data sizes (e.g., VSFA 208*540*960 vs MGQA 10*160*160), not a measured latency or FLOP comparison; this is a validity concern about an efficiency claim, not a circular derivation. The paper also states an honest limitation: the proposed approach may not generalize to high-motion or highly dynamic videos (Section IV-B). No uniqueness theorem from the authors' prior work is invoked to force a choice, and the adopted baseline (VSFA) and sampling method (GMS) are cited from external groups. Accordingly, there is no circular step to report.
Assumptions & free parameters
free parameters (4)
- Spatial sampling configuration =
S1: 4 patches of 32x32; S2: 4 of 64x64; S3: 25 of 32x32; S4: 36 of 32x32
- Temporal sampling configuration =
TSN: 10 segments, 4 frames per segment; TSM: 10 segments; ECO: 20 segments
- MGQA default sampling =
S3 spatial + ECO temporal
- Logistic mapping parameters =
tau1-tau4 fitted per database
assumptions (4)
- domain assumption Natural videos contain large redundant information in spatial and temporal dimensions
- domain assumption The six public databases are representative of video content in online processing scenarios
- domain assumption Random 60/20/20 splits with three repetitions yield reliable performance estimates
- ad hoc to paper Input data ratio is a valid proxy for computational cost
Cite this review
Pith. "Pith review of Video Quality Assessment for Online Processing: From Spatial to Temporal Sampling." pith.science (2026). https://pith.science/paper/F4NKO7HU
@misc{pith2026250107087,
author = {Pith},
title = {Pith review of: Video Quality Assessment for Online Processing: From Spatial to Temporal Sampling},
year = {2026},
howpublished = {\url{https://pith.science/paper/F4NKO7HU}},
note = {Machine review of arXiv:2501.07087}
}
read the original abstract
With the rapid development of multimedia processing and deep learning technologies, especially in the field of video understanding, video quality assessment (VQA) has achieved significant progress. Although researchers have moved from designing efficient video quality mapping models to various research directions, in-depth exploration of the effectiveness-efficiency trade-offs of spatio-temporal modeling in VQA models is still less sufficient. Considering the fact that videos have highly redundant information, this paper investigates this problem from the perspective of joint spatial and temporal sampling, aiming to seek the answer to how little information we should keep at least when feeding videos into the VQA models while with acceptable performance sacrifice. To this end, we drastically sample the video's information from both spatial and temporal dimensions, and the heavily squeezed video is then fed into a stable VQA model. Comprehensive experiments regarding joint spatial and temporal sampling are conducted on six public video quality databases, and the results demonstrate the acceptable performance of the VQA model when throwing away most of the video information. Furthermore, with the proposed joint spatial and temporal sampling strategy, we make an initial attempt to design an online VQA model, which is instantiated by as simple as possible a spatial feature extractor, a temporal feature fusion module, and a global quality regression module. Through quantitative and qualitative experiments, we verify the feasibility of online VQA model by simplifying itself and reducing input.
Figures
Figures from the paper (2 more)
Reference graph
Works this paper leans on
-
[1]
A survey on recent advances in video quality assessment,
J. Yan, Y . Fang, X. Liu, Y . Yao, and X. Sui, “A survey on recent advances in video quality assessment,” Chinese Journal of Computers , vol. 46, no. 10, pp. 2196–2224, 2023
work page 2023
-
[2]
The review of distortion-related image quality assessment,
J. Yan, Y . Fang, and X. Liu, “The review of distortion-related image quality assessment,” Journal of Image and Graphics , vol. 27, no. 05, pp. 1430–1466, 2022
2022
-
[3]
Superpixel-based quality assessment of multi-exposure image fusion for both static and dynamic scenes,
Y . Fang, Y . Zeng, W. Jiang, H. Zhu, and J. Yan, “Superpixel-based quality assessment of multi-exposure image fusion for both static and dynamic scenes,” IEEE Transactions on Image Processing , vol. 30, pp. 2526–2537, 2021
2021
-
[4]
Perceptually optimized deep high- dynamic-range image tone mapping,
C. Le, J. Yan, Y . Fang, and K. Ma, “Perceptually optimized deep high- dynamic-range image tone mapping,” in IEEE International Conference on Virtual Reality and Visualization , 2021
work page 2021
-
[5]
Applications of objective image quality assessment methods,
Z. Wang, “Applications of objective image quality assessment methods,” IEEE Signal Processing Magazine , vol. 28, no. 6, pp. 137–142, 2011
work page 2011
-
[6]
Perceptual video quality assessment: A survey,
X. Min, H. Duan, W. Sun, Y . Zhu, and G. Zhai, “Perceptual video quality assessment: A survey,” ArXiv Preprint ArXiv:2402.03413 , 2024
arXiv 2024
-
[7]
A human visual system-based objective video distortion measurement system,
Z. Wang and A. C. Bovik, “A human visual system-based objective video distortion measurement system,” in International Conference on Multimedia Processing and System , 2001, pp. 1–4
work page 2001
-
[8]
Image quality assessment: From error visibility to structural similarity,
Z. Wang, A. C. Bovik, H. R. Sheikh, and E. P. Simoncelli, “Image quality assessment: From error visibility to structural similarity,” IEEE Transactions on Image Processing , vol. 13, no. 4, pp. 600–612, 2004
2004
Show all 59 references
-
[9]
Video quality assessment using a statistical model of human visual speed perception,
Z. Wang and Q. Li, “Video quality assessment using a statistical model of human visual speed perception,” Journal of the Optical Society of America. A, vol. 24, no. 12, 2007
2007
-
[10]
Viewer response to time-varying video quality,
D. E. Pearson, “Viewer response to time-varying video quality,” in Human Vision and Electronic Imaging III , vol. 3299, 1998, pp. 16–25
1998
-
[11]
Efficient feature extraction, encoding and classification for action recognition,
V . Kantorov and I. Laptev, “Efficient feature extraction, encoding and classification for action recognition,” in IEEE Conference on Computer Vision and Pattern Recognition , 2014, pp. 2593–2600
2014
-
[12]
Subjective and objective quality of experience of free viewpoint videos,
J. Yan, J. Li, Y . Fang, Z. Che, X. Xia, and Y . Liu, “Subjective and objective quality of experience of free viewpoint videos,” IEEE Transactions on Image Processing , vol. 31, pp. 3896–3907, 2022
2022
-
[13]
Study of spatio-temporal modeling in video quality assessment,
Y . Fang, Z. Li, J. Yan, X. Sui, and H. Liu, “Study of spatio-temporal modeling in video quality assessment,” IEEE Transactions on Image Processing, vol. 32, pp. 2693–2702, 2023
2023
-
[14]
Quality assessment of in-the-wild videos,
D. Li, T. Jiang, and M. Jiang, “Quality assessment of in-the-wild videos,” in ACM International Conference on Multimedia , 2019, pp. 2351–2359
2019
-
[15]
FAST-VQA: Efficient end-to-end video quality assessment with fragment sampling,
H. Wu, C. Chen, J. Hou, L. Liao, A. Wang, W. Sun, Q. Yan, and W. Lin, “FAST-VQA: Efficient end-to-end video quality assessment with fragment sampling,” in European Conference on Computer Vision, 2022, pp. 538–554
2022
-
[16]
Analysis of video quality datasets via design of minimalistic video quality models,
W. Sun, W. Wen, X. Min, L. Lan, G. Zhai, and K. Ma, “Analysis of video quality datasets via design of minimalistic video quality models,” IEEE Transactions on Pattern Analysis and Machine Intelligence , 2024
2024
-
[17]
Blind prediction of natural video quality,
M. Saad, A. Bovik, and C. Charrier, “Blind prediction of natural video quality,” IEEE Transactions on Image Processing , vol. 23, no. 3, pp. 1352–1365, 2014
2014
-
[18]
Learning spatiotemporal inter- actions for user-generated video quality assessment,
H. Zhu, B. Chen, L. Zhu, and S. Wang, “Learning spatiotemporal inter- actions for user-generated video quality assessment,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 33, no. 3, pp. 1031– 1042, 2022
2022
-
[19]
A completely blind video integrity oracle,
A. Mittal, M. A. Saad, and A. C. Bovik, “A completely blind video integrity oracle,” IEEE Transactions on Image Processing, vol. 25, no. 1, pp. 289–300, 2016
2016
-
[20]
Two-level approach for no-reference consumer video quality assessment,
J. Korhonen, “Two-level approach for no-reference consumer video quality assessment,” IEEE Transactions on Image Processing , vol. 28, no. 12, pp. 5923–5938, 2019
2019
-
[21]
ChipQA: No-reference video quality prediction via space-time chips,
J. Ebenezer, Z. Shang, Y . Wu, H. Wei, S. Sethuraman, and A. C. Bovik, “ChipQA: No-reference video quality prediction via space-time chips,” IEEE Transactions on Image Processing, vol. 30, pp. 8059–8074, 2021
2021
-
[22]
Spatiotemporal representation learning for blind video quality assessment,
Y . Liu, J. Wu, L. Li, W. Dong, J. Zhang, and G. Shi, “Spatiotemporal representation learning for blind video quality assessment,” IEEE Trans- actions on Circuits and Systems for Video Technology , vol. 32, no. 6, pp. 3500–3513, 2021
2021
-
[23]
Blindly assess quality of in-the-wild videos via quality-aware pre-training and motion perception,
B. Li, W. Zhang, M. Tian, G. Zhai, and X. Wang, “Blindly assess quality of in-the-wild videos via quality-aware pre-training and motion perception,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 32, no. 9, pp. 5944–5958, 2022
2022
-
[24]
Deep residual learning for image recognition,
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in IEEE Conference on Computer Vision and Pattern Recognition, 2016, pp. 770–778
2016
-
[25]
Learning phrase representations using RNN encoder-decoder for statistical machine translation,
K. Cho, B. Van Merri ¨enboer, C. Gulcehre, D. Bahdanau, F. Bougares, H. Schwenk, and Y . Bengio, “Learning phrase representations using RNN encoder-decoder for statistical machine translation,” in Conference on Empirical Methods in Natural Language Processing , 2014
2014
-
[26]
Exploring video quality assessment on user generated contents from aesthetic and technical perspectives,
H. Wu, E. Zhang, L. Liao, C. Chen, J. Hou, A. Wang, W. Sun, Q. Yan, and W. Lin, “Exploring video quality assessment on user generated contents from aesthetic and technical perspectives,” in International Conference on Computer Vision , 2023, pp. 20 144–20 154
2023
-
[27]
RIRNet: Recurrent-in- recurrent network for video quality assessment,
P. Chen, L. Li, L. Ma, J. Wu, and G. Shi, “RIRNet: Recurrent-in- recurrent network for video quality assessment,” in ACM International Conference on Multimedia , 2020, pp. 834–842
2020
-
[28]
End-to-end blind quality assessment of compressed videos using deep neural networks
W. Liu, Z. Duanmu, and Z. Wang, “End-to-end blind quality assessment of compressed videos using deep neural networks.” inACM International Conference on Multimedia , 2018, pp. 546–554
2018
-
[29]
DisCoVQA: Temporal distortion-content transformers for video quality assessment,
H. Wu, C. Chen, L. Liao, J. Hou, W. Sun, Q. Yan, and W. Lin, “DisCoVQA: Temporal distortion-content transformers for video quality assessment,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 33, pp. 4840–4854, 2023
2023
-
[30]
TSM: Temporal shift module for efficient video understanding,
J. Lin, C. Gan, and S. Han, “TSM: Temporal shift module for efficient video understanding,” in International Conference on Computer Vision , 2019, pp. 7083–7093
2019
-
[31]
Temporal segment networks for action recognition in videos,
L. Wang, Y . Xiong, Y . Wang, Z.and Qiao, D. Lin, X. Tang, and L. Van Gool, “Temporal segment networks for action recognition in videos,” IEEE Transactions on Pattern Analysis and Machine Intelli- gence, vol. 41, no. 11, pp. 2740–2755, 2018
2018
-
[32]
Temporal relational reasoning in videos,
B. Zhou, A. Andonian, A. Oliva, and A. Torralba, “Temporal relational reasoning in videos,” inEuropean Conference on Computer Vision, 2018, pp. 831–846
2018
-
[33]
ECO: Efficient convolutional network for online video understanding,
M. Zolfaghari, K. Singh, and T. Brox, “ECO: Efficient convolutional network for online video understanding,” in European Conference on Computer Vision, 2018, pp. 695–712
2018
-
[34]
Video swin transformer,
Z. Liu, J. Ning, Y . Cao, Y . Wei, Z. Zhang, S. Lin, and H. Hu, “Video swin transformer,” in IEEE Conference on Computer Vision and Pattern Recognition, 2022, pp. 3192–3201
2022
-
[35]
Neighbourhood representative sampling for efficient end-to- end video quality assessment,
H. Wu, C. Chen, L. Liao, J. Hou, W. Sun, J. Yan, Qi.and Gu, and W. Lin, “Neighbourhood representative sampling for efficient end-to- end video quality assessment,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 45, pp. 15 185–15 202, 2023
2023
-
[36]
Zoom-VQA: Patches, frames and clips integration for video quality assessment,
K. Zhao, K. Yuan, M. Sun, and X. Wen, “Zoom-VQA: Patches, frames and clips integration for video quality assessment,” in IEEE Conference on Computer Vision and Pattern Recognition , 2023, pp. 1302–1310
2023
-
[37]
Scaling and masking: A new paradigm of data sampling for image and video quality assessment,
Y . Liu, Y . Quan, G. Xiao, A. Li, and J. Wu, “Scaling and masking: A new paradigm of data sampling for image and video quality assessment,” in AAAI Conference on Artificial Intelligence , 2024
2024
-
[38]
Revisiting the robustness of spatio-temporal modeling in video quality assessment,
J. Yan, L. Wu, W. Jiang, C. Liu, and F. Shen, “Revisiting the robustness of spatio-temporal modeling in video quality assessment,” Displays, vol. 81, p. 102585, 2024
2024
-
[39]
The Konstanz natural video database (KoNViD-1k),
V . Hosu, F. Hahn, M. Jenadeleh, H. Lin, H. Men, T. Szir ´anyi, S. Li, and D. Saupe, “The Konstanz natural video database (KoNViD-1k),” in Ninth International Conference on Quality of Multimedia Experience , 2017, pp. 1–6
2017
-
[40]
Large-scale study of perceptual video quality,
Z. Sinno and A. C. Bovik, “Large-scale study of perceptual video quality,” IEEE Transactions on Image Processing , vol. 28, no. 2, pp. 612–627, 2018
2018
-
[41]
CVD2014—A database for evaluating no-reference 11 video quality assessment algorithms,
M. Nuutinen, M. Virtanen, T.and Vaahteranoksa, T. Vuori, P. Oittinen, and J. H ¨akkinen, “CVD2014—A database for evaluating no-reference 11 video quality assessment algorithms,” IEEE Transactions on Image Processing, vol. 25, no. 7, pp. 3073–3086, 2016
2016
-
[42]
In-capture mobile video distortions: A study of subjective behavior and objective algorithms,
D. Ghadiyaram, J. Pan, A. Bovik, A. Moorthy, P. Panda, and K. Yang, “In-capture mobile video distortions: A study of subjective behavior and objective algorithms,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 28, no. 9, pp. 2061–2077, 2017
2017
-
[43]
Patch-VQ: ‘Patch- ing up’ the video quality problem,
Z. Ying, M. Mandal, D. Ghadiyaram, and A. Bovik, “Patch-VQ: ‘Patch- ing up’ the video quality problem,” in IEEE Conference on Computer Vision and Pattern Recognition , 2021, pp. 14 019–14 029
2021
-
[44]
NTIRE 2023 quality assessment of video enhancement challenge,
X. Liu, R. Timofte, Y . Dong, Z. Ma, H. Fan, C. Zhu, X. Min, G. Zhai, Z. Jia, M. Agarla et al. , “NTIRE 2023 quality assessment of video enhancement challenge,” in IEEE Conference on Computer Vision and Pattern Recognition, 2023, pp. 1551–1569
2023
-
[45]
Attention is all you need,
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in Advances in Neural Information Processing Systems , vol. 30, 2017
2017
-
[46]
ImageNet large scale visual recognition challenge,
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein et al., “ImageNet large scale visual recognition challenge,” International Journal of Computer Vision, vol. 115, pp. 211–252, 2015
2015
-
[47]
Final report from the Video Quality Experts Group on the validation of objective models of video quality assessment,
VQEG, “Final report from the Video Quality Experts Group on the validation of objective models of video quality assessment,” [online]. Available: http://www.vqeg.org, 2003
2003
-
[48]
FASTER recurrent networks for efficient video classification,
L. Zhu, D. Tran, L. Sevilla-Lara, Y . Yang, M. Feiszli, and H. Wang, “FASTER recurrent networks for efficient video classification,” in AAAI Conference on Artificial Intelligence , vol. 34, no. 07, 2020, pp. 13 098– 13 105
2020
-
[49]
Qoe evaluation for live broadcasting video,
P. Chen, L. Li, Y . Huang, F. Tan, and W. Chen, “Qoe evaluation for live broadcasting video,” in IEEE International Conference on Image Processing, 2019, pp. 454–458
2019
-
[50]
Study of the subjective and objective quality of high motion live streaming videos,
Z. Shang, J. P. Ebenezer, Y . Wu, H. Wei, S. Sethuraman, and A. C. Bovik, “Study of the subjective and objective quality of high motion live streaming videos,” IEEE Transactions on Image Processing, vol. 31, pp. 1027–1041, 2021
2021
-
[51]
A study of subjective and objec- tive quality assessment of hdr videos,
Z. Shang, J. P. Ebenezer, A. K. Venkataramanan, Y . Wu, H. Wei, S. Sethuraman, and A. C. Bovik, “A study of subjective and objec- tive quality assessment of hdr videos,” IEEE Transactions on Image Processing, vol. 33, pp. 42–57, 2023
2023
-
[52]
Hidro-vqa: High dynamic range oracle for video quality assessment,
S. Saini, A. Saha, and A. C. Bovik, “Hidro-vqa: High dynamic range oracle for video quality assessment,” in IEEE/CVF Winter Conference on Applications of Computer Vision , 2024, pp. 469–479
2024
-
[53]
MobileNets: Efficient convolutional neural networks for mobile vision applications,
A. Howard, M. Zhu, B. Chen, D. Kalenichenko, W. Wang, T. Weyand, M. Andreetto, and H. Adam, “MobileNets: Efficient convolutional neural networks for mobile vision applications,” ArXiv Preprint ArXiv:1704.04861, 2017
2017 arXiv
-
[54]
GraphIQA: Learning distortion graph representations for blind image quality assessment,
S. Sun, T. Yu, J. Xu, W. Zhou, and Z. Chen, “GraphIQA: Learning distortion graph representations for blind image quality assessment,” IEEE Transactions on Multimedia , vol. 25, pp. 2912–2925, 2022
2022
-
[55]
Run, don’t walk: Chasing higher flops for faster neural networks,
J. Chen, S. Kao, H. He, W. Zhuo, S. Wen, C. Lee, and S. Chan, “Run, don’t walk: Chasing higher flops for faster neural networks,” in IEEE Conference on Computer Vision and Pattern Recognition , 2023, pp. 12 021–12 031
2023
-
[56]
EfficientNet: Rethinking model scaling for con- volutional neural networks,
M. Tan and Q. Le, “EfficientNet: Rethinking model scaling for con- volutional neural networks,” in International Conference on Machine Learning, 2019, pp. 6105–6114
2019
-
[57]
ShuffleNet: An extremely efficient convolutional neural network for mobile devices,
X. Zhang, X. Zhou, M. Lin, and J. Sun, “ShuffleNet: An extremely efficient convolutional neural network for mobile devices,” in IEEE Conference on Computer Vision and Pattern Recognition , 2018, pp. 6848–6856
2018
-
[58]
MobileOne: An improved one millisecond mobile backbone,
P. Vasu, J. Gabriel, J. Zhu, O. Tuzel, and A. Ranjan, “MobileOne: An improved one millisecond mobile backbone,” in IEEE Conference on Computer Vision and Pattern Recognition , 2023, pp. 7907–7917
2023
-
[59]
Long short-term memory,
A. Graves and A. Graves, “Long short-term memory,” Supervised Sequence Labelling with Recurrent Neural Networks , pp. 37–45, 2012. Jiebin Yan received the Ph.D. degree from Jiangxi University of Finance and Economics, Nanchang, China. He was a computer vision engineer with MT-...
2012
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.