Pith. sign in

REVIEW 4 major objections 6 minor 59 references

Video Quality Assessment for Online Processing: From Spatial to Temporal Sampling

T0 review · 4 major / 6 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read Heavily squeezed video—a few patches from a few frames—still predicts video quality on common in-the-wild databases.

desk verdict A useful empirical grid on joint spatial/temporal sampling for VQA, but the 99.83% cost-reduction headline is an input-pixel ratio, not measured compute. read the letter →

arxiv 2501.07087 v1 pith:F4NKO7HU submitted 2025-01-13 cs.CV cs.AI

classification cs.CVcs.AI
keywords videoqualityassessmentno-referenceVQAspatial-temporalsamplingtemporalkeyframegridmini-patchlightweightmodelMGQAgraph-basedregression
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper asks how little of a video a quality-assessment model actually needs to see. It argues that videos are so redundant in space and time that a heavily squeezed video—a few small patches from a few sampled frames—can predict the perceptual quality of the original video with acceptable accuracy on existing in-the-wild databases. Across six databases, jointly sampling temporal keyframes and spatial grid patches keeps performance close to feeding the full video, and the resulting MGQA model cuts the amount of processed pixel data by an average of $99.83\%$ relative to the VSFA baseline. If true, online and mobile video-quality monitoring could run on roughly $1/500$ of the pixel data without needing full-resolution video as input.

What carries the argument

The mechanism is a two-stage spatio-temporal sampling pipeline. Temporally, the video is divided into segments and keyframes are drawn by one of three strategies (TSN, TSM, or ECO); spatially, each keyframe is divided into uniform grids and small patches (e.g., $32\times32$ or $64\times64$) are randomly sampled from every grid and stitched into a fragment, following the GMS method. This produces stacked spatio-temporal blocks that retain enough local distortion evidence for quality regression while discarding most pixels. The MGQA model then pairs MobileNet as a lightweight spatial distortion capturer with a graph-based temporal fusion module and fully connected global regression.

What would settle it

Run VSFA and MGQA on the same hardware with the exact input sizes of Table IX and measure end-to-end latency and FLOPs; if the runtime reduction is far smaller than the $99.83\%$ data reduction, the paper's computational-cost claim fails as stated. Alternatively, evaluate the trained MGQA on a high-motion video database with rapidly moving objects or frequent scene changes; if PLCC/SRCC drop below acceptable levels, the claim that heavily squeezed video predicts original quality does not generalize to dynamic content.

Watch

Extended reading notes

Core claim

The central claim is that the heavily squeezed video can be used to predict the quality of the original video: after aggressive temporal sampling (TSN, TSM, or ECO) and spatial grid-patch sampling (GMS), a handful of small fragments—for example $10 \times 160 \times 160$ pixel patches on KoNViD-1k instead of $208 \times 540 \times 960$—suffices to keep correlation with human scores close to the full-input model. The paper shows this for VSFA and a transformer variant across six databases, and then instantiates an online model, MGQA, whose MobileNet spatial extractor and graph-network temporal fusion achieve PLCC/SRCC values such as $0.87/0.89$ on CVD2014 and $0.83/0.86$ on LIVE-Qualcomm while processing about $0.17\%$ of the original pixel data on average. The authors read this as demonstrating the feasibility of online VQA through joint sampling and a deliberately simple architecture.

Load-bearing premise

The load-bearing premise is that reducing the amount of input data by $99.83\%$ reduces computational cost by roughly the same proportion, since no wall-clock time, FLOPs, or energy is measured and the architecture changes from ResNet50 plus GRU to MobileNet plus Graph.

Editorial extensions

If this is right

  • Video quality monitoring can be shifted to tiny inputs: on the six tested databases, roughly $0.17\%$ of the pixel data suffices to keep correlation within a small margin of the full-input VSFA baseline.
  • Aggressive joint sampling is safest where spatial distortion dominates (CVD2014, KoNViD-1k, LSVQ); on temporal-distortion-heavy content (LIVE-VQC, LIVE-Qualcomm) the paper reports larger drops, so sampling density should be content-adaptive.
  • The order of sampled frames barely matters, so online systems can process short, shuffled fragments rather than long ordered sequences.
  • A graph-based temporal fusion regressor is the key to MGQA's accuracy: replacing it with GRU, LSTM, transformer, or plain FC degrades LIVE-VQC performance substantially in the paper's ablation.
  • Lightweight spatial backbones are not interchangeable: among MobileNet, FasterNet, EfficientNet, ShuffleNet, and MobileOne, only MobileNet keeps PLCC above 0.65 on LIVE-VQC, which points to backbone choice as a major efficiency-accuracy lever.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper's 'computational cost' is measured as processed data volume, not runtime; an equal-hardware latency and energy benchmark is the natural next experiment and would either support or qualify the $99.83\%$ claim.
  • The orderless robustness suggests that much of VQA's difficulty is spatial; an implied testable extension is to combine aggressive sampling with a motion-aware light module for high-motion and HDR content, which the paper explicitly flags as an open problem.
  • The large backbone gap implies that data reduction alone does not determine efficiency; a future study could treat sampling density and backbone width as a joint trade-off curve rather than fixing one.
  • The graph regressor's advantage over recurrent and transformer fusion on tiny inputs hints that relational pooling over patch positions is well matched to fragment-based input, a hypothesis the paper does not directly test.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper investigates how much spatial and temporal information can be removed from videos before a no-reference VQA model degrades unacceptably. It first measures VSFA and a Transformer-based variant under combinations of four spatial sampling schemes (grid-based patch extraction) and three temporal sampling strategies (TSN, TSM, ECO) on six public databases. It then proposes MGQA, a lightweight model using MobileNet and a graph network, and reports that its processed data is about 99.8% smaller than VSFA's, which is claimed as a computational cost reduction. The paper concludes that heavily squeezed video can predict original video quality and that an online VQA model is feasible.

Significance. If the efficiency claim is validated with actual compute measurements, the MGQA design would be a practical contribution to low-resource VQA. The systematic sampling study itself is a useful empirical reference for understanding the trade-off between input size and quality prediction accuracy, and the paper's explicit limitation statement about high-motion videos is honest. However, the headline 99.83% figure currently rests on an unvalidated proxy, and the absence of comparisons with other efficient VQA models limits the assessment of the method's relative value.

major comments (4)
  1. [Section IV-C, Table IX, and Introduction contribution 3] The '99.83% computational cost reduction' claim is computed as the ratio of input tensor element counts (VSFA input dimensions versus MGQA input dimensions), not as any measured computational cost. The paper reports no wall-clock time, FLOPs, or energy measurements anywhere, even though MGQA replaces ResNet50+GRU with MobileNet plus a graph module and adds the GMS patch-extraction pipeline. As written, the headline efficiency claim is an assertion about processed data volume, not about computational cost, and the paper provides no evidence that the two are interchangeable. Please provide direct efficiency measurements (latency, FLOPs, or energy) under the same hardware and protocol, or revise the claim to state precisely that what is reduced is input pixel count.
  2. [Section IV-A3 and Table IX] The MGQA default temporal sampler is stated to be ECO, which the text describes as extracting M=20 segments, with a possible extra keyframe when N mod M is nonzero (i.e., 20 or 21 keyframes). However, Table IX lists MGQA input frame counts of 10, 16, 22, 15, 12, and 14 for the six databases. These numbers are inconsistent with the described ECO default, making the reported processed-data ratios non-reproducible as written. Please clarify the actual frame counts and explain how they were derived.
  3. [Section IV-C (page 7)] The sentence 'The model consistently performs well, demonstrating state-of-the-art results on the CVD2014 and LIVE-Qualcomm datasets' is not supported by any comparison with existing BVQA methods under the same training/test protocol; Tables V-VIII only compare MGQA variants against each other. The 'state-of-the-art' claim should be substantiated with a table of recent lightweight or efficient VQA methods evaluated on the same splits, or removed.
  4. [Section IV-B, Tables II and III] The claim that joint spatial and temporal sampling 'does not lead to a significant drop in performance across the datasets' is not uniformly supported by the tables. For example, on LIVE-Qualcomm, VSFA with TSM and S1 drops from PLCC 0.73 to 0.57; on LIVE-VQC, the TSM/S1 setting drops from PLCC 0.72 to 0.63. The paper's own text concedes 'a significant performance drop on the temporal distortion dominated databases.' The conclusion should be scoped accordingly, e.g., to spatial-distortion-dominated databases, and the abstract and conclusion should be amended to match.
minor comments (6)
  1. [Section IV-A3] The phrase 'we set the sampling stepstep to' contains a typo; it should be 'step'.
  2. [Section II-B] 'sptaiotemporal sampling' appears to be a typo for 'spatiotemporal sampling'.
  3. [Table VII] The row labeled 'Transofrmer' should be 'Transformer'.
  4. [Section IV-C] The text 'the Grap and fully connected layers' has an incomplete word; it should likely be 'the Graph and fully connected layers'.
  5. [Figure 3 caption] The caption states that 'TSM extracts four frames ... while TSM only samples one frame from each segment,' which confuses TSN and TSM; the second occurrence should presumably refer to TSN.
  6. [Section IV-D] 'Followed by our previous study [13]' should be 'Following our previous study [13]'.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the sampling study and MGQA evaluation are empirical and self-contained.

full rationale

This paper is an empirical study of joint spatial and temporal sampling for VQA. The central claims—(i) aggressively squeezed video inputs yield acceptable quality prediction on six public databases (Tables II–V), and (ii) the MGQA model built on MobileNet plus a graph regressor achieves competitive accuracy—are established by training on human-annotated quality scores and testing on held-out splits. No prediction is defined in terms of a fitted constant, and no equation equates the output with an input by construction. The redundancy premise is motivated by prior self-citations ([12], [13], [38]) but is re-tested in the present settings: Tables II–V and Table X directly measure the effect of sampling, so the self-citations are not load-bearing. The 99.83% 'computational cost reduction' (Table IX) is explicitly a ratio of processed data sizes (e.g., VSFA 208*540*960 vs MGQA 10*160*160), not a measured latency or FLOP comparison; this is a validity concern about an efficiency claim, not a circular derivation. The paper also states an honest limitation: the proposed approach may not generalize to high-motion or highly dynamic videos (Section IV-B). No uniqueness theorem from the authors' prior work is invoked to force a choice, and the adopted baseline (VSFA) and sampling method (GMS) are cited from external groups. Accordingly, there is no circular step to report.

Assumptions & free parameters 4 free parameters · 4 assumptions · 0 invented entities

The central claim rests on the redundancy of video content, the representativeness of six public databases, and the unstated equivalence between processed data and computational cost. The sampling schemes and default model settings are hand-chosen rather than fitted, and the logistic mapping parameters are fitted per database as standard practice.

free parameters (4)
  • Spatial sampling configuration = S1: 4 patches of 32x32; S2: 4 of 64x64; S3: 25 of 32x32; S4: 36 of 32x32
    Hand-chosen combinations from GMS; S3 is later used as the MGQA default without a formal selection rule.
  • Temporal sampling configuration = TSN: 10 segments, 4 frames per segment; TSM: 10 segments; ECO: 20 segments
    Adopted from prior action recognition work; not re-derived for VQA.
  • MGQA default sampling = S3 spatial + ECO temporal
    Selected after the sensitivity study in Table VIII; the paper states it is the default but does not justify the choice beyond empirical stability.
  • Logistic mapping parameters = tau1-tau4 fitted per database
    Standard VQEG nonlinear mapping before computing PLCC; performance numbers therefore depend on per-dataset calibration.
assumptions (4)
  • domain assumption Natural videos contain large redundant information in spatial and temporal dimensions
    Invoked in the introduction and Section II-B; the entire sampling study relies on this premise.
  • domain assumption The six public databases are representative of video content in online processing scenarios
    Conclusions about acceptable performance are drawn from KoNViD-1k, LIVE-VQC, CVD2014, LIVE-Qualcomm, LSVQ, and NTIRE; the paper itself warns that high-motion and HDR videos may violate this.
  • domain assumption Random 60/20/20 splits with three repetitions yield reliable performance estimates
    No standard fixed splits are used and no standard deviations are reported.
  • ad hoc to paper Input data ratio is a valid proxy for computational cost
    Contribution 3 and Table IX compute the 99.83% reduction from input pixel counts rather than measured runtime, FLOPs, or energy.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Video Quality Assessment for Online Processing: From Spatial to Temporal Sampling." pith.science (2026). https://pith.science/paper/F4NKO7HU

@misc{pith2026250107087,
  author       = {Pith},
  title        = {Pith review of: Video Quality Assessment for Online Processing: From Spatial to Temporal Sampling},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/F4NKO7HU}},
  note         = {Machine review of arXiv:2501.07087}
}
read the original abstract

With the rapid development of multimedia processing and deep learning technologies, especially in the field of video understanding, video quality assessment (VQA) has achieved significant progress. Although researchers have moved from designing efficient video quality mapping models to various research directions, in-depth exploration of the effectiveness-efficiency trade-offs of spatio-temporal modeling in VQA models is still less sufficient. Considering the fact that videos have highly redundant information, this paper investigates this problem from the perspective of joint spatial and temporal sampling, aiming to seek the answer to how little information we should keep at least when feeding videos into the VQA models while with acceptable performance sacrifice. To this end, we drastically sample the video's information from both spatial and temporal dimensions, and the heavily squeezed video is then fed into a stable VQA model. Comprehensive experiments regarding joint spatial and temporal sampling are conducted on six public video quality databases, and the results demonstrate the acceptable performance of the VQA model when throwing away most of the video information. Furthermore, with the proposed joint spatial and temporal sampling strategy, we make an initial attempt to design an online VQA model, which is instantiated by as simple as possible a spatial feature extractor, a temporal feature fusion module, and a global quality regression module. Through quantitative and qualitative experiments, we verify the feasibility of online VQA model by simplifying itself and reducing input.

Figures

Figures reproduced from arXiv: 2501.07087 by the authors.

Figure 1
Figure 1. An illustration of spatio-temporal sampling paradigm, which extracts [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. The whole framework of this study (a). Given a video sequence, we first squeeze the input video by joint spatial and temporal sampling (b and c), [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. The intuitive comparison of different temporal sampling methods. [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: Comparison of the proposed MGQA with different spatial feature [PITH_FULL_IMAGE:figures/full_fig_p008_4.png]
Figure 5
Figure 5. Figure 5: The visual examples of spatio-temporal local quality maps, where blue areas refer to relatively low quality scores and red areas refer to high scores. [PITH_FULL_IMAGE:figures/full_fig_p009_5.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

59 extracted references · 51 canonical work pages

  1. [1]

    A survey on recent advances in video quality assessment,

    J. Yan, Y . Fang, X. Liu, Y . Yao, and X. Sui, “A survey on recent advances in video quality assessment,” Chinese Journal of Computers , vol. 46, no. 10, pp. 2196–2224, 2023

  2. [2]

    The review of distortion-related image quality assessment,

    J. Yan, Y . Fang, and X. Liu, “The review of distortion-related image quality assessment,” Journal of Image and Graphics , vol. 27, no. 05, pp. 1430–1466, 2022

  3. [3]

    Superpixel-based quality assessment of multi-exposure image fusion for both static and dynamic scenes,

    Y . Fang, Y . Zeng, W. Jiang, H. Zhu, and J. Yan, “Superpixel-based quality assessment of multi-exposure image fusion for both static and dynamic scenes,” IEEE Transactions on Image Processing , vol. 30, pp. 2526–2537, 2021

  4. [4]

    Perceptually optimized deep high- dynamic-range image tone mapping,

    C. Le, J. Yan, Y . Fang, and K. Ma, “Perceptually optimized deep high- dynamic-range image tone mapping,” in IEEE International Conference on Virtual Reality and Visualization , 2021

  5. [5]

    Applications of objective image quality assessment methods,

    Z. Wang, “Applications of objective image quality assessment methods,” IEEE Signal Processing Magazine , vol. 28, no. 6, pp. 137–142, 2011

  6. [6]

    Perceptual video quality assessment: A survey,

    X. Min, H. Duan, W. Sun, Y . Zhu, and G. Zhai, “Perceptual video quality assessment: A survey,” ArXiv Preprint ArXiv:2402.03413 , 2024

  7. [7]

    A human visual system-based objective video distortion measurement system,

    Z. Wang and A. C. Bovik, “A human visual system-based objective video distortion measurement system,” in International Conference on Multimedia Processing and System , 2001, pp. 1–4

  8. [8]

    Image quality assessment: From error visibility to structural similarity,

    Z. Wang, A. C. Bovik, H. R. Sheikh, and E. P. Simoncelli, “Image quality assessment: From error visibility to structural similarity,” IEEE Transactions on Image Processing , vol. 13, no. 4, pp. 600–612, 2004

Show all 59 references
  1. [9]

    Video quality assessment using a statistical model of human visual speed perception,

    Z. Wang and Q. Li, “Video quality assessment using a statistical model of human visual speed perception,” Journal of the Optical Society of America. A, vol. 24, no. 12, 2007

  2. [10]

    Viewer response to time-varying video quality,

    D. E. Pearson, “Viewer response to time-varying video quality,” in Human Vision and Electronic Imaging III , vol. 3299, 1998, pp. 16–25

  3. [11]

    Efficient feature extraction, encoding and classification for action recognition,

    V . Kantorov and I. Laptev, “Efficient feature extraction, encoding and classification for action recognition,” in IEEE Conference on Computer Vision and Pattern Recognition , 2014, pp. 2593–2600

  4. [12]

    Subjective and objective quality of experience of free viewpoint videos,

    J. Yan, J. Li, Y . Fang, Z. Che, X. Xia, and Y . Liu, “Subjective and objective quality of experience of free viewpoint videos,” IEEE Transactions on Image Processing , vol. 31, pp. 3896–3907, 2022

  5. [13]

    Study of spatio-temporal modeling in video quality assessment,

    Y . Fang, Z. Li, J. Yan, X. Sui, and H. Liu, “Study of spatio-temporal modeling in video quality assessment,” IEEE Transactions on Image Processing, vol. 32, pp. 2693–2702, 2023

  6. [14]

    Quality assessment of in-the-wild videos,

    D. Li, T. Jiang, and M. Jiang, “Quality assessment of in-the-wild videos,” in ACM International Conference on Multimedia , 2019, pp. 2351–2359

  7. [15]

    FAST-VQA: Efficient end-to-end video quality assessment with fragment sampling,

    H. Wu, C. Chen, J. Hou, L. Liao, A. Wang, W. Sun, Q. Yan, and W. Lin, “FAST-VQA: Efficient end-to-end video quality assessment with fragment sampling,” in European Conference on Computer Vision, 2022, pp. 538–554

  8. [16]

    Analysis of video quality datasets via design of minimalistic video quality models,

    W. Sun, W. Wen, X. Min, L. Lan, G. Zhai, and K. Ma, “Analysis of video quality datasets via design of minimalistic video quality models,” IEEE Transactions on Pattern Analysis and Machine Intelligence , 2024

  9. [17]

    Blind prediction of natural video quality,

    M. Saad, A. Bovik, and C. Charrier, “Blind prediction of natural video quality,” IEEE Transactions on Image Processing , vol. 23, no. 3, pp. 1352–1365, 2014

  10. [18]

    Learning spatiotemporal inter- actions for user-generated video quality assessment,

    H. Zhu, B. Chen, L. Zhu, and S. Wang, “Learning spatiotemporal inter- actions for user-generated video quality assessment,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 33, no. 3, pp. 1031– 1042, 2022

  11. [19]

    A completely blind video integrity oracle,

    A. Mittal, M. A. Saad, and A. C. Bovik, “A completely blind video integrity oracle,” IEEE Transactions on Image Processing, vol. 25, no. 1, pp. 289–300, 2016

  12. [20]

    Two-level approach for no-reference consumer video quality assessment,

    J. Korhonen, “Two-level approach for no-reference consumer video quality assessment,” IEEE Transactions on Image Processing , vol. 28, no. 12, pp. 5923–5938, 2019

  13. [21]

    ChipQA: No-reference video quality prediction via space-time chips,

    J. Ebenezer, Z. Shang, Y . Wu, H. Wei, S. Sethuraman, and A. C. Bovik, “ChipQA: No-reference video quality prediction via space-time chips,” IEEE Transactions on Image Processing, vol. 30, pp. 8059–8074, 2021

  14. [22]

    Spatiotemporal representation learning for blind video quality assessment,

    Y . Liu, J. Wu, L. Li, W. Dong, J. Zhang, and G. Shi, “Spatiotemporal representation learning for blind video quality assessment,” IEEE Trans- actions on Circuits and Systems for Video Technology , vol. 32, no. 6, pp. 3500–3513, 2021

  15. [23]

    Blindly assess quality of in-the-wild videos via quality-aware pre-training and motion perception,

    B. Li, W. Zhang, M. Tian, G. Zhai, and X. Wang, “Blindly assess quality of in-the-wild videos via quality-aware pre-training and motion perception,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 32, no. 9, pp. 5944–5958, 2022

  16. [24]

    Deep residual learning for image recognition,

    K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in IEEE Conference on Computer Vision and Pattern Recognition, 2016, pp. 770–778

  17. [25]

    Learning phrase representations using RNN encoder-decoder for statistical machine translation,

    K. Cho, B. Van Merri ¨enboer, C. Gulcehre, D. Bahdanau, F. Bougares, H. Schwenk, and Y . Bengio, “Learning phrase representations using RNN encoder-decoder for statistical machine translation,” in Conference on Empirical Methods in Natural Language Processing , 2014

  18. [26]

    Exploring video quality assessment on user generated contents from aesthetic and technical perspectives,

    H. Wu, E. Zhang, L. Liao, C. Chen, J. Hou, A. Wang, W. Sun, Q. Yan, and W. Lin, “Exploring video quality assessment on user generated contents from aesthetic and technical perspectives,” in International Conference on Computer Vision , 2023, pp. 20 144–20 154

  19. [27]

    RIRNet: Recurrent-in- recurrent network for video quality assessment,

    P. Chen, L. Li, L. Ma, J. Wu, and G. Shi, “RIRNet: Recurrent-in- recurrent network for video quality assessment,” in ACM International Conference on Multimedia , 2020, pp. 834–842

  20. [28]

    End-to-end blind quality assessment of compressed videos using deep neural networks

    W. Liu, Z. Duanmu, and Z. Wang, “End-to-end blind quality assessment of compressed videos using deep neural networks.” inACM International Conference on Multimedia , 2018, pp. 546–554

  21. [29]

    DisCoVQA: Temporal distortion-content transformers for video quality assessment,

    H. Wu, C. Chen, L. Liao, J. Hou, W. Sun, Q. Yan, and W. Lin, “DisCoVQA: Temporal distortion-content transformers for video quality assessment,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 33, pp. 4840–4854, 2023

  22. [30]

    TSM: Temporal shift module for efficient video understanding,

    J. Lin, C. Gan, and S. Han, “TSM: Temporal shift module for efficient video understanding,” in International Conference on Computer Vision , 2019, pp. 7083–7093

  23. [31]

    Temporal segment networks for action recognition in videos,

    L. Wang, Y . Xiong, Y . Wang, Z.and Qiao, D. Lin, X. Tang, and L. Van Gool, “Temporal segment networks for action recognition in videos,” IEEE Transactions on Pattern Analysis and Machine Intelli- gence, vol. 41, no. 11, pp. 2740–2755, 2018

  24. [32]

    Temporal relational reasoning in videos,

    B. Zhou, A. Andonian, A. Oliva, and A. Torralba, “Temporal relational reasoning in videos,” inEuropean Conference on Computer Vision, 2018, pp. 831–846

  25. [33]

    ECO: Efficient convolutional network for online video understanding,

    M. Zolfaghari, K. Singh, and T. Brox, “ECO: Efficient convolutional network for online video understanding,” in European Conference on Computer Vision, 2018, pp. 695–712

  26. [34]

    Video swin transformer,

    Z. Liu, J. Ning, Y . Cao, Y . Wei, Z. Zhang, S. Lin, and H. Hu, “Video swin transformer,” in IEEE Conference on Computer Vision and Pattern Recognition, 2022, pp. 3192–3201

  27. [35]

    Neighbourhood representative sampling for efficient end-to- end video quality assessment,

    H. Wu, C. Chen, L. Liao, J. Hou, W. Sun, J. Yan, Qi.and Gu, and W. Lin, “Neighbourhood representative sampling for efficient end-to- end video quality assessment,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 45, pp. 15 185–15 202, 2023

  28. [36]

    Zoom-VQA: Patches, frames and clips integration for video quality assessment,

    K. Zhao, K. Yuan, M. Sun, and X. Wen, “Zoom-VQA: Patches, frames and clips integration for video quality assessment,” in IEEE Conference on Computer Vision and Pattern Recognition , 2023, pp. 1302–1310

  29. [37]

    Scaling and masking: A new paradigm of data sampling for image and video quality assessment,

    Y . Liu, Y . Quan, G. Xiao, A. Li, and J. Wu, “Scaling and masking: A new paradigm of data sampling for image and video quality assessment,” in AAAI Conference on Artificial Intelligence , 2024

  30. [38]

    Revisiting the robustness of spatio-temporal modeling in video quality assessment,

    J. Yan, L. Wu, W. Jiang, C. Liu, and F. Shen, “Revisiting the robustness of spatio-temporal modeling in video quality assessment,” Displays, vol. 81, p. 102585, 2024

  31. [39]

    The Konstanz natural video database (KoNViD-1k),

    V . Hosu, F. Hahn, M. Jenadeleh, H. Lin, H. Men, T. Szir ´anyi, S. Li, and D. Saupe, “The Konstanz natural video database (KoNViD-1k),” in Ninth International Conference on Quality of Multimedia Experience , 2017, pp. 1–6

  32. [40]

    Large-scale study of perceptual video quality,

    Z. Sinno and A. C. Bovik, “Large-scale study of perceptual video quality,” IEEE Transactions on Image Processing , vol. 28, no. 2, pp. 612–627, 2018

  33. [41]

    CVD2014—A database for evaluating no-reference 11 video quality assessment algorithms,

    M. Nuutinen, M. Virtanen, T.and Vaahteranoksa, T. Vuori, P. Oittinen, and J. H ¨akkinen, “CVD2014—A database for evaluating no-reference 11 video quality assessment algorithms,” IEEE Transactions on Image Processing, vol. 25, no. 7, pp. 3073–3086, 2016

  34. [42]

    In-capture mobile video distortions: A study of subjective behavior and objective algorithms,

    D. Ghadiyaram, J. Pan, A. Bovik, A. Moorthy, P. Panda, and K. Yang, “In-capture mobile video distortions: A study of subjective behavior and objective algorithms,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 28, no. 9, pp. 2061–2077, 2017

  35. [43]

    Patch-VQ: ‘Patch- ing up’ the video quality problem,

    Z. Ying, M. Mandal, D. Ghadiyaram, and A. Bovik, “Patch-VQ: ‘Patch- ing up’ the video quality problem,” in IEEE Conference on Computer Vision and Pattern Recognition , 2021, pp. 14 019–14 029

  36. [44]

    NTIRE 2023 quality assessment of video enhancement challenge,

    X. Liu, R. Timofte, Y . Dong, Z. Ma, H. Fan, C. Zhu, X. Min, G. Zhai, Z. Jia, M. Agarla et al. , “NTIRE 2023 quality assessment of video enhancement challenge,” in IEEE Conference on Computer Vision and Pattern Recognition, 2023, pp. 1551–1569

  37. [45]

    Attention is all you need,

    A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in Advances in Neural Information Processing Systems , vol. 30, 2017

  38. [46]

    ImageNet large scale visual recognition challenge,

    O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein et al., “ImageNet large scale visual recognition challenge,” International Journal of Computer Vision, vol. 115, pp. 211–252, 2015

  39. [47]

    Final report from the Video Quality Experts Group on the validation of objective models of video quality assessment,

    VQEG, “Final report from the Video Quality Experts Group on the validation of objective models of video quality assessment,” [online]. Available: http://www.vqeg.org, 2003

  40. [48]

    FASTER recurrent networks for efficient video classification,

    L. Zhu, D. Tran, L. Sevilla-Lara, Y . Yang, M. Feiszli, and H. Wang, “FASTER recurrent networks for efficient video classification,” in AAAI Conference on Artificial Intelligence , vol. 34, no. 07, 2020, pp. 13 098– 13 105

  41. [49]

    Qoe evaluation for live broadcasting video,

    P. Chen, L. Li, Y . Huang, F. Tan, and W. Chen, “Qoe evaluation for live broadcasting video,” in IEEE International Conference on Image Processing, 2019, pp. 454–458

  42. [50]

    Study of the subjective and objective quality of high motion live streaming videos,

    Z. Shang, J. P. Ebenezer, Y . Wu, H. Wei, S. Sethuraman, and A. C. Bovik, “Study of the subjective and objective quality of high motion live streaming videos,” IEEE Transactions on Image Processing, vol. 31, pp. 1027–1041, 2021

  43. [51]

    A study of subjective and objec- tive quality assessment of hdr videos,

    Z. Shang, J. P. Ebenezer, A. K. Venkataramanan, Y . Wu, H. Wei, S. Sethuraman, and A. C. Bovik, “A study of subjective and objec- tive quality assessment of hdr videos,” IEEE Transactions on Image Processing, vol. 33, pp. 42–57, 2023

  44. [52]

    Hidro-vqa: High dynamic range oracle for video quality assessment,

    S. Saini, A. Saha, and A. C. Bovik, “Hidro-vqa: High dynamic range oracle for video quality assessment,” in IEEE/CVF Winter Conference on Applications of Computer Vision , 2024, pp. 469–479

  45. [53]

    MobileNets: Efficient convolutional neural networks for mobile vision applications,

    A. Howard, M. Zhu, B. Chen, D. Kalenichenko, W. Wang, T. Weyand, M. Andreetto, and H. Adam, “MobileNets: Efficient convolutional neural networks for mobile vision applications,” ArXiv Preprint ArXiv:1704.04861, 2017

  46. [54]

    GraphIQA: Learning distortion graph representations for blind image quality assessment,

    S. Sun, T. Yu, J. Xu, W. Zhou, and Z. Chen, “GraphIQA: Learning distortion graph representations for blind image quality assessment,” IEEE Transactions on Multimedia , vol. 25, pp. 2912–2925, 2022

  47. [55]

    Run, don’t walk: Chasing higher flops for faster neural networks,

    J. Chen, S. Kao, H. He, W. Zhuo, S. Wen, C. Lee, and S. Chan, “Run, don’t walk: Chasing higher flops for faster neural networks,” in IEEE Conference on Computer Vision and Pattern Recognition , 2023, pp. 12 021–12 031

  48. [56]

    EfficientNet: Rethinking model scaling for con- volutional neural networks,

    M. Tan and Q. Le, “EfficientNet: Rethinking model scaling for con- volutional neural networks,” in International Conference on Machine Learning, 2019, pp. 6105–6114

  49. [57]

    ShuffleNet: An extremely efficient convolutional neural network for mobile devices,

    X. Zhang, X. Zhou, M. Lin, and J. Sun, “ShuffleNet: An extremely efficient convolutional neural network for mobile devices,” in IEEE Conference on Computer Vision and Pattern Recognition , 2018, pp. 6848–6856

  50. [58]

    MobileOne: An improved one millisecond mobile backbone,

    P. Vasu, J. Gabriel, J. Zhu, O. Tuzel, and A. Ranjan, “MobileOne: An improved one millisecond mobile backbone,” in IEEE Conference on Computer Vision and Pattern Recognition , 2023, pp. 7907–7917

  51. [59]

    Long short-term memory,

    A. Graves and A. Graves, “Long short-term memory,” Supervised Sequence Labelling with Recurrent Neural Networks , pp. 37–45, 2012. Jiebin Yan received the Ph.D. degree from Jiangxi University of Finance and Economics, Nanchang, China. He was a computer vision engineer with MT-...

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.