REVIEW 2 major objections 1 minor 1 cited by
3DGS-VBench: A Comprehensive Video Quality Evaluation Benchmark for 3DGS Compression
T0 review · 2 major / 1 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read 3DGS-VBench provides a large-scale video quality benchmark for 3D Gaussian Splatting compression, with 660 compressed models and MOS scores from 50 raters.
desk verdict The abstract promises a 3DGS video-quality benchmark, but the submitted full text is an unrelated math.NA paper—there is no supporting body for any claim. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the 3DGS-VBench dataset/benchmark itself: a pipeline that starts from 11 3DGS scenes, applies 6 compression algorithms at designed parameter levels, renders the compressed models into video sequences, and annotates them with MOS from 50 raters. This machinery simultaneously provides training data for VQA models and a controlled testbed for ranking algorithms and metrics.
What would settle it
Compare the benchmark's MOS-based ranking of the 6 compression algorithms against a perceptual study using a different set of scenes and bitrates; if the relative quality order changes, the benchmark's representativeness claim fails. Alternatively, measure inter-rater agreement (e.g., Krippendorff's alpha) on the collected scores; low agreement would undermine the MOS reliability.
Extended reading notes
Core claim
The central discovery the paper aims to establish is that a purpose-built VQA benchmark for 3DGS compression is feasible and reliable: 660 compressed 3DGS models from 11 scenes across 6 algorithms, rated by 50 participants, yield MOS scores that, after outlier removal, allow trustworthy comparison of compression methods and quality metrics. The 15 metrics are evaluated across paradigms to see which ones align with human perception on the unique distortions 3DGS compression produces.
Load-bearing premise
The benchmark's rankings transfer only if the 11 scenes and chosen compression parameter levels represent the types of content and distortion levels that matter in real 3DGS deployments, and if MOS from 50 raters, after outlier removal, is a stable ground truth for these novel distortions.
Editorial extensions
If this is right
- If correct, 3DGS compression researchers can use 3DGS-VBench as a standard testbed, making algorithm comparisons in new papers directly comparable.
- The MOS data can train a specialized 3DGS VQA model that outperforms generic metrics on these distortions.
- The metric evaluation tells which existing metrics are trustworthy for 3DGS compression, so practitioners can use them without costly human studies.
- The benchmark's release creates a common ground for rate-distortion trade-off analysis in 3DGS.
- The methodology can be extended to other view-synthesis formats.
Reading between the lines
- The 15-metric rankings may shift when more content types or higher distortion levels are added, since the 11 scenes and parameter levels are a sample of a larger distortion space.
- MOS from 50 participants, though standard, may not capture preference heterogeneity; a future benchmark with more raters per video could yield different algorithm rankings.
- A VQA model trained on this benchmark may overfit to the 6 included compression algorithms and not generalize to unseen codecs.
- Editorial note: the full text accompanying this abstract is a different manuscript on disk-domain interpolation, so the claims above are based on the abstract alone and may not match the paper's actual experiments.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The submission, as identified by its abstract and title, claims to introduce 3DGS-VBench, a large-scale Video Quality Assessment (VQA) benchmark for 3D Gaussian Splatting compression, comprising 660 compressed 3DGS models, 11 scenes, 6 compression algorithms, MOS scores from 50 participants with outlier removal, and an evaluation of 15 quality metrics. The full text supplied, however, is a completely different manuscript titled "A novel interpolation–regression approach for function approximation on the disk and its application to cubature formulas" (arXiv:2508.07047), with no content related to 3DGS, compression, VQA, MOS, or the claimed benchmark.
Significance. If the claimed benchmark existed as described, it could be a valuable community resource for training and evaluating 3DGS-specific VQA models and for comparing compression algorithms. The stated scale (660 models, 50 participants, 15 metrics) is plausible and would fill a gap in the literature. However, the submitted manuscript contains no evidence of the dataset's existence, design, or validation. The claimed contributions are therefore unverifiable from the submitted text, and the work as presented has no discernible scientific content that can be assessed.
major comments (2)
- [Full text (entire manuscript)] The submitted full text is arXiv:2508.07047, a mathematics paper on interpolation–regression on the disk and cubature formulas. It contains no mention of 3DGS, compression, video quality assessment, MOS, or the benchmark. The central load-bearing claims of the abstract—the existence of the dataset, its 660 compressed models, 11 scenes, 6 algorithms, 50-participant MOS annotations, outlier removal, and reliability validation—are entirely unsupported by the manuscript body. This is a fundamental mismatch, not a local omission.
- [Abstract / Dataset availability] The abstract states that the dataset is available at a GitHub URL, but no code, data, or documentation is provided in the manuscript. The reliability claim ("validated dataset reliability") is made without any supporting statistics, protocols, or analysis. No inter-rater agreement, MOS distribution, or outlier-removal criterion is given. Even if the body mismatch were corrected, these methodological details would be necessary for a benchmark paper.
minor comments (1)
- [Title and metadata] The title, author list, and abstract do not match the content of the full text. The arXiv identifier of the full text is 2508.07047, not 2508.07038 as referenced. This suggests a submission error that must be resolved by the authors.
Circularity Check
No circular derivation found: the submitted body is an unrelated numerical-analysis paper, so no claimed reduction can be checked.
full rationale
The abstract of arXiv:2508.07038 describes a 3DGS compression benchmark with 660 compressed models, MOS from 50 participants, and evaluations of 15 quality metrics. The supplied 'full text,' however, is arXiv:2508.07047, 'A novel interpolation–regression approach for function approximation on the disk and its application to cubature formulas,' which contains no mention of 3DGS, video quality assessment, compression, MOS, or the benchmark. Because the body provides none of the benchmark's construction details, no derivation chain, fitted parameter, self-citation argument, or definitional reduction can be inspected. The absence of supporting evidence for the abstract's claims is a serious completeness/correctness problem, but it is not circularity under the required standard: there is no quotable equation or construction showing that a stated result reduces to its own inputs. Standard benchmark practice would not make the metric-vs-MOS evaluation circular even if the full benchmark text were present, since ranking metrics against human labels is the intended use of a VQA dataset. Therefore the circularity score is 0.
Assumptions & free parameters
free parameters (3)
- per-algorithm compression parameter levels
- MOS outlier-removal criterion
- scene selection (11 scenes)
assumptions (3)
- domain assumption MOS from 50 participants after outlier removal is a reliable ground truth for perceived quality of 3DGS compression distortions.
- domain assumption The 11 scenes and the 6 compression algorithms span the practical distortion space of 3DGS compression.
- domain assumption The 15 evaluated quality metrics represent the relevant metric paradigms for 3DGS-generated video.
Cite this review
Pith. "Pith review of 3DGS-VBench: A Comprehensive Video Quality Evaluation Benchmark for 3DGS Compression." pith.science (2026). https://pith.science/paper/6R2T2WKQ
@misc{pith2026250807038,
author = {Pith},
title = {Pith review of: 3DGS-VBench: A Comprehensive Video Quality Evaluation Benchmark for 3DGS Compression},
year = {2026},
howpublished = {\url{https://pith.science/paper/6R2T2WKQ}},
note = {Machine review of arXiv:2508.07038}
}
read the original abstract
3D Gaussian Splatting (3DGS) enables real-time novel view synthesis with high visual fidelity, but its substantial storage requirements hinder practical deployment, prompting state-of-the-art (SOTA) 3DGS methods to incorporate compression modules. However, these 3DGS generative compression techniques introduce unique distortions lacking systematic quality assessment research. To this end, we establish 3DGS-VBench, a large-scale Video Quality Assessment (VQA) Dataset and Benchmark with 660 compressed 3DGS models and video sequences generated from 11 scenes across 6 SOTA 3DGS compression algorithms with systematically designed parameter levels. With annotations from 50 participants, we obtained MOS scores with outlier removal and validated dataset reliability. We benchmark 6 3DGS compression algorithms on storage efficiency and visual quality, and evaluate 15 quality assessment metrics across multiple paradigms. Our work enables specialized VQA model training for 3DGS, serving as a catalyst for compression and quality assessment research. The dataset is available at https://github.com/YukeXing/3DGS-VBench.
Forward citations
Cited by 1 Pith paper
-
3DGSI-Assessor: A Large-Scale Dataset and An LMM-based Method for 3D Gaussian Splatting Image Quality Assessment
A new 15,200-image human-annotated dataset and an LMM-based metric that jointly predicts overall, geometry, and color quality of compressed 3D Gaussian Splatting images.
Reference graph
Works this paper leans on
- [1]
-
[2]
B. Mildenhall, P. P. Srinivasan, M. Tancik, J. T. Barron, R. Ramamoorthi, and R. Ng, ``Nerf: representing scenes as neural radiance fields for view synthesis,'' Commun. ACM , vol. 65, no. 1, p. 99–106, 2021
work page 2021
-
[3]
T. Lu, M. Yu, L. Xu, Y. Xiangli, L. Wang, D. Lin, and B. Dai, ``Scaffold-gs: Structured 3d gaussians for view-adaptive rendering,'' in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pp. 20654--20664, 2024
work page 2024
-
[4]
Y. Chen, Q. Wu, W. Lin, M. Harandi, and J. Cai, ``Hac: Hash-grid assisted context for 3d gaussian splatting compression,'' in European Conference on Computer Vision (ECCV) , pp. 422--438, Springer, 2024
work page 2024
-
[5]
Z. Fan, K. Wang, K. Wen, Z. Zhu, D. Xu, Z. Wang, et al. , ``Lightgaussian: Unbounded 3d gaussian compression with 15x reduction and 200+ fps,'' Advances in neural information processing systems , vol. 37, pp. 140138--140158, 2024
work page 2024
-
[6]
K. Navaneet, K. Pourahmadi Meibodi, S. Abbasi Koohpayegani, and H. Pirsiavash, ``Compgs: Smaller and faster gaussian splatting with vector quantization,'' in European Conference on Computer Vision , pp. 330--349, Springer, 2024
work page 2024
-
[7]
S. Niedermayr, J. Stumpfegger, and R. Westermann, ``Compressed 3d gaussian splatting for accelerated novel view synthesis,'' in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pp. 10349--10358, June 2024
work page 2024
-
[8]
J. C. Lee, D. Rho, X. Sun, J. H. Ko, and E. Park, ``Compact 3d gaussian splatting for static and dynamic radiance fields,'' arXiv preprint arXiv:2408.03822 , 2024
arXiv 2024
Show all 40 references
-
[9]
Girish, K
S. Girish, K. Gupta, and A. Shrivastava, ``Eagles: Efficient accelerated 3d gaussians with lightweight encodings,'' in European Conference on Computer Vision , pp. 54--71, Springer, 2024
2024
-
[10]
X. Liu, X. Wu, P. Zhang, S. Wang, Z. Li, and S. Kwong, ``Compgs: Efficient 3d scene representation via compressed gaussian splatting,'' in Proceedings of the 32nd ACM International Conference on Multimedia , pp. 2936--2944, 2024
2024
-
[11]
Q. Yang, L. Yang, G. Van Der Auwera, and Z. Li, ``Hybridgs: High-efficiency gaussian splatting data compression using dual-channel sparse representation and point cloud encoder,'' arXiv preprint arXiv:2505.01938 , 2025
2025 arXiv
-
[12]
Martin, A
P. Martin, A. Rodrigues, J. Ascenso, and M. P. Queluz, ``Nerf-qa: Neural radiance fields quality assessment database,'' in 2023 15th International Conference on Quality of Multimedia Experience (QoMEX) , pp. 107--110, 2023
2023
-
[13]
Martin, A
P. Martin, A. Rodrigues, J. Ascenso, and M. Paula Queluz, ``Nerf view synthesis: Subjective quality assessment and objective metrics evaluation,'' IEEE Access , vol. 13, pp. 26--41, 2025
2025
-
[14]
Liang, T
H. Liang, T. Wu, P. Hanji, F. Banterle, H. Gao, R. Mantiuk, and C. \"O ztireli, ``Perceptual quality assessment of nerf and neural view synthesis methods for front-facing views,'' in Computer Graphics Forum , vol. 43, p. e15036, Wiley Online Library, 2024
2024
-
[15]
Y. Xing, Q. Yang, K. Yang, Y. Xu, and Z. Li, ``Explicit-nerf-qa: A quality assessment database for explicit nerf model compression,'' in 2024 IEEE International Conference on Visual Communications and Image Processing (VCIP) , pp. 1--5, 2024
2024
-
[16]
Q. Yang, K. Yang, Y. Xing, Y. Xu, and Z. Li, ``A benchmark for gaussian splatting compression and quality assessment study,'' in Proceedings of the 6th ACM International Conference on Multimedia in Asia , pp. 1--8, 2024
2024
-
[18]
Zhang, J
Y. Zhang, J. Maraval, Z. Zhang, N. Ramin, S. Tian, and L. Zhang, ``Evaluating human perception of novel view synthesis: Subjective quality assessment of gaussian splatting and nerf in dynamic scenes,'' arXiv preprint arXiv:2501.08072 , 2025
2025 arXiv
-
[19]
J. T. Barron, B. Mildenhall, D. Verbin, P. P. Srinivasan, and P. Hedman, ``Mip-nerf 360: Unbounded anti-aliased neural radiance fields,'' in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition (CVPR) , pp. 5470--5479, 2022
2022
-
[20]
Knapitsch, J
A. Knapitsch, J. Park, Q.-Y. Zhou, and V. Koltun, ``Tanks and temples: Benchmarking large-scale scene reconstruction,'' ACM Transactions on Graphics (ToG) , vol. 36, no. 4, pp. 1--13, 2017
2017
-
[21]
Hedman, J
P. Hedman, J. Philip, T. Price, J.-M. Frahm, G. Drettakis, and G. Brostow, ``Deep blending for free-viewpoint image-based rendering,'' ACM Transactions on Graphics (ToG) , vol. 37, no. 6, pp. 1--15, 2018
2018
-
[22]
Zheng, L
X. Zheng, L. Liao, X. Li, J. Jiao, R. Wang, F. Gao, S. Wang, and R. Wang, ``Pku-dymvhumans: A multi-view video benchmark for high-fidelity dynamic human modeling,'' in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pp. 22530--22540, 2024
2024
-
[23]
ITU-T RECOMMENDATION, ``Subjective video quality assessment methods for multimedia applications,'' 1999
P. ITU-T RECOMMENDATION, ``Subjective video quality assessment methods for multimedia applications,'' 1999
1999
-
[24]
Z. Wang, A. Bovik, H. Sheikh, and E. Simoncelli, ``Image quality assessment: From error visibility to structural similarity,'' IEEE Transactions on Image Processing , vol. 13, no. 4, pp. 600--612, 2004
2004
-
[25]
Zhang, P
R. Zhang, P. Isola, A. A. Efros, E. Shechtman, and O. Wang, ``The unreasonable effectiveness of deep features as a perceptual metric,'' in 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition , pp. 586--595, 2018
2018
-
[26]
K. Ding, K. Ma, S. Wang, and E. P. Simoncelli, ``Image quality assessment: Unifying structure and texture similarity,'' IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 44, no. 5, pp. 2567--2581, 2022
2022
-
[27]
Sheikh and A
H. Sheikh and A. Bovik, ``Image information and visual quality,'' IEEE Transactions on Image Processing , vol. 15, no. 2, pp. 430--444, 2006
2006
-
[28]
Zhang, L
L. Zhang, L. Zhang, X. Mou, and D. Zhang, ``Fsim: A feature similarity index for image quality assessment,'' IEEE Transactions on Image Processing , vol. 20, no. 8, pp. 2378--2386, 2011
2011
-
[29]
Wang and Q
Z. Wang and Q. Li, ``Information content weighting for perceptual image quality assessment,'' IEEE Transactions on Image Processing , vol. 20, no. 5, pp. 1185--1198, 2011
2011
-
[30]
Z. Wang, E. Simoncelli, and A. Bovik, ``Multiscale structural similarity for image quality assessment,'' in The Thrity-Seventh Asilomar Conference on Signals, Systems & Computers, 2003 , vol. 2, pp. 1398--1402 Vol.2, 2003
2003
-
[31]
J. Wang, K. C. Chan, and C. C. Loy, ``Exploring clip for assessing the look and feel of images,'' in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 37, pp. 2555--2563, 2023
2023
-
[32]
Mittal, A
A. Mittal, A. K. Moorthy, and A. C. Bovik, ``No-reference image quality assessment in the spatial domain,'' IEEE Transactions on Image Processing , vol. 21, no. 12, pp. 4695--4708, 2012
2012
-
[33]
H. Wu, E. Zhang, L. Liao, C. Chen, J. Hou, A. Wang, W. Sun, Q. Yan, and W. Lin, ``Exploring video quality assessment on user generated contents from aesthetic and technical perspectives,'' in Proceedings of the IEEE/CVF International Conference on Computer Vision , pp. 20144--...
2023
-
[34]
H. Wu, C. Chen, J. Hou, L. Liao, A. Wang, W. Sun, Q. Yan, and W. Lin, ``Fast-vqa: Efficient end-to-end video quality assessment with fragment sampling,'' in European conference on computer vision , pp. 538--554, Springer, 2022
2022
-
[35]
W. Sun, X. Min, W. Lu, and G. Zhai, ``A deep learning based no-reference quality assessment model for ugc videos,'' in Proceedings of the 30th ACM International Conference on Multimedia , pp. 856--865, 2022
2022
-
[36]
D. Li, T. Jiang, and M. Jiang, ``Quality assessment of in-the-wild videos,'' in Proceedings of the 27th ACM international conference on multimedia , pp. 2351--2359, 2019
2019
-
[37]
H. Wu, Z. Zhang, W. Zhang, C. Chen, L. Liao, C. Li, Y. Gao, A. Wang, E. Zhang, W. Sun, et al. , ``Q-align: Teaching lmms for visual scoring via discrete text-defined levels,'' arXiv preprint arXiv:2312.17090 , 2023
2023 arXiv
-
[38]
Q. Yang, Y. Zhang, S. Chen, Y. Xu, J. Sun, and Z. Ma, ``Mped: Quantifying point cloud distortion based on multiscale potential energy discrepancy,'' IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 45, no. 5, pp. 6037--6054, 2023
2023
-
[39]
Eason, B
G. Eason, B. Noble, and I. N. Sneddon, ``On certain integrals of Lipschitz-Hankel type involving products of Bessel functions,'' Phil. Trans. Roy. Soc. London, vol. A247, pp. 529--551, April 1955
1955
-
[40]
Clerk Maxwell, A Treatise on Electricity and Magnetism, 3rd ed., vol
J. Clerk Maxwell, A Treatise on Electricity and Magnetism, 3rd ed., vol. 2. Oxford: Clarendon, 1892, pp.68--73
-
[41]
completely blind
11em plus .33em minus .07em 4000 4000 100 4000 4000 500 `\.=1000 = #1 \@IEEEnotcompsoconly \@IEEEcompsoconly #1 * [1] 0pt [0pt][0pt] #1 * [1] 0pt [0pt][0pt] #1 * \| ** #1 \@IEEEauthorblockNstyle \@IEEEcompsocnotconfonly \@IEEEauthorblockAstyle \@IEEEcompsocnotconfonly \@IEEEco...
2023 arXiv
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.