REVIEW 3 major objections 4 minor 65 references
Puzzle Similarity: A Perceptually-guided Cross-Reference Metric for Artifact Detection in 3D Scene Reconstructions
T0 review · 3 major / 4 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read Puzzle Similarity claims to outperform every tested full-reference, cross-reference, and no-reference metric at localizing artifacts in novel views of 3D reconstructions, using only training-view patches as references.
desk verdict Useful dataset and a sensible metric, but the reported SOTA comes from tuning on the test set. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the similarity map $S(I)$ defined by Eq. (1) through Eq. (3): for each CNN layer $\ell$, each feature vector of the test embedding is matched to its nearest neighbor, meaning the maximum cosine similarity, among all spatially flattened feature vectors of all $N$ training views, and the per-layer maps are upsampled and affinely combined with weights $w_2=0.67$, $w_3=0.2$, $w_4=0.13$ in the SqueezeNet backbone. This single maximum-match-to-any-reference-patch operation carries the entire argument, because it defines an artifact as a region with no close puzzle piece in the training views. The layered CNN embedding provides perceptual alignment and lets the metric localize artifacts at multiple scales.
What would settle it
Take a well-reconstructed novel view, duplicate a small texture patch and paste it at a different location, or swap two similarly textured regions, then have human observers mark artifacts while Puzzle Similarity scores the same image. If humans flag the duplication but the metric's similarity stays near its maximum because each pasted patch exactly matches a reference patch, the assumption that artifact regions never match reference patches fails.
Extended reading notes
Core claim
The central claim is that Puzzle Similarity outperforms all tested full-reference, cross-reference and no-reference metrics in capturing artifacts aligned with human perception, as measured on the authors' human-labeled dataset. The metric computes, for each pixel of the embedded test view, the maximum cosine similarity against all feature vectors extracted from all training views across several CNN layers, then combines the per-layer maps with fixed weights. Because the search ignores spatial position, it is robust to camera shifts; because it never needs a reference aligned to the test view, it applies in the setting where reconstruction-quality assessment is hardest. The claim is supported by an average Pearson correlation of 0.615 (std 0.120), ahead of the cross-reference baseline at 0.510 (std 0.204) and the best no-reference method at 0.402 (std 0.178), with lower variance indicating more consistent performance across artifact types.
Load-bearing premise
The metric assumes that every correctly reconstructed region has at least one closely matching patch somewhere in the training views and that every artifact region has none, with no check on whether the matched patches are spatially consistent; artifacts made of valid patches in wrong places (duplicated texture, ghosting, swapped parts) can therefore score as high as clean regions.
Editorial extensions
If this is right
- On the new human-labeled dataset, Puzzle Similarity's artifact maps correlate with human segmentations better on average than every no-reference, cross-reference, and full-reference metric tested, with smaller variance across the 12 scenes.
- Because the metric needs only the training views of the scene and a pretrained CNN, it can flag artifacts in any new view without a ground-truth image, which is the setting where reconstruction quality is otherwise hardest to assess.
- The resulting artifact masks can drive automatic inpainting: the paper's iterative thresholding framework uses Puzzle Similarity to select masks and reports monotone improvement in similarity.
- Swapping the CNN backbone adapts the metric to a new domain with no retraining, so the same cross-reference principle transfers beyond 3D reconstruction to any image set that defines a distribution.
- The released human-labeled artifact dataset gives other cross-reference metrics a benchmark with ground-truth maps, which previously did not exist.
Reading between the lines
- The metric's blind spot is spatial rearrangement: because Eq. (1) maximizes similarity over all reference patches regardless of location, a view with duplicated or swapped textures can receive high scores even where humans see artifacts; adding a spatial-consistency or epipolar check would close this gap.
- One can extend the evaluation to synthetic perturbations: replaying the same scene with known pasted or ghosted artifacts would quantify exactly how much of Puzzle Similarity's human correlation comes from local texture mismatch rather than geometric correctness.
- A natural relaxation of the hard max operation is a softmax over reference patches, which could make the metric more suitable for gradient-based optimization; the paper notes that the current max operation is unlikely to produce useful gradients.
- The reported standard deviations (0.120 vs 0.204 for the cross-reference baseline) suggest the advantage is as much about consistency across artifact types as about average accuracy, and scenes with blurry or unnatural textures are where the gap appears largest.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes Puzzle Similarity, a cross-reference metric for detecting and localizing artifacts in novel views of 3D scene reconstructions. The metric embeds image patches using a pretrained CNN (SqueezeNet), then for each patch in the test view computes the maximum cosine similarity against all patches in a set of unaligned reference views, across all spatial positions. The similarity maps from layers 2, 3, and 4 are weighted (0.67, 0.20, 0.13) and combined into a final artifact map. The authors collect a new human-labeled dataset of artifact masks for 36 renderings from 12 scenes reconstructed with 3D Gaussian Splatting, and evaluate the metric's correlation with averaged human masks, comparing against no-reference, cross-reference, and full-reference baselines. They report that Puzzle Similarity achieves the highest average Pearson and Spearman correlations (Tab. 2), and demonstrate an iterative inpainting application in Sec. 5.
Significance. If the reported results hold under a properly controlled evaluation, Puzzle Similarity would be a practical and cheap artifact-localization tool for 3D reconstruction, with the advantage of not requiring aligned references. The paper also contributes a public human-labeled dataset for cross-reference metric evaluation, which is a useful resource for the community. The method's strengths are its clarity, simplicity, differentiability, and the efficient blockwise implementation described in Sec. 3. However, the central state-of-the-art claim is currently weakened by the evaluation protocol, in which the method's design choices were selected on the same test examples used for the final comparison, and by the small hand-picked evaluation set. The significance of the claimed advance is therefore conditional on additional validation.
major comments (3)
- [Sec. 3 (Pre-trained Model Choice) and Sec. 4.2, Tab. 2] The backbone, layer set, and layer weights are selected on the same data used for the final evaluation. The paper states in Sec. 3 that SqueezeNet 'aligned best with our test examples, specifically using layers ℓ∈{2,3,4} with the weights w2=0.67, w3=0.2, and w4=0.13, which we found heuristically.' Since these are the same 36 renderings (Sec. 4.1) used to compute the headline correlations in Tabs. 1 and 2, the reported 0.615 vs. 0.510 average Pearson advantage over CrossScore may reflect test-set tuning rather than a genuine property of the metric. The baselines are applied off-the-shelf without dataset-specific tuning, so the comparison is asymmetric. The authors should either introduce a separate validation split for choosing hyperparameters, or report a sensitivity analysis showing that the advantage persists across a range of weights, layer sets, and backbones. This is load-bearing for the conclusion that Puzzle Similarity 'outperforms all tested' metrics.
- [Sec. 3, Eq. (1); Sec. 4.1 dataset composition] The metric's core assumption is that every well-reconstructed region has a near-perfect patch match somewhere in the reference views and that every artifact region lacks such a match. However, Eq. (1) takes a spatial max over all reference patches, so artifacts composed of valid patches in invalid spatial arrangements (e.g., ghosting, duplicated texture, swapped scene parts) can match a reference patch very well and be incorrectly labeled as clean. The paper even notes in Sec. 3 that the max search 'relinquishes spatial relations.' The current dataset, which the authors state is dominated by blur, holes, and texture mismatches, cannot exercise this failure mode. The authors should test the metric on examples with spatial rearrangement artifacts, or explicitly acknowledge and bound this limitation. This is a principal risk for the claimed perceptual alignment.
- [Sec. 4.1 and Sec. 4.5; Conclusion] The conclusion claims that Puzzle Similarity outperforms 'all tested full-reference, cross-reference and no-reference metrics,' but the full-reference comparison is not reported in the main text. Sec. 4.5 only states that 'we also provide an extensive comparison' and mentions results in qualitative terms; the actual FR numbers are deferred to the Supplementary. Since the central claim explicitly includes FR metrics, the main text should display the FR results (at least a summary table comparable to Tab. 2). Without this, the stated scope of the claim cannot be verified from the paper as presented. Additionally, the evaluation is based on only 36 hand-picked images (3 per scene, 12 scenes), and the per-scene correlations in Tab. 1 are computed from just three images each, which limits the statistical strength of the average comparison.
minor comments (4)
- [Sec. 3, Eq. (4)] The outer-product formulation in Eq. (4) is a useful implementation detail, but the notation is slightly inconsistent: the text defines Sℓ(I) as a map, while the equation appears to express the row-max of a matrix; please clarify the dimensions and the flattening order (N, Hℓ, Wℓ, Cℓ) to make the memory-efficient tiling easier to follow.
- [Sec. 4.2, Eq. (5)] The 5-parameter logistic fit is applied per scene and per metric; please state explicitly whether the logistic parameters are fit on the same human masks used for the reported correlation values. If so, this is a standard but potentially optimistic procedure, and it should be described in sufficient detail to assess whether all metrics benefit equally.
- [Sec. 6 (Limitations)] The Limitations section already acknowledges that the layer weights are empirically calibrated, but it does not mention that the calibration was performed on the test data. Please add an explicit statement about the validation protocol and, ideally, point to a sensitivity analysis that reassures readers the result is not an artifact of overfitting to the 36 images.
- [Sec. 4.1, participant details] The paper reports 22 participants but only states that gender and age distributions are in the Supplementary. Please move at least the participant count and any exclusion criteria into the main text, as the quality of the ground-truth labels is central to the evaluation.
Circularity Check
Evaluation-set tuning of the backbone and layer weights undermines the reported state-of-the-art correlation, while the inpainting 'monotonic improvement' is a self-referential selection rule.
-
fitted input called prediction
[Section 3 (Pre-trained Model Choice) with Section 4.1 and Tables 1-2]
"We opted for SqueezeNet as it aligned best with our test examples, specifically using layers ℓ∈{2,3,4} with the weights w2 = 0.67, w3 = 0.2, and w4 = 0.13, which we found heuristically. ... For each dataset, we selected three renderings that demonstrated a mix of well-reconstructed areas, strong artifacts, and subtle artifacts, resulting in 36 samples across 12 datasets."
The 'test examples' used to select SqueezeNet, the layer set, and the weights are the same 36 human-labeled renderings on which Tables 1 and 2 report the state-of-the-art correlations. No validation split or sensitivity analysis is described. The reported 0.615 average Pearson correlation is therefore an in-sample selection, not an out-of-sample prediction of human alignment: a few free choices were adjusted to maximize agreement with these exact labels, while the baselines were applied off-the-shelf as published. The conclusion that Puzzle Similarity 'outperforms all tested full-reference, cross-reference and no-reference metrics' is thus not independently supported by the experiment as reported.
-
self definitional
[Section 5, Progressive Inpainting]
"The quality of each inpainted candidate is evaluated by calculating the average similarity difference before and after inpainting, denoted as δi. ... We then select the candidate that maximizes δ. ... This framework guarantees a monotonic improvement in PuzzleSim similarity."
The 'guarantee' is a restatement of the selection rule: δi is defined as the PuzzleSim difference produced by candidate i, and the algorithm chooses the candidate with maximal δ, terminating when max δi ≤ 0. The reported monotonic improvement is therefore true by construction in the metric being optimized, not an independent empirical validation that the inpaintings align better with human perception. This is a minor, self-referential application detail rather than the main evidence for the paper's headline claim.
full rationale
The core of Puzzle Similarity is self-contained: Eqs. (1)-(3) define the score map as a maximum cosine similarity over reference-patch embeddings, and no human label or benchmark outcome enters the map computation itself. The paper's central claim, however, is supported by a comparison in which the method's design choices were selected on the benchmark used for evaluation. Section 3 states that SqueezeNet, the layer set {2,3,4}, and weights (0.67, 0.20, 0.13) were chosen because SqueezeNet 'aligned best with our test examples' and the weights were 'found heuristically.' Section 4.1 identifies those test examples as the 36 renderings used for Tables 1 and 2. No validation split or sensitivity analysis is described, and the competing metrics were applied as published without equivalent tuning. The reported average Pearson correlation of 0.615 versus CrossScore's 0.510 is therefore an in-sample selection rather than an out-of-sample prediction; this is a partial circularity of the 'fitted input called prediction' type. Separately, the inpainting application's 'guaranteed monotonic improvement in PuzzleSim similarity' is a restatement of the selection rule (choose the candidate maximizing the PuzzleSim delta), so it is a self-referential optimization rather than an independent validation. These issues are partial, not total: the metric itself does not reduce to the human labels by construction, no load-bearing self-citation chain is present, and the tuned choices involve only a few free parameters. The headline superiority claim nonetheless is not supported by the experiment as reported.
Assumptions & free parameters
free parameters (4)
- Layer combination weights w2, w3, w4 =
0.67, 0.20, 0.13
- Layer set {2,3,4} and downsampling limit =
three halvings
- Feature backbone (SqueezeNet) =
SqueezeNet
- 5-parameter logistic fit (a1..a5) per scene =
fit per scene to maximize PCC/SRCC
assumptions (5)
- domain assumption ImageNet-pretrained CNN features are valid perceptual embeddings for reconstruction artifacts.
- domain assumption Training views used as references are artifact-free (or their artifacts are negligible).
- domain assumption Maximum over all spatial positions yields viewpoint invariance without harming artifact detection.
- domain assumption Average of 22 binary human masks is a reliable ground truth for artifact localization.
- domain assumption 5-parameter logistic mapping is an appropriate nonlinear calibration for all compared metrics.
Cite this review
Pith. "Pith review of Puzzle Similarity: A Perceptually-guided Cross-Reference Metric for Artifact Detection in 3D Scene Reconstructions." pith.science (2026). https://pith.science/paper/Q7ZFLVKW
@misc{pith2026241117489,
author = {Pith},
title = {Pith review of: Puzzle Similarity: A Perceptually-guided Cross-Reference Metric for Artifact Detection in 3D Scene Reconstructions},
year = {2026},
howpublished = {\url{https://pith.science/paper/Q7ZFLVKW}},
note = {Machine review of arXiv:2411.17489}
}
read the original abstract
Modern reconstruction techniques can effectively model complex 3D scenes from sparse 2D views. However, automatically assessing the quality of novel views and identifying artifacts is challenging due to the lack of ground truth images and the limitations of no-reference image metrics in predicting reliable artifact maps. The absence of such metrics hinders assessment of the quality of novel views and limits the adoption of post-processing techniques, such as inpainting, to enhance reconstruction quality. To tackle this, recent work has established a new category of metrics (cross-reference), predicting image quality solely by leveraging context from alternate viewpoint captures (arXiv:2404.14409). In this work, we propose a new cross-reference metric, Puzzle Similarity, which is designed to localize artifacts in novel views. Our approach utilizes image patch statistics from the training views to establish a scene-specific distribution, later used to identify poorly reconstructed regions in the novel views. Given the lack of good measures to evaluate cross-reference methods in the context of 3D reconstruction, we collected a novel human-labeled dataset of artifact and distortion maps in unseen reconstructed views. Through this dataset, we demonstrate that our method achieves state-of-the-art localization of artifacts in novel views, correlating with human assessment, even without aligned references. We can leverage our new metric to enhance applications like automatic image restoration, guided acquisition, or 3D reconstruction from sparse inputs. Find the project page at https://nihermann.github.io/puzzlesim/ .
Figures
Figures from the paper (2 more)
Reference graph
Works this paper leans on
-
[1]
Mantiuk, Karol Myszkowski, Hans-Peter Seidel, and Piotr Didyk
Vamsi Kiran Adhikarla, Marek Vinkler, Denis Sumin, Rafal K. Mantiuk, Karol Myszkowski, Hans-Peter Seidel, and Piotr Didyk. Towards a Quality Metric for Dense Light Fields. In2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 3720–3729, Honolulu, HI, 2017. IEEE. 5
work page 2017
- [2]
-
[3]
Barron, Ben Mildenhall, Matthew Tancik, Peter Hedman, Ricardo Martin-Brualla, and Pratul P
Jonathan T. Barron, Ben Mildenhall, Matthew Tancik, Peter Hedman, Ricardo Martin-Brualla, and Pratul P. Srinivasan. Mip-NeRF: A Multiscale Representation for Anti-Aliasing Neural Radiance Fields, 2021. 1, 5
work page 2021
-
[4]
Fast Feedforward 3D Gaussian Splatting Compression, 2024
Yihang Chen, Qianyi Wu, Mengyao Li, Weiyao Lin, Mehrtash Harandi, and Jianfei Cai. Fast Feedforward 3D Gaussian Splatting Compression, 2024. arXiv:2410.08017. 1
arXiv 2024
-
[5]
Depth-Regularized Optimization for 3D Gaussian Splatting in Few-Shot Images, 2024
Jaeyoung Chung, Jeongtaek Oh, and Kyoung Mu Lee. Depth-Regularized Optimization for 3D Gaussian Splatting in Few-Shot Images, 2024. arXiv:2311.13398 [cs]. 1
arXiv 2024
-
[6]
Scott J. Daly. Visible differences predictor: an algorithm for the assessment of image fidelity. InHuman Vision, Visual Processing, and Digital Display III, pages 2–15. SPIE, 1992. 2
work page 1992
-
[7]
Image inpainting through neural networks hallucinations
Alhussein Fawzi, Horst Samulowitz, Deepak Turaga, and Pascal Frossard. Image inpainting through neural networks hallucinations. In2016 IEEE 12th Image, Video, and Mul- tidimensional Signal Processing Workshop (IVMSP), pages 1–5, Bordeaux, France, 2016. IEEE. 2
work page 2016
-
[8]
Bayes’ Rays: Uncertainty Quantifica- tion for Neural Radiance Fields, 2023
Lily Goli, Cody Reading, Silvia Sell ´an, Alec Jacobson, and Andrea Tagliasacchi. Bayes’ Rays: Uncertainty Quantifica- tion for Neural Radiance Fields, 2023. arXiv:2309.03185 [cs]. 1
arXiv 2023
Show all 65 references
-
[9]
MCNeRF: Monte Carlo Rendering and Denoising for Real- Time NeRFs
Kunal Gupta, Milos Hasan, Zexiang Xu, Fujun Luan, Kalyan Sunkavalli, Xin Sun, Manmohan Chandraker, and Sai Bi. MCNeRF: Monte Carlo Rendering and Denoising for Real- Time NeRFs. InSIGGRAPH Asia 2023 Conference Papers, pages 1–11, Sydney NSW Australia, 2023. ACM. 2
2023
-
[10]
Deep blending for free-viewpoint image-based rendering.ACM Transactions on Graphics, 37(6):1–15, 2018
Peter Hedman, Julien Philip, True Price, Jan-Michael Frahm, George Drettakis, and Gabriel Brostow. Deep blending for free-viewpoint image-based rendering.ACM Transactions on Graphics, 37(6):1–15, 2018. 5
2018
-
[11]
Iandola, Song Han, Matthew W
Forrest N. Iandola, Song Han, Matthew W. Moskewicz, Khalid Ashraf, William J. Dally, and Kurt Keutzer. SqueezeNet: AlexNet-level accuracy with 50x fewer param- eters and<0.5MB model size, 2016. arXiv:1602.07360 [cs]. 3, 4
2016 arXiv
-
[12]
Billion- Scale Similarity Search with GPUs.IEEE Transactions on Big Data, 7(3):535–547, 2021
Jeff Johnson, Matthijs Douze, and Herv ´e J ´egou. Billion- Scale Similarity Search with GPUs.IEEE Transactions on Big Data, 7(3):535–547, 2021. Conference Name: IEEE Transactions on Big Data. 8
2021
-
[13]
Convo- lutional Neural Networks for No-Reference Image Quality Assessment
Le Kang, Peng Ye, Yi Li, and David Doermann. Convo- lutional Neural Networks for No-Reference Image Quality Assessment. In2014 IEEE Conference on Computer Vision and Pattern Recognition, pages 1733–1740, Columbus, OH, USA, 2014. IEEE. 2, 3, 6, 7
2014
-
[14]
MUSIQ: Multi-scale Image Quality Trans- former
Junjie Ke, Qifei Wang, Yilin Wang, Peyman Milanfar, and Feng Yang. MUSIQ: Multi-scale Image Quality Trans- former. In2021 IEEE/CVF International Conference on Computer Vision (ICCV), pages 5128–5137, Montreal, QC, Canada, 2021. IEEE. 2
2021
-
[15]
3D Gaussian Splatting for Real-Time Radiance Field Rendering.ACM Transactions on Graphics, 42(4):1–14, 2023
Bernhard Kerbl, Georgios Kopanas, Thomas Leimkuehler, and George Drettakis. 3D Gaussian Splatting for Real-Time Radiance Field Rendering.ACM Transactions on Graphics, 42(4):1–14, 2023. 1, 2, 5
2023
-
[16]
Tanks and temples: benchmarking large-scale scene reconstruction.ACM Transactions on Graphics, 36(4):1–13,
Arno Knapitsch, Jaesik Park, Qian-Yi Zhou, and Vladlen Koltun. Tanks and temples: benchmarking large-scale scene reconstruction.ACM Transactions on Graphics, 36(4):1–13,
-
[17]
Improving NeRF Quality by Progressive Camera Placement for Un- restricted Navigation in Complex Environments, 2023
Georgios Kopanas and George Drettakis. Improving NeRF Quality by Progressive Camera Placement for Un- restricted Navigation in Complex Environments, 2023. arXiv:2309.00014 [cs, eess]. 1, 2
2023 arXiv
-
[18]
Im- ageNet Classification with Deep Convolutional Neural Net- works
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. Im- ageNet Classification with Deep Convolutional Neural Net- works. InAdvances in Neural Information Processing Sys- tems. Curran Associates, Inc., 2012. 3, 4
2012
-
[19]
Deep Shape-Texture Statistics for Com- pletely Blind Image Quality Evaluation.ACM Transactions on Multimedia Computing, Communications, and Applica- tions, page 3694977, 2024
Yixuan Li, Peilin Chen, Hanwei Zhu, Keyan Ding, Leida Li, and Shiqi Wang. Deep Shape-Texture Statistics for Com- pletely Blind Image Quality Evaluation.ACM Transactions on Multimedia Computing, Communications, and Applica- tions, page 3694977, 2024. 5
2024
-
[20]
Mantiuk, K
R. Mantiuk, K. Myszkowski, and H.-P. Seidel. Visible dif- ference predicator for high dynamic range images. In2004 IEEE International Conference on Systems, Man and Cyber- netics (IEEE Cat. No.04CH37583), pages 2763–2769 vol.3,
-
[21]
Mantiuk, Gyorgy Denes, Alexandre Chapiro, An- ton Kaplanyan, Gizem Rufo, Romain Bachy, Trisha Lian, and Anjul Patney
Rafał K. Mantiuk, Gyorgy Denes, Alexandre Chapiro, An- ton Kaplanyan, Gizem Rufo, Romain Bachy, Trisha Lian, and Anjul Patney. FovVideoVDP: a visible difference pre- dictor for wide field-of-view video.ACM Transactions on Graphics, 40(4):1–19, 2021. 2
2021
-
[22]
Mantiuk, Dounia Hammou, and Param Hanji
Rafal K. Mantiuk, Dounia Hammou, and Param Hanji. HDR-VDP-3: A multi-metric for predicting image differ- ences, quality and contrast distortions in high dynamic range and regular content, 2023. arXiv:2304.13625. 2
2023 arXiv
-
[23]
Mantiuk, Param Hanji, Maliha Ashraf, Yuta Asano, and Alexandre Chapiro
Rafal K. Mantiuk, Param Hanji, Maliha Ashraf, Yuta Asano, and Alexandre Chapiro. ColorVideoVDP: A visual differ- ence predictor for image, video and display distortions, 2024. arXiv:2401.11485. 2 9
2024 arXiv
-
[24]
Srinivasan, Matthew Tancik, Jonathan T
Ben Mildenhall, Pratul P. Srinivasan, Matthew Tancik, Jonathan T. Barron, Ravi Ramamoorthi, and Ren Ng. NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis, 2020. arXiv:2003.08934 [cs]. 1, 2
2020 arXiv
-
[25]
No-Reference Image Quality Assessment in the Spa- tial Domain.IEEE Transactions on Image Processing, 21 (12):4695–4708, 2012
Anish Mittal, Anush Krishna Moorthy, and Alan Conrad Bovik. No-Reference Image Quality Assessment in the Spa- tial Domain.IEEE Transactions on Image Processing, 21 (12):4695–4708, 2012. Conference Name: IEEE Transac- tions on Image Processing. 1, 2
2012
-
[26]
Completely Blind
Anish Mittal, Rajiv Soundararajan, and Alan C. Bovik. Mak- ing a “Completely Blind” Image Quality Analyzer.IEEE Signal Processing Letters, 20(3):209–212, 2013. Conference Name: IEEE Signal Processing Letters. 1
2013
-
[27]
Blind Im- age Quality Assessment: From Natural Scene Statistics to Perceptual Quality.IEEE Transactions on Image Processing, 20(12):3350–3364, 2011
Anush Krishna Moorthy and Alan Conrad Bovik. Blind Im- age Quality Assessment: From Natural Scene Statistics to Perceptual Quality.IEEE Transactions on Image Processing, 20(12):3350–3364, 2011. Conference Name: IEEE Transac- tions on Image Processing. 2
2011
-
[28]
Instant Neural Graphics Primitives with a Mul- tiresolution Hash Encoding.ACM Transactions on Graphics, 41(4):1–15, 2022
Thomas M ¨uller, Alex Evans, Christoph Schied, and Alexan- der Keller. Instant Neural Graphics Primitives with a Mul- tiresolution Hash Encoding.ACM Transactions on Graphics, 41(4):1–15, 2022. arXiv:2201.05989 [cs]. 1, 2
2022 arXiv
-
[29]
Channappayya, and Swarup S
Venkatanath N, Praneeth D, Maruthi Chandrasekhar Bh, Sumohana S. Channappayya, and Swarup S. Medasani. Blind image quality evaluation using perception based fea- tures. In2015 Twenty First National Conference on Commu- nications (NCC), pages 1–6, 2015. 3, 6, 7
2015
-
[30]
Mueller, Chakravarty R
Thomas Neff, Pascal Stadlbauer, Mathias Parger, Andreas Kurz, Joerg H. Mueller, Chakravarty R. Alla Chaitanya, An- ton Kaplanyan, and Markus Steinberger. DONeRF: Towards Real-Time Rendering of Compact Neural Radiance Fields using Depth Oracle Networks, 2021. arXiv:2103.03231. 2
2021 arXiv
-
[31]
Understanding SSIM, 2020
Jim Nilsson and Tomas Akenine-M ¨oller. Understanding SSIM, 2020. arXiv:2006.13846 [eess]. 3, 7
2020 arXiv
-
[32]
Limitations of the SSIM quality metric in the context of diagnostic imaging
Jean-Franc ¸ois Pambrun and Rita Noumeir. Limitations of the SSIM quality metric in the context of diagnostic imaging. In 2015 IEEE International Conference on Image Processing (ICIP), pages 2960–2963, 2015. 3, 7
2015
-
[33]
Sch ¨onberger and Jan-Michael Frahm
Johannes L. Sch ¨onberger and Jan-Michael Frahm. Structure- from-Motion Revisited. In2016 IEEE Conference on Com- puter Vision and Pattern Recognition (CVPR), pages 4104– 4113, 2016. ISSN: 1063-6919. 1
2016
-
[34]
Sheikh, M.F
H.R. Sheikh, M.F. Sabir, and A.C. Bovik. A Statistical Eval- uation of Recent Full Reference Image Quality Assessment Algorithms.IEEE Transactions on Image Processing, 15 (11):3440–3451, 2006. Conference Name: IEEE Transac- tions on Image Processing. 5
2006
-
[35]
Very Deep Convo- lutional Networks for Large-Scale Image Recognition, 2015
Karen Simonyan and Andrew Zisserman. Very Deep Convo- lutional Networks for Large-Scale Image Recognition, 2015. arXiv:1409.1556 [cs]. 3, 4
2015 arXiv
-
[36]
Deep- V oxels: Learning Persistent 3D Feature Embeddings, 2019
Vincent Sitzmann, Justus Thies, Felix Heide, Matthias Nießner, Gordon Wetzstein, and Michael Zollh ¨ofer. Deep- V oxels: Learning Persistent 3D Feature Embeddings, 2019. arXiv:1812.01024 [cs]. 2
2019 arXiv
-
[37]
Resolution-robust Large Mask Inpainting with Fourier Convolutions
Roman Suvorov, Elizaveta Logacheva, Anton Mashikhin, Anastasia Remizova, Arsenii Ashukha, Aleksei Silvestrov, Naejin Kong, Harshith Goka, Kiwoong Park, and Victor Lempitsky. Resolution-robust Large Mask Inpainting with Fourier Convolutions. In2022 IEEE/CVF Winter Conference on...
2022
-
[38]
Why Are Deep Representations Good Perceptual Quality Features? InComputer Vision – ECCV 2020, pages 445–461
Taimoor Tariq, Okan Tarhan Tursun, Munchurl Kim, and Pi- otr Didyk. Why Are Deep Representations Good Perceptual Quality Features? InComputer Vision – ECCV 2020, pages 445–461. Springer International Publishing, Cham, 2020. Series Title: Lecture Notes in Computer Science. 3
2020
-
[39]
Perceptu- ally Adaptive Real-Time Tone Mapping
Taimoor Tariq, Nathan Matsuda, Eric Penner, Jerry Jia, Dou- glas Lanman, Ajit Ninan, and Alexandre Chapiro. Perceptu- ally Adaptive Real-Time Tone Mapping. InSIGGRAPH Asia 2023 Conference Papers, pages 1–10, Sydney NSW Aus- tralia, 2023. ACM. 2
2023
-
[40]
Perceptual Visibility Model for Temporal Contrast Changes in Periphery.ACM Trans
Cara Tursun and Piotr Didyk. Perceptual Visibility Model for Temporal Contrast Changes in Periphery.ACM Trans. Graph., 42(2):20:1–20:16, 2022. 2
2022
-
[41]
Luminance-contrast-aware foveated rendering.ACM Transactions on Graphics, 38(4): 1–14, 2019
Okan Tarhan Tursun, Elena Arabadzhiyska-Koleva, Marek Wernikowski, Radosław Mantiuk, Hans-Peter Seidel, Karol Myszkowski, and Piotr Didyk. Luminance-contrast-aware foveated rendering.ACM Transactions on Graphics, 38(4): 1–14, 2019. 2
2019
-
[42]
Tuy and Lee Tan Tuy
Heang K. Tuy and Lee Tan Tuy. Direct 2-D display of 3-D objects.IEEE Computer Graphics and Applications, 4(10): 29–34, 1984. Conference Name: IEEE Computer Graphics and Applications. 2
1984
-
[43]
Gomez, Lukasz Kaiser, and Illia Polosukhin
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszko- reit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention Is All You Need, 2017. arXiv:1706.03762 [cs]. 3
2017 arXiv
-
[44]
Chan, and Chen Change Loy
Jianyi Wang, Kelvin C.K. Chan, and Chen Change Loy. Ex- ploring CLIP for assessing the look and feel of images. In Proceedings of the Thirty-Seventh AAAI Conference on Arti- ficial Intelligence and Thirty-Fifth Conference on Innovative Applications of Artificial Intelligence a...
2023
-
[45]
Wang, E.P
Z. Wang, E.P. Simoncelli, and A.C. Bovik. Multiscale struc- tural similarity for image quality assessment. InThe Thrity- Seventh Asilomar Conference on Signals, Systems & Com- puters, 2003, pages 1398–1402 V ol.2, 2003. 2
2003
-
[46]
Bovik, H.R
Zhou Wang, A.C. Bovik, H.R. Sheikh, and E.P. Simoncelli. Image quality assessment: from error visibility to structural similarity.IEEE Transactions on Image Processing, 13(4): 600–612, 2004. Conference Name: IEEE Transactions on Image Processing. 2
2004
-
[47]
CrossScore: Towards Multi-View Image Evaluation and Scoring
Zirui Wang, Wenjing Bian, and Victor Adrian Prisacariu. CrossScore: Towards Multi-View Image Evaluation and Scoring. InComputer Vision – ECCV 2024, pages 492–510, Cham, 2025. Springer Nature Switzerland. 1, 3, 6, 7
2024
-
[48]
Nerfbusters: Re- moving Ghostly Artifacts from Casually Captured NeRFs,
Frederik Warburg, Ethan Weber, Matthew Tancik, Alek- sander Holynski, and Angjoo Kanazawa. Nerfbusters: Re- moving Ghostly Artifacts from Casually Captured NeRFs,
-
[49]
Krzysztof Wolski, Daniele Giunchi, Nanyang Ye, Piotr Didyk, Karol Myszkowski, Radosław Mantiuk, Hans-Peter 10 Seidel, Anthony Steed, and Rafał K. Mantiuk. Dataset and Metrics for Predicting Local Visible Differences.ACM Trans. Graph., 37(5):172:1–172:14, 2018. 5
2018
-
[50]
WaveNeRF: Wavelet-based Generalizable Neural Radiance Fields, 2023
Muyu Xu, Fangneng Zhan, Jiahui Zhang, Yingchen Yu, Xi- aoqin Zhang, Christian Theobalt, Ling Shao, and Shijian Lu. WaveNeRF: Wavelet-based Generalizable Neural Radiance Fields, 2023. arXiv:2308.04826. 2
2023 arXiv
-
[51]
Learning with- out Human Scores for Blind Image Quality Assessment
Wufeng Xue, Lei Zhang, and Xuanqin Mou. Learning with- out Human Scores for Blind Image Quality Assessment. In2013 IEEE Conference on Computer Vision and Pattern Recognition, pages 995–1002, Portland, OR, USA, 2013. IEEE. 2
2013
-
[52]
Beyond Human Opinion Scores: Blind Image Quality Assessment Based on Synthetic Scores
Peng Ye, Jayant Kumar, and David Doermann. Beyond Human Opinion Scores: Blind Image Quality Assessment Based on Synthetic Scores. In2014 IEEE Conference on Computer Vision and Pattern Recognition, pages 4241– 4248, 2014. ISSN: 1063-6919. 2
2014
-
[53]
From Patches to Pic- tures (PaQ-2-PiQ): Mapping the Perceptual Space of Picture Quality
Zhenqiang Ying, Haoran Niu, Praful Gupta, Dhruv Mahajan, Deepti Ghadiyaram, and Alan Bovik. From Patches to Pic- tures (PaQ-2-PiQ): Mapping the Perceptual Space of Picture Quality. In2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 3572–3582, S...
2020
-
[54]
Transformer for Image Quality Assessment, 2021
Junyong You and Jari Korhonen. Transformer for Image Quality Assessment, 2021. arXiv:2101.01097 [cs]. 2
2021 arXiv
-
[55]
Plenox- els: Radiance Fields without Neural Networks, 2021
Alex Yu, Sara Fridovich-Keil, Matthew Tancik, Qinhong Chen, Benjamin Recht, and Angjoo Kanazawa. Plenox- els: Radiance Fields without Neural Networks, 2021. arXiv:2112.05131 [cs]. 2
2021 arXiv
-
[56]
PlenOctrees for Real-time Rendering of Neural Radiance Fields, 2021
Alex Yu, Ruilong Li, Matthew Tancik, Hao Li, Ren Ng, and Angjoo Kanazawa. PlenOctrees for Real-time Rendering of Neural Radiance Fields, 2021. arXiv:2103.14024 [cs]. 2
2021 arXiv
-
[57]
pixelNeRF: Neural Radiance Fields from One or Few Im- ages, 2021
Alex Yu, Vickie Ye, Matthew Tancik, and Angjoo Kanazawa. pixelNeRF: Neural Radiance Fields from One or Few Im- ages, 2021. arXiv:2012.02190 [cs]. 1
2021 arXiv
-
[58]
Learning to Com- pare Image Patches via Convolutional Neural Networks
Sergey Zagoruyko and Nikos Komodakis. Learning to Com- pare Image Patches via Convolutional Neural Networks. pages 4353–4361, 2015. 6
2015
-
[59]
FSIM: A Feature Similarity Index for Image Quality Assess- ment.IEEE Transactions on Image Processing, 20(8):2378– 2386, 2011
Lin Zhang, Lei Zhang, Xuanqin Mou, and David Zhang. FSIM: A Feature Similarity Index for Image Quality Assess- ment.IEEE Transactions on Image Processing, 20(8):2378– 2386, 2011. Conference Name: IEEE Transactions on Image Processing. 2
2011
-
[60]
Lin Zhang, Lei Zhang, and Alan C. Bovik. A Feature- Enriched Completely Blind Image Quality Evaluator.IEEE Transactions on Image Processing, 24(8):2579–2591, 2015. Conference Name: IEEE Transactions on Image Processing. 5, 8
2015
-
[61]
Perceptual Artifacts Local- ization for Image Synthesis Tasks
Lingzhi Zhang, Zhengjie Xu, Connelly Barnes, Yuqian Zhou, Qing Liu, He Zhang, Sohrab Amirghodsi, Zhe Lin, Eli Shechtman, and Jianbo Shi. Perceptual Artifacts Local- ization for Image Synthesis Tasks. In2023 IEEE/CVF In- ternational Conference on Computer Vision (ICCV), pages 7...
2023
-
[62]
Efros, Eli Shecht- man, and Oliver Wang
Richard Zhang, Phillip Isola, Alexei A. Efros, Eli Shecht- man, and Oliver Wang. The Unreasonable Effectiveness of Deep Features as a Perceptual Metric. pages 586–595, 2018. 2, 3, 7
2018
-
[63]
Blind Image Quality Assessment via Vision- Language Correspondence: A Multitask Learning Perspec- tive
Weixia Zhang, Guangtao Zhai, Ying Wei, Xiaokang Yang, and Kede Ma. Blind Image Quality Assessment via Vision- Language Correspondence: A Multitask Learning Perspec- tive. In2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 14071–14081, Van- couv...
2023
-
[2004]
ISSN: 1062-922X. 1, 2
-
[2023]
arXiv:2304.10532 [cs]. 1
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.