{"id":"2a7a6929-4bbf-4673-8349-6f76620c85df","arxiv_id":"2412.08029","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"NeRF-NQA reports a no-reference quality score for dense-viewpoint NeRF scenes, outperforming PSNR, SSIM, LPIPS, VMAF, and light-field metrics on three labeled datasets.","lead":"NeRF-NQA is a no-reference quality metric for scenes generated by NeRF and other neural view synthesis methods, combining per-view spatial features with per-point angular features. It reports large gains over 23 existing image, video, and light-field quality metrics on three datasets with human perceptual labels.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Evaluation relies on a single subjective dataset with no significance testing; reported gains may reflect noise or label bias, not perceptual superiority.","rationale":"The reader's weakest assumption correctly identifies the JOD subjective scores as load-bearing. My stress-test agrees and sharpens the point: the problem is not only label noise but the complete absence of statistical inference on the reported differences. The abstract's 'significantly' is used without any significance test, and the small number of independent evaluation units (4 test scenes per dataset) makes the large percentage improvements in Tables 2-6 potentially compatible with sampling variation. The paper provides no confidence intervals, no bootstrap, and no permutation test, so the empirical superiority claim is not yet quantitatively established. The Limitations section is honest about missing 360-degree and 3DGS coverage, which narrows the scope but does not fix the statistical gap. I also note an internal tension in Table 1: the Pointwise module makes RMSE worse on Fieldwork (0.9202 without it vs 1.1969 with it), so the architecture's claimed consistent benefit is not uniformly supported. Despite these concerns, the method is plausible and the reported correlations are large, so the appropriate verdict remains conditional rather than reject; the proposed bootstrap and independent-subject check would determine whether the concern actually lands.","tokens_in":21521,"tokens_out":5095,"duration_ms":57545,"concrete_test":"Compute paired bootstrap or permutation confidence intervals for the difference in RMSE and SRCC between NeRF-NQA and the second-best method on each dataset, resampling at the level of scenes or NVS-method groups (not individual score samples). If any 95% confidence interval for the difference includes zero for Fieldwork, LLFF, or Lab, the claimed significant superiority is not established and the conclusion should be weakened. As a further validation, test NeRF-NQA on a new subjective study with NVS methods not used in training (e.g., 3D Gaussian Splatting, Instant-NGP) and fresh observers to verify that the learned mapping generalizes beyond the specific JOD campaign.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim of significant superiority is entirely empirical and rests on the JOD labels of Liang et al. [23] being valid, transferable ground truth. NeRF-NQA is trained to predict those labels and then evaluated on held-out scenes from the same three datasets and the same subjective campaign (39 volunteers, ASAP sampling, Thurstone V scaling). No independent subjective benchmark, no confidence intervals, and no significance test are reported. With only 4 test scenes per dataset in Table 2, and 12 scene-level or 10 NVS-method-level comparisons in Tables 3-4, the large RMSE/SRCC gaps could be driven by a few samples or by dataset-specific biases in the observer panel. The word 'significantly' in the abstract is therefore not supported by any statistical procedure; if the JOD labels are noisy, biased, or non-representative, every table ranking collapses. The paper's own Limitation section narrows scope by excluding 360-degree scenes and recent methods such as 3DGS, but it does not address the label-validity problem. An additional internal inconsistency is visible in Table 1: on Fieldwork, the final model with the Pointwise module has RMSE 1.1969, which is worse than the variant without it (0.9202), so the central design choice is not even consistently beneficial on its own reported numbers.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes NeRF-NQA, a no-reference quality assessment method for densely observed scenes synthesized by NeRF and other neural view synthesis methods. It combines a viewwise module that extracts per-view spatial quality features and inter-view consistency along the camera path with a pointwise module that computes normalized spherical gradients over COLMAP surface points and aggregates them with PointNet; the two feature streams are fused by an MLP to predict a JOD score. The method is trained and tested on the Lab, LLFF, and Fieldwork datasets using subjective JOD labels from Liang et al. [23], and it is compared with 23 image, video, and light-field quality assessment methods under RMSE, SRCC, PLCC, and outlier-ratio metrics, including held-out-scene and cross-dataset protocols. The central claim is that NeRF-NQA significantly outperforms all existing methods and is the first no-reference method for this setting.","tokens_in":21729,"tokens_out":4704,"duration_ms":50536,"significance":"The task is timely and practically relevant: NVS systems are increasingly used to generate dense viewpoint content, and no-reference quality assessment for such content is genuinely underdeveloped. The paper's pointwise angular-quality idea, embodied in the PNSG feature, is a sensible and nontrivial contribution, and the comparison against 23 established methods is broad. The release of an implementation is also a strength. If the reported gains were accompanied by proper statistical support, this would be a useful paper for the NVS and quality-assessment communities. However, as written, the central claim of 'significant' superiority rests entirely on point estimates over very small test sets, with no confidence intervals, significance tests, or repeated-split variance, and one of the three datasets shows a large RMSE regression when the pointwise module is added.","major_comments":[{"comment":"The abstract and Section 4.6 claim that NeRF-NQA 'outperforms the existing assessment methods significantly' and shows 'substantial superiority,' but these claims are supported only by point estimates computed over four test scenes per dataset, with tenfold surface sampling that does not create independent scenes. No confidence intervals, bootstrap estimates, per-split variance, or significance tests are reported anywhere in the manuscript. Because the headline gains in Table 2 could change with a different random scene split, the authors should provide repeated-split results, bootstrap confidence intervals over scenes, or a paired significance test across scenes/methods, and they should temper the word 'significantly' unless such a procedure supports it.","section":"§4.6, Table 2; Abstract"},{"comment":"The ablation study in Table 1 does not support the claim that the pointwise module is consistently beneficial. On the Fieldwork dataset, adding the pointwise module increases RMSE from 0.9202 to 1.1969, a roughly 30% regression, even though SRCC improves from 0.9343 to 0.9701. The text states that 'with the exception of RMSE on the Fieldwork dataset, where the results are closely aligned,' but a 0.28 difference on this scale is not close, and the full model's Fieldwork RMSE of 1.1969 is the very number used as a headline result in Table 2. The authors need to explain this trade-off, report which metric is primary, or show that the RMSE regression is not systematic before claiming that the pointwise design is validated.","section":"Table 1, §4.5"},{"comment":"The ground-truth labels are JOD scores from a single subjective campaign with 39 volunteers, and NeRF-NQA is trained and tested on scenes from that same campaign. While the held-out-scene split prevents direct circularity, the paper provides no analysis of label reliability: there are no bootstrap confidence intervals for the JOD values, no observer-variability metrics, and no independent perceptual benchmark. Since every ranking in Tables 2-6 is measured against these labels, the authors should report the uncertainty in the ground truth (e.g., by bootstrapping the pairwise-comparison data) or evaluate on an external subjective dataset. The Limitation section should also acknowledge this dependence rather than only listing 360-degree scenes and 3DGS as future work.","section":"§4.2, §5, Table 2"},{"comment":"The caption of Table 6 says that 'each method is trained on two datasets and tested on the third,' but this cannot be true for the full-reference methods, which are not trainable, and it is unclear whether the no-reference baselines were retrained or used with default weights. This ambiguity matters for the fairness of the cross-dataset comparison, especially because FR-IQA methods that use references from the test set are compared with a method that sees no references. Please clarify the exact training protocol for each baseline and, if no-reference baselines were not retrained, state that explicitly.","section":"§4.9, Table 6"}],"minor_comments":[{"comment":"The sentence 'using a image set' should read 'using an image set.'","section":"Abstract"},{"comment":"The phrase 'The implementation replied on the PyTorch' should read 'The implementation relied on PyTorch.'","section":"§4.3"},{"comment":"The PNSG description mixes terminology: it says the polar axis is partitioned into bins but then refers to a 'specific azimuthal bin,' and the notation NSGazi/NSGpol could be defined more clearly. Please align the axis names and bin indices.","section":"§3.3"},{"comment":"The term 'Just-Objectionable-Difference' in Section 4.2 and 'Just-Noticeable Differences' in the Figure 7 caption should be harmonized, preferably with the terminology used in the original subjective study [23].","section":"§4.2, Figure 7"},{"comment":"Several table cells in the supplied text show repeated digit strings such as '0.92020.92020.9202' and '1.19691.19691.1969.' Please ensure the final PDF renders single values and that the source tables do not contain duplicated numeric tokens.","section":"Tables 1-4"},{"comment":"The scatter plots have no labeled axes and no legend for the scene markers; adding axis labels, units (JOD vs. predicted score), and a legend would make the figure interpretable.","section":"Figure 6"},{"comment":"The paper says four scenes are randomly designated for testing in each dataset, but no random seed or explicit split is provided. Releasing the exact scene split (or seeds) would make the results reproducible.","section":"§4.1"},{"comment":"Reference [23] is cited as an arXiv preprint even though the paper's header shows a TVCG publication with a DOI; please cite the published version.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The core idea is sound and the reported gains are large, but the statistical support for the central claim is missing and the Fieldwork ablation regression is more serious than the text admits. I would like to see the authors either add proper uncertainty quantification and a significance analysis or substantially weaken the 'significantly' claim in the abstract and conclusions. The paper appears to be an already-published TVCG article posted to arXiv; if this is being considered as a new submission, the reviewer should be aware that the statistical concerns apply to the published version as well."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper does something new: it targets no-reference quality assessment specifically for dense-viewpoint NVS scenes, and the PNSG pointwise angular feature is a real design contribution. The joint viewwise/pointwise architecture is sensible, the cross-dataset evaluation is a plus, and the code release helps reproducibility. Credit where due: the authors benchmark against 23 methods, include an ablation, and their limitations section honestly flags the COLMAP dependency and missing 360-degree/3DGS coverage. That is more than many papers do.\n\nThe soft spots are statistical, not architectural. The claim of \"significant superiority\" rests entirely on point estimates from a single subjective dataset—the JOD labels from Liang et al. (39 volunteers, ASAP sampling). There are no confidence intervals, no significance tests, and only four test scenes per dataset in the main table. If those labels are noisy or biased for certain scene types, the rankings could shift. That is a load-bearing assumption, and the paper does not address it. The ablation also has an internal inconsistency: on Fieldwork, the version with the pointwise module has worse RMSE (1.1969) than without it (0.9202), though SRCC improves. The text calls this \"closely aligned,\" which is not accurate. Minor reporting gaps—angular bins b, number of sampled surface points, and other hyperparameters—make reproduction harder.\n\nNone of this kills the contribution. The metric is plausible, the gains are large and consistent across many comparisons, and the cross-dataset results suggest generalization beyond a single split. But \"promising\" and \"proven\" are different. The paper deserves a serious referee, and a revision that adds confidence intervals, per-split variance, or an independent subjective benchmark would substantially strengthen it. As submitted, treat the quantitative superiority as suggestive rather than conclusive.","headline":"Plausible first no-reference NVS quality metric with a genuinely new angular feature, but the 'significant superiority' claim needs statistical support before it can be taken at face value.","tokens_in":22317,"tokens_out":1216,"would_cite":true,"duration_ms":15099,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"NeRF-NQA is the first no-reference quality assessment method for dense-view scenes synthesized by NeRF and other neural view synthesis methods, reporting consistent performance gains over 23 existing image, video, and light-field metrics…","keywords":["no-reference quality assessment","neural radiance fields","neural view synthesis","pointwise angular quality","spherical gradient","view consistency","JOD subjective scores","dense viewpoint scenes"],"falsifier":"Run NeRF-NQA on a new set of dense-view scenes, including 360-degree scenes and outputs from methods not in the training set, with freshly collected human opinion scores from a different group of observers, and compare SRCC and PLCC; if NeRF-NQA no longer ranks first or its correlations drop markedly, the claimed superiority and no-reference generalization are not established.","tokens_in":21285,"feed_emoji":"🖼️","tokens_out":5750,"duration_ms":56335,"temperature":0.7,"pith_summary":"NeRF-NQA is proposed as the first no-reference quality assessment method for densely observed scenes produced by neural view synthesis (NVS) and NeRF-style renderers. Standard metrics such as PSNR, SSIM, and LPIPS compare each synthesized view against a ground-truth view, which is unavailable or too sparse for most NVS output, and they miss the consistency of a scene across many viewpoints. The paper argues that a scene should be judged both view by view and point by point: each view contributes spatial quality, while each reconstructed surface point contributes angular quality, observable as how much pixel values change with viewing angle. On the Fieldwork, LLFF, and Lab datasets, with human JOD scores as ground truth, the method reports the best RMSE, SRCC, PLCC, and outlier-ratio results among 24 assessed methods, including image, video, and light-field metrics.","feed_headline":"First no-reference quality score for NeRF scenes beats 23 metrics","feed_subtitle":"Viewwise and pointwise features match human judgments on Fieldwork, LLFF, and Lab without any reference view.","key_machinery":"The load-bearing object is the Pointwise Normalized Spherical Gradient (PNSG): for a surface point $p$ and two pixels $x_i, x_j$ observing it from different views, the normalized spherical gradient is $\\mathrm{NSG}(x_i,x_j)=(I(x_i)-I(x_j))/\\measuredangle x_i p x_j$, the RGB change per unit angular separation. Aggregating these gradients over azimuthal and polar bins for many surface points yields a feature that encodes how consistently the scene appears from different directions, which is exactly the angular quality that per-image metrics cannot see. The viewwise module supplies the complementary spatial signal: per-view quality features, produced by an EfficientNetV2-style backbone, are processed along the camera path with (Fused) MBConv layers and max pooling so that inter-view consistency enters the score. The final quality score is a learned MLP fusion of the two feature streams.","core_discovery":"On the paper's own terms, the central discovery is that no-reference quality assessment for NVS scenes becomes accurate when spatial and angular evidence are combined. The viewwise module scores each synthesized frame and then reads quality along the camera path, capturing whether adjacent views stay consistent. The pointwise module samples sparse surface points with COLMAP, collects the pixels from different views that observe each point, and computes Pointwise Normalized Spherical Gradients (PNSG), i.e., pixel-value differences divided by the angular separation between viewing directions; high gradients at a surface point indicate angular inconsistency and thus perceived distortion. An MLP fuses the two feature streams into a JOD-scaled quality score. The paper reports that this joint design cuts RMSE by 33.0%, 34.9%, and 20.0% against the second-best baseline on Fieldwork, LLFF, and Lab respectively, and that it also wins on most individual scenes and NVS methods and in cross-dataset tests.","pith_inferences":["Because NeRF-NQA reads only synthesized views and camera poses, the same machinery should extend to any dense-view renderer, including 3D Gaussian Splatting and other rasterizers, despite being trained only on NeRF-family outputs; this is an extension the paper names as future work rather than a demonstrated result.","The PNSG signal could be reused as a diagnostic for view consistency in multi-view reconstruction tasks beyond quality scoring, such as detecting floaters or flicker in free-viewpoint video, though the paper does not test this.","The reliance on COLMAP sparse points could be replaced by depth or ray-marching information from the renderer itself, which would help in textureless or specular scenes where structure-from-motion points are scarce; the paper only argues empirically that the current reliance is not fatal.","A larger subjective study with more scenes and observers, including 360-degree content, would be the natural stress test; if NeRF-NQA's margin shrinks on that data, the method's generality would need to be revised."],"forward_implications":["If the reported results hold, NVS quality can be assessed without any reference views, so datasets like LLFF that provide only sparse captures become fully evaluable.","Because both per-view and cross-view artifacts are penalized, NeRF-NQA should be a stronger predictor of human preference than single-image metrics, especially for blur and artifacts visible only across an image sequence.","The same trained model transfers across datasets, trained on two and tested on the third, without scene-specific fine-tuning, which the paper reports as consistent gains over baselines.","Scene- and method-level analyses show the largest wins on complex shapes and specular surfaces, where conventional metrics are weakest.","A practical consequence is that immersive VR/AR content rendered from neural view synthesis can be monitored for quality at runtime without storing reference imagery."],"supporting_citations":[{"why":"Supplies the JOD subjective quality labels from 39 volunteers that serve as ground truth for training and evaluation.","marker":"[23]"},{"why":"Provides COLMAP structure-from-motion points used by the pointwise module to sample surface points and poses.","marker":"[43]"},{"why":"Defines the LLFF dataset and its sparse-view protocol, the motivating case where full-reference assessment is impossible.","marker":"[26]"},{"why":"Introduces NeRF, the central rendering method whose outputs NeRF-NQA assesses.","marker":"[27]"},{"why":"Describes the ASAP active sampling used to collect the pairwise comparisons behind the JOD labels.","marker":"[25]"},{"why":"Provides the Thurstone Case V observer model that converts pairwise comparisons into JOD scores.","marker":"[35]"},{"why":"Supplies PointNet, used to extract inter-point features from the sampled surface points.","marker":"[38]"},{"why":"Supplies the EfficientNetV2 and MBConv design used in the viewwise quality module.","marker":"[50]"},{"why":"SSIM is one of the main full-reference baselines NeRF-NQA must beat.","marker":"[56]"},{"why":"LPIPS is the deep perceptual baseline most commonly used for NeRF evaluation and is a key comparison in the experiments.","marker":"[64]"}],"fun_headline_variants":["No-reference NeRF quality metric cuts error by 34% vs baselines","Viewwise+pointwise scoring nails NeRF scene quality without references","NeRF-NQA: first blind quality score for neural view synthesis","Beat 23 visual metrics on NeRF scenes with no reference views","Spatial plus angular cues make NeRF quality assessment reference-free"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The entire claim of superiority rests on the human quality labels from one subjective study being accurate and representative; if those labels are noisy or biased, the reported performance gains over other metrics would not generalise.","fun_headline_variants_meta":{"raw":{"variants":["No-reference NeRF quality metric cuts error by 34% vs baselines","Viewwise+pointwise scoring nails NeRF scene quality without references","NeRF-NQA: first blind quality score for neural view synthesis","Beat 23 visual metrics on NeRF scenes with no reference views","Spatial plus angular cues make NeRF quality assessment reference-free"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000744,"raw_usage":{"total_tokens":3368,"prompt_tokens":1048,"completion_tokens":2320,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":664,"completion_tokens_details":{"reasoning_tokens":2227}},"tokens_in":664,"tokens_out":2320,"duration_ms":16356,"temperature":1.0,"reasoning_tokens":2227,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T18:17:30.173448+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run NeRF-NQA on a new set of dense-view scenes, including 360-degree scenes and outputs from methods not in the training set, with freshly collected human opinion scores from a different group of observers, and compare SRCC and PLCC; if NeRF-NQA no longer ranks first or its correlations drop markedly, the claimed superiority and no-reference generalization are not established.","supporting_citations":[{"cited_title":"Perceptual Quality Assessment of NeRF and Neural View Synthesis Methods for Front-Facing Views","cited_arxiv_id":"2303.15206","evidence_quote":"Supplies the JOD subjective quality labels from 39 volunteers that serve as ground truth for training and evaluation."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides COLMAP structure-from-motion points used by the pointwise module to sample surface points and poses."},{"cited_title":"Mildenhall, P","cited_arxiv_id":null,"evidence_quote":"Defines the LLFF dataset and its sparse-view protocol, the motivating case where full-reference assessment is impossible."},{"cited_title":"Mildenhall, P","cited_arxiv_id":null,"evidence_quote":"Introduces NeRF, the central rendering method whose outputs NeRF-NQA assesses."},{"cited_title":"Mikhailiuk, C","cited_arxiv_id":null,"evidence_quote":"Describes the ASAP active sampling used to collect the pairwise comparisons behind the JOD labels."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies PointNet, used to extract inter-point features from the sampled surface points."},{"cited_title":"Tan and Q","cited_arxiv_id":null,"evidence_quote":"Supplies the EfficientNetV2 and MBConv design used in the viewwise quality module."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"SSIM is one of the main full-reference baselines NeRF-NQA must beat."},{"cited_title":"Zhang, P","cited_arxiv_id":null,"evidence_quote":"LPIPS is the deep perceptual baseline most commonly used for NeRF evaluation and is a key comparison in the experiments."}],"review_version":1}