{"id":"9cf4ec3c-e016-4fe5-bb0e-3786e9237895","arxiv_id":"2506.06710","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A new real-world degraded 360-degree image and video dataset, 360Insta, shows that current omnidirectional super-resolution models perform markedly worse on authentic degradations than on synthetic bicubic degradation.","lead":"This paper surveys deep learning methods for super-resolving 360-degree images and videos, and introduces a new real-world degraded dataset called 360Insta for evaluating them. The benchmark results show current methods lose substantial quality on authentic motion blur, low light, and exposure problems, which matters for VR and AR immersion.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"360Insta generalization claim rests on unvalidated NR-IQA metrics without ground truth or LR baseline; Table 8 cannot distinguish metric bias from model failure.","rationale":"Good-faith reading: the paper is a survey plus a 360Insta dataset and an empirical evaluation. Its main novel claim is that methods that work well on bicubic and fisheye synthetic degradations generalize poorly to authentic degraded omnidirectional content. For that claim to hold, the 360Insta evaluation must be able to measure quality of SR outputs on unpaired real images. The crucial condition is therefore that the NR-IQA metrics are valid for ERP omnidirectional images and that the comparison is fair (LR baseline and enough methods). Section 6.2.5 is the only place this is tested, and it is exactly where the assumption is least secure. The reader identified this as the weakest assumption, and I agree. Additional observations reinforce this: the 360Insta video link is a personal Google Drive placeholder ('drive/my-drive'), the method coverage in Table 8 is sparse, and the BPOSR numbers in the Section 6.2.1 text do not match Table 4 (the x16 text quotes x8 values). None of these are fraud concerns; they are evidence that the empirical section needs care. The survey taxonomy and standard benchmark tables are useful, and the dataset idea addresses a genuine gap if released and validated. My recommendation is UNCHANGED relative to the reader's CONDITIONAL verdict: the paper can be accepted conditionally, but the generalization claim should be stated with appropriate uncertainty until the NR-IQA metrics are calibrated for omnidirectional content or the evaluation is supplemented with a perceptual study and an LR baseline.","tokens_in":26033,"tokens_out":5664,"duration_ms":56106,"concrete_test":"Recompute the four NR-IQA metrics on the 100 ground-truth HR images of the ODISR-clean test set and on their bicubic LR (and x4 OSRT/BPOSR outputs). If the HR ground truths receive scores in the same low range as Table 8 (e.g., NIQE above 7 or MANIQA near 0.3), the metrics are not calibrated for ERP content and Table 8 cannot support the generalization-failure conclusion. If HR images score clearly better, then add an LR-input row to Table 8 for 360Insta: SR outputs must beat the LR inputs on the same metrics for the claim that methods fail on real degradations to be meaningful.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 6.2.5 and Table 8 conclude that current ODISR methods 'do not perform well on the real degradation dataset' using only four no-reference metrics (NIQE, MUSIQ, MANIQA, CLIPIQA). Since 360Insta has no paired high-resolution ground truth, the entire conclusion depends on these metrics being valid for ERP omnidirectional images. This assumption is unverified and questionable: the metrics were designed for perspective images with common distortions, while ERP content has latitude-dependent stretching and seams; the paper itself (Sec. 6.1.2) concedes there is no universally accepted metric. The evaluation protocol also splits 3840x1920 images into 16 blocks, which can perturb natural-scene statistics and degrade NR-IQA reliability. Table 8 contains no LR-input baseline and covers only two methods at x2 (360-SS, OSRT) and three at x4 (adding BPOSR), so 'uniformly low' is not robustly established. Low absolute NR-IQA values could reflect the inherently degraded content of 360Insta or metric bias rather than a failure of SR generalization. Without calibration of these metrics on omnidirectional content, the central empirical claim is unsupported.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper is a systematic survey of deep learning-based omnidirectional image and video super-resolution (ODISR/ODVSR). It taxonomizes existing methods by design principle (distortion maps, projection designs, position encoding, training strategies, diffusion), describes common datasets and metrics, and introduces a new dataset, 360Insta, of authentically degraded omnidirectional images and videos captured with a consumer 360-degree camera. The authors benchmark several published ODISR methods on established synthetic-degradation datasets (ODISR, SUN360 and clean variants) under both ERP and fisheye downsampling, compare model complexity, and evaluate a subset of methods on 360Insta using no-reference IQA metrics. The central empirical claim is that current ODISR methods, which perform well on synthetic bicubic degradation, achieve uniformly low no-reference quality scores on 360Insta, indicating poor generalization to real-world degradations.","tokens_in":26194,"tokens_out":3589,"duration_ms":34737,"significance":"If the 360Insta evaluation is trustworthy, the dataset and the generalization finding would be a useful contribution to the omnidirectional SR community, where most benchmarks rely on synthetic bicubic or fisheye downsampling. The paper also provides a valuable organized testbed: it compiles representative methods, reports results using official implementations, includes parameter/FLOP comparisons, and makes code and data links publicly available (with the caveat noted below). These are concrete strengths. However, the paper's main empirical novelty—the real-world generalization claim—is currently supported only by unvalidated no-reference metrics without a low-resolution baseline, which is a load-bearing weakness. The paper does not derive new theory; the value rests on the reliability of the new benchmark, so the evaluation protocol must be robust.","major_comments":[{"comment":"The central claim that current ODISR methods 'do not perform well on the real degradation dataset' rests entirely on four no-reference metrics (NIQE, MUSIQ, MANIQA, CLIPIQA) applied to 360Insta, which has no paired high-resolution ground truth. This is problematic for two reasons. First, the manuscript itself notes in Section 6.1.2 that 'there is still no universally accepted metric' for this field, and none of the four chosen metrics was developed or validated for equirectangular projection (ERP) content, which has latitude-dependent stretching, seams, and large unpopulated sky/ground regions. Second, Table 8 contains no low-resolution input baseline: it reports only SR outputs, so low NIQE or MANIQA values could reflect the inherently degraded content of 360Insta or systematic metric bias rather than a failure of SR generalization. The evaluation also splits 3840x1920 images into 16 blocks, which can alter natural-scene statistics on which NIQE relies. To make the conclusion load-bearing, the authors should report the same NR-IQA scores for the LR inputs and for a trivial baseline (e.g., bicubic upsampling), and should add at least a small paired real-data subset with full-reference metrics or a human rating study. Without such calibration, Section 6.2.5's conclusion is not supported.","section":"Section 6.2.1, text after Table 4"},{"comment":"The text states: 'At a scaling factor of 16, the BPOSR method performs well in PSNR and SSIM, reaching 24.01 and 0.6730, respectively.' In Table 4, these exact values (24.01 PSNR, 0.6730 SSIM on ODISR) are the BPOSR x8 results; the actual x16 values are 22.05 PSNR and 0.6258 SSIM. This misattribution affects the discussion of large-scale SR performance and should be corrected, along with the following sentence about WS-PSNR/WS-SSIM, which similarly appears to reference the x8 row.","section":"Section 6.2.5, Table 8"},{"comment":"The dataset contribution is incompletely documented. Section 5.3 gives only image counts for 360Insta and no count, resolution, or duration details for the claimed 360Insta video dataset, and Section 6.2.6 provides the video dataset link as 'https://drive.google.com/drive/my-drive', which is a private placeholder rather than an accessible public resource. Since the abstract and introduction promise that 'all datasets... are publicly available,' this unverifiable availability undermines a core contribution and needs to be fixed before the dataset claim can be accepted.","section":"Section 5.3 and Section 6.2.6"}],"minor_comments":[{"comment":"The dataset table lists '360HUD' for the video dataset while Section 5.2 and reference [39] use '360UHD'; this inconsistency should be reconciled.","section":"Section 5.1"},{"comment":"There is a typographical error: '2048°×1024' includes a degree symbol where none is intended; it should read '2048×1024'.","section":"Section 5.2"},{"comment":"The sentence 'The ODV-SR data set is then divided into training, validation, and testing sets, which include 270 clips, 20 clips, and 25 clips, respectively' is correct, but the adjacent description of ODV360 says 'all videos were downsampled to a resolution of 2K (2160×1080), with each clip consisting of 100 frames'; the numbers for ODV360 (210/20/20) do not match the earlier statement of '90 HR videos' plus '160 videos' collected with Insta360 cameras—please clarify the total counts.","section":"Section 6.2.4"},{"comment":"The table lists a method named 'Aalign' [68]; the standard name in the cited reference is 'A2N' (or 'Align' in some papers). Please use the canonical name to avoid confusion.","section":"Section 6.2.6"},{"comment":"In the visual-quality discussion, the text says 'OSRT [23] reveals pronounced artifacts that compromise visual coherence while it still performs well'; this sentence contradicts the stronger quantitative claims about OSRT's superiority in the same section. Please rephrase to state which artifacts are observed and how they interact with the quantitative results.","section":"Section 1"}],"recommendation":"major_revision","confidential_remarks":"The paper is a survey plus benchmark contribution rather than a new algorithm. The editor should consider whether the journal's scope welcomes such a contribution; the empirical claims will need to be re-supported with the baseline and calibration experiments described in the major comments before the work can be considered further."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line: this is a serviceable survey with a genuinely useful benchmark idea, but the paper's headline empirical claim about poor real-world generalization is not backed by the evidence as presented. The survey part is the safest contribution: it organizes 2018-2025 ODISR/ODVSR methods into sensible categories, unifies test protocols, and reruns public code under ERP and fisheye downsampling. Tables 4-7 and the visual comparisons look internally coherent and will save people time. The 360Insta dataset idea is right—existing paired datasets are bicubic or synthetic and do not reflect real capture conditions—and the authors are honest about the diversity of degradations. That is enough to make the paper worth engaging.\n\nThe soft spot is Section 6.2.5/Table 8. The conclusion that current methods 'do not perform well' on 360Insta rests entirely on NIQE, MUSIQ, MANIQA, and CLIPIQA with no ground truth, no LR-input baseline, and only two methods at x2 and three at x4. The paper itself says in Sec 6.1.2 there is no universally accepted metric, and these metrics were not designed for ERP with block splitting. Low absolute values might mean metric bias on panoramic content, not SR failure. The stress-test note is right: Table 8 cannot distinguish metric bias from model failure. Add an LR baseline, calibrate the metrics on known distortions, or collect paired HR captures.\n\nTwo smaller issues. Section 6.2.1 says BPOSR at x16 reaches PSNR 24.01 and SSIM 0.6730, but those are the x8 values in Table 4; the x16 values are 22.05 and 0.6258. And the video half of 360Insta is not in the paper; the provided link is a placeholder. I cannot verify the central dataset claim until the data is actually downloadable. Citation coverage looks solid for the subfield, and the projection math is routine but correct.\n\nThe paper should go to peer review, not desk rejection, but acceptance should be conditional: release the dataset and evaluation code, fix the x16 text, and strengthen the 360Insta evaluation protocol.","headline":"Useful survey and a benchmark idea worth taking seriously, but the real-world generalization claim is not yet supported by the NR-IQA-only evaluation.","tokens_in":26756,"tokens_out":2655,"would_cite":true,"duration_ms":29867,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper argues that deep learning-based omnidirectional image and video super-resolution methods, which perform well on synthetic bicubic-downsampling benchmarks, generalize poorly to authentically captured degradations, and it…","keywords":["omnidirectional image super-resolution","omnidirectional video super-resolution","360Insta dataset","real-world degradation","no-reference image quality assessment","virtual reality","deep learning","panoramic distortion"],"falsifier":"A subjective study would settle the claim: show viewers the original low-resolution frames, bicubic upscaling, and each model's scale-factor-2 and scale-factor-4 outputs from a subset of 360Insta on a head-mounted display, then correlate their quality ratings with NIQE, MUSIQ, MANIQA, and CLIPIQA. If human preference ranks the models differently, or rates the super-resolved frames as acceptable quality, the paper's conclusion that current methods do not perform well on real degradation would not stand.","tokens_in":25812,"feed_emoji":"🌐","tokens_out":13887,"duration_ms":134035,"temperature":0.7,"pith_summary":"This paper is a systematic survey of deep learning-based omnidirectional image and video super-resolution, and its central contribution is a real-world evaluation. The authors argue that standard benchmarks, which create low-resolution panoramas by bicubic or fisheye downsampling, do not represent what happens when 360-degree cameras record actual scenes. To close that gap, they introduce 360Insta, a dataset of 1,500 authentically degraded omnidirectional images plus video, spanning indoor, outdoor, day, night, motion blur, dim lighting, and varied exposure. Current models that look strong on synthetic benchmarks receive uniformly low scores on four no-reference quality metrics when run on 360Insta. The paper concludes that current methods do not generalize well to real-world degradations and that the field needs more realistic benchmarks.","feed_headline":"Real captures defeat 360-degree super-resolution models","feed_subtitle":"New 1,500-image benchmark of authentically degraded panoramas shows synthetic-trained models score low on real footage","key_machinery":"The load-bearing instrument is 360Insta, a benchmark of 1,500 authentically degraded 360-degree images in equirectangular projection, the common flat panoramic format, at 3840 by 1920 resolution, plus a video set, captured with a consumer 360-degree camera under real conditions and deliberately including multi-scene blur, different lighting, dim conditions, motion blur, and exposure variation. Because 360Insta has no high-resolution ground truth, evaluation runs through four no-reference image quality metrics—NIQE, MUSIQ, MANIQA, and CLIPIQA—computed on 16 blocks of each image after super-resolution at scale factors 2 and 4. The comparison uses the publicly released code of existing ODISR models such as 360-SS, OSRT, BPOSR, and OmniSSR. The argument's force comes from the contrast: the same methods score well under synthetic ERP and fisheye downsampling protocols in Tables 4 and 5 and then produce uniformly low no-reference scores on real captures in Table 8.","core_discovery":"The central claim is that deep learning-based omnidirectional super-resolution has been tuned to synthetic degradations and has not been shown to work on real-world panoramic captures. Trained and tested on datasets like ODISR, ODISR-clean, SUN360, and SUN360-clean, where low-resolution inputs are produced by bicubic or fisheye downsampling, the surveyed methods achieve competitive scores on full-reference metrics. On 360Insta, which contains images with motion blur, dim conditions, changing exposure, and varied scenes, the same methods score uniformly low on NIQE, MUSIQ, MANIQA, and CLIPIQA at both scale factors 2 and 4. The paper treats this as evidence that the field requires real-degradation benchmarks and models robust to them, rather than only higher numbers on synthetic pairs.","pith_inferences":["A natural experiment the paper does not run is to fine-tune or train an ODISR model on real-degradation examples like 360Insta and see whether the no-reference scores improve; if they do, the failure is a distribution-shift problem rather than a hard limit of super-resolution.","Because 360Insta has no paired high-resolution ground truth, it cannot distinguish a model that genuinely recovers detail from one that merely changes texture in ways the metrics reward; pairing the benchmark with matched high-quality captures of the same scenes would convert it from a diagnostic set into a trainable one.","The authors' conclusion depends on four no-reference metrics, and those metrics were not designed for equirectangular panoramas; a head-mounted-display subjective study comparing super-resolved outputs against bicubic upscaling would test whether viewers actually share the metrics' verdict.","The taxonomy in the survey suggests that real-degradation robustness might require combining distortion-aware losses, projection fusion, and temporal alignment, but the paper does not propose such a combination."],"forward_implications":["On 360Insta, every tested model receives low NIQE, MUSIQ, MANIQA, and CLIPIQA scores at scale factors 2 and 4, so claiming real-world readiness for VR and AR applications from synthetic-benchmark results is not supported by this evaluation.","Increasing the upscaling factor makes real-degradation quality worse for the tested models, and no method in the comparison escapes that trend.","The survey's unified test protocols and public dataset make it possible for future ODISR and ODVSR work to be measured for real-world robustness rather than only for synthetic fidelity.","A method that is strong on synthetic benchmarks is not automatically strong on authentic captures; the rankings change when the evaluation moves to 360Insta.","Because 360Insta includes videos, the same real-degradation gap can be checked for video super-resolution, although the paper notes its video set could not be included in the main text."],"supporting_citations":[{"why":"It supplies the OSRT distortion-aware transformer and the ODISR-clean and SUN360-clean fisheye-downsampling protocols that define the synthetic side of the comparison.","marker":"[23]"},{"why":"It supplies the LAU-Net latitude-adaptive method and the ODISR and SUN360 ERP bicubic-downsampling benchmarks used in the synthetic evaluations.","marker":"[29]"},{"why":"It supplies the 360-SS adversarial baseline, a compared method whose outputs receive some of the worst no-reference scores on 360Insta.","marker":"[38]"},{"why":"It supplies the BPOSR bi-projection fusion baseline, a strong synthetic-benchmark performer that still scores low on 360Insta.","marker":"[27]"},{"why":"It supplies the OmniSSR zero-shot diffusion baseline included in both the synthetic and real-degradation comparisons.","marker":"[33]"},{"why":"It represents the prior challenge datasets built on synthetic downsampling that the survey positions 360Insta against.","marker":"[36]"},{"why":"It is one of the four no-reference metrics used to judge the super-resolved outputs on 360Insta.","marker":"[75]"},{"why":"It is one of the four no-reference metrics used to judge the super-resolved outputs on 360Insta.","marker":"[76]"},{"why":"It is one of the four no-reference metrics used to judge the super-resolved outputs on 360Insta.","marker":"[77]"},{"why":"It is one of the four no-reference metrics used to judge the super-resolved outputs on 360Insta.","marker":"[78]"}],"fun_headline_variants":["Real-world panorama super-resolution models underperform","Synthetic-trained 360° SR fails on real captures","New benchmark exposes 360° super-resolution gap","Authentic degradation dataset stumps omnidirectional SR"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The real-world generalization conclusion rests on the assumption that the four no-reference quality metrics (NIQE, MUSIQ, MANIQA, and CLIPIQA) faithfully capture perceived quality for distorted 360-degree images; since 360Insta has no high-resolution ground truth, there is no direct check of what the super-resolved outputs actually recover.","fun_headline_variants_meta":{"raw":{"variants":["Real-world panorama super-resolution models underperform","Synthetic-trained 360° SR fails on real captures","New benchmark exposes 360° super-resolution gap","Authentic degradation dataset stumps omnidirectional SR"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000196,"raw_usage":{"total_tokens":1373,"prompt_tokens":968,"completion_tokens":405,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":584,"completion_tokens_details":{"reasoning_tokens":345}},"tokens_in":584,"tokens_out":405,"duration_ms":4462,"temperature":1.0,"reasoning_tokens":345,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T05:50:17.346296+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A subjective study would settle the claim: show viewers the original low-resolution frames, bicubic upscaling, and each model's scale-factor-2 and scale-factor-4 outputs from a subset of 360Insta on a head-mounted display, then correlate their quality ratings with NIQE, MUSIQ, MANIQA, and CLIPIQA. If human preference ranks the models differently, or rates the super-resolved frames as acceptable quality, the paper's conclusion that current methods do not perform well on real degradation would not stand.","supporting_citations":[{"cited_title":"Osrt: Om- nidirectional image super-resolution with distortion-aware trans- former,","cited_arxiv_id":null,"evidence_quote":"It supplies the OSRT distortion-aware transformer and the ODISR-clean and SUN360-clean fisheye-downsampling protocols that define the synthetic side of the comparison."},{"cited_title":"Lau-net: Latitude adaptive upscaling network for omnidirectional image super-resolution,","cited_arxiv_id":null,"evidence_quote":"It supplies the LAU-Net latitude-adaptive method and the ODISR and SUN360 ERP bicubic-downsampling benchmarks used in the synthetic evaluations."},{"cited_title":"Super-resolution of omnidi- rectional images using adversarial learning,","cited_arxiv_id":null,"evidence_quote":"It supplies the 360-SS adversarial baseline, a compared method whose outputs receive some of the worst no-reference scores on 360Insta."},{"cited_title":"Omnidirectional image super-resolution via bi-projection fusion,","cited_arxiv_id":null,"evidence_quote":"It supplies the BPOSR bi-projection fusion baseline, a strong synthetic-benchmark performer that still scores low on 360Insta."},{"cited_title":"Omnissr: Zero-shot omnidi- rectional image super-resolution using stable diffusion model,","cited_arxiv_id":null,"evidence_quote":"It supplies the OmniSSR zero-shot diffusion baseline included in both the synthetic and real-degradation comparisons."},{"cited_title":"Ntire 2023 challenge on 360deg omni- directional image and video super-resolution: Datasets, methods and results,","cited_arxiv_id":null,"evidence_quote":"It represents the prior challenge datasets built on synthetic downsampling that the survey positions 360Insta against."},{"cited_title":"Making a “com- pletely blind","cited_arxiv_id":null,"evidence_quote":"It is one of the four no-reference metrics used to judge the super-resolved outputs on 360Insta."},{"cited_title":"Musiq: Multi- scale image quality transformer,","cited_arxiv_id":null,"evidence_quote":"It is one of the four no-reference metrics used to judge the super-resolved outputs on 360Insta."},{"cited_title":"Maniqa: Multi-dimension attention network for no- reference image quality assessment,","cited_arxiv_id":null,"evidence_quote":"It is one of the four no-reference metrics used to judge the super-resolved outputs on 360Insta."}],"review_version":1}