{"id":"b02d7525-d9ba-4bd0-9bba-0e9bc35fe94b","arxiv_id":"2506.15276","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"MSNeRV is an implicit neural representation video codec that combines temporal-window fusion, GoP-level background grids, multi-resolution supervision, and multi-scale feature blocks, reporting strong compression results on HEVC ClassB and UVG.","lead":"MSNeRV is a neural video representation that fuses temporal and spatial multi-scale features to improve reconstruction of fast-moving, detailed video. It reports 4.5% bitrate savings over the VTM-23.7 codec on HEVC ClassB under PSNR, while on UVG it remains about 7.8% worse under PSNR.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 4.5% ClassB BD-rate advantage over VTM-23.7 RA is not yet established: aggregate-only results on five sequences, no QP ladder or per-sequence BDBR, and no released code; the UVG result is 7.8% worse, so the 'superior on average' claim is overbroad.","rationale":"The reader's CONDITIONAL verdict and weakest assumption are correct. The manuscript's own data show a split: ClassB PSNR BDBR is -4.5% while UVG is +7.8%, so the 'superior on average' wording is already too broad. Because the claimed advantage is small and unaccompanied by per-sequence or RD-point detail, it is impossible to distinguish a genuine gain from an artifact of QP selection or a single outlier. The architecture and ablations are detailed and plausible enough to warrant conditional acceptance, but the headline comparison needs the requested check. I also considered Equation (5) as a candidate concern; it is mathematically unusual as a training target, but it is a heuristic rather than the load-bearing step for the VTM comparison. No change to the reader's verdict is needed.","tokens_in":14425,"tokens_out":5798,"duration_ms":63629,"concrete_test":"Ask the authors to release the trained models, code, and the exact RD points behind Table 2. Independently recompute per-sequence BDBR for the five ClassB sequences with VTM-23.7 RA at QPs {22, 27, 32, 37} and again with an extended ladder {20, 25, 30, 35, 40}; if any individual sequence has a positive BDBR or the ClassB average changes sign under the extended ladder, the 'outperforms VTM' claim fails. If all five sequences remain negative and the average stays below roughly -2%, the concern is settled.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is empirical, and the load-bearing premise is that the ClassB comparison to VTM-23.7 RA is both stable and fair. The paper reports only aggregate BDBR in Table 2: -4.5% PSNR on ClassB but +7.8% on UVG. No per-sequence BDBR is given, no error bars, and the VTM configuration is described only as 'Random Access' with no QP ladder, no intra-period, and no bit-depth statement. The RD curves in Figure 6 are not tabulated. The measured gap is 4.5%, which is smaller than typical cross-sequence BD-rate dispersion for such comparisons, so a single outlier sequence could dominate the average. The contribution bullet also overstates the result by claiming 'superior compression performance to VTM on average' when the UVG PSNR row is 7.8% worse. Since code is not released, an independent check of the protocol is impossible. This does not invalidate the architecture's representation gains over HiNeRV shown in Table 1, but it makes the headline claim conditional.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes MSNeRV, an implicit neural representation (INR) framework for video compression. The method introduces a multi-scale temporal encoder that fuses base grids with sliding-window temporal features and GoP-level background grids, a multi-scale spatial decoder with multi-resolution supervision and a high-frequency boosting target, a scale-adaptive loss, and multi-scale feature blocks with cross-depth fusion. Experiments on HEVC ClassB and UVG datasets report PSNR gains over existing INR methods (e.g., 0.9–1.4 dB over HiNeRV at comparable model size) and claim BD-rate savings of 4.5% in PSNR and 30.5% in MS-SSIM against VTM-23.7 (Random Access) on ClassB, plus 42% and 28% bitrate savings versus HiNeRV on ClassB and UVG, respectively. The central claim is that MSNeRV outperforms the state-of-the-art conventional codec VTM on dynamic 1080p content.","tokens_in":14646,"tokens_out":7219,"duration_ms":69507,"significance":"If the reported results are reproducible, MSNeRV would be among the first INR-based video codecs to beat VTM random-access compression on 1080p dynamic content under PSNR, which is a notable advance for the field. The per-sequence representation results in Table 1 and the systematic ablations in Tables 3, 5, and 6 are useful empirical contributions, and the architecture components are clearly motivated. However, the headline VTM comparison rests on aggregate BD-rate numbers over only five ClassB sequences, with no per-sequence breakdown, no confidence intervals, and an incompletely specified codec configuration; the UVG PSNR result is 7.8% worse than VTM, which contradicts the unqualified 'superior on average' claim. The absence of released code and the selection of hyperparameters on the same evaluation set further reduce confidence. These issues make the central empirical claim conditional pending a more rigorous evaluation.","major_comments":[{"comment":"The central claim that MSNeRV outperforms VTM-23.7 (RA) is based solely on aggregate BDBR values over five HEVC ClassB sequences, with no per-sequence BDBR, no confidence intervals, and no description of the VTM configuration (QP ladder, intra period, bit depth, encoder options). The PSNR BDBR gain is only -4.5% on ClassB while the UVG PSNR row is +7.8% (worse), so the observed advantage is small relative to typical cross-sequence dispersion. Please provide per-sequence BDBR tables, specify the exact VTM settings used, report error bars or the per-sequence range, and explicitly scope the claim to the ClassB PSNR comparison.","section":"§4.2, Table 2"},{"comment":"The rate-distortion comparison does not state how many RD points were used for each codec, which QPs were selected for VTM and DCVC, or how the bitrates of INR-based methods were matched to those points. BDBR is sensitive to the placement and number of RD points, so without this information the -4.5% value is not reproducible. Please tabulate the RD points and describe the bitrate-control protocol for each method.","section":"§4.2, Figure 6"},{"comment":"Several design choices—temporal window size (Appendix C.2), layer depths (Appendix C.3), SA-loss coefficients (Table 4), and downsampling type (Table 5)—are tuned on the same HEVC ClassB sequences used for the headline VTM comparison. This selection on the test set can bias the measured BDBR gain. Please either validate on a held-out set or report a sensitivity analysis showing how the ClassB BDBR changes when these hyperparameters are varied within reasonable ranges.","section":"§4.1, §4.2, Appendix C"},{"comment":"The high-frequency boosting target Vboost = Vgt + H(|Vgt - XN|) is defined using the absolute value of the prediction residual, which is a non-smooth, phase-rectified signal; filtering this quantity does not extract the true high-frequency content in a phase-correct manner, and the target is non-stationary because XN changes during training. The high-pass filter is not specified (filter type, kernel, normalization). Please justify this target, provide the exact HPF implementation, and show that the 0.10 dB gain in the V3 ablation is robust across bitrates and training seeds; if it is not, consider removing the component or replacing it with a more conventional high-frequency loss.","section":"§3.2.2, Eq. (5)"},{"comment":"The ablation study reports a single operating point per variant, and several differences are 0.05–0.10 dB PSNR, which is within typical run-to-run variation for neural representation training. Without multiple seeds or confidence intervals, the claim that each component contributes meaningfully is not supported. Please add variance estimates or demonstrate deterministic training with a fixed seed.","section":"§4.3, Table 3"},{"comment":"The claim of 'superior compression performance to VTM on average' is overbroad: the UVG PSNR row in Table 2 is +7.8% worse, and the average over ClassB and UVG is not defined. The abstract's statement that MSNeRV 'surpasses VTM-23.7 (Random Access) in dynamic scenarios' should be scoped explicitly to the HEVC ClassB dataset and to the PSNR metric; otherwise the claim is misleading.","section":"Abstract, §1, §4.2"}],"minor_comments":[{"comment":"The notation in the temporal encoder is inconsistent: the summation in Eq. (2) is written as 'i=t+l−1X i=t', and Eq. (4) introduces 'γ base.(t)' while Eq. (2) uses 'X i base'; please unify the notation and define all variables (e.g., w_t^i, γ_GoP(k)) clearly. Also, 'By apply a sliding window mechanism' should be 'By applying a sliding window mechanism'.","section":"§3.1, Eq. (2), Eq. (4)"},{"comment":"The subscript in '3x3 convolution layer Nr' should be N_r, not 'Nr'. Additionally, the sentence 'Since this metric primarily focuses on the perceptual similarity that is crucial for the final reconstructed video.' is a sentence fragment; please revise it.","section":"§3.2.1"},{"comment":"The text says 'MS-SSIM is used to evaluate the subjective quality of reconstructed videos.' MS-SSIM is an objective metric, not a subjective one; please replace 'subjective quality' with 'perceptual quality'.","section":"§4.2"},{"comment":"Reference [11] appears to be an unrelated paper on measuring ocular torsion, and reference [50] describes VTM 17 while the text reports VTM-23.7; please correct the citations to match the VTM version actually used.","section":"References"},{"comment":"The figure caption for Figure 6 does not state which dataset or bitrate/quality axes are shown, nor whether the RD curves are averaged over sequences or represent a single sequence; please add a descriptive caption.","section":"Figure 6"}],"recommendation":"major_revision","confidential_remarks":"The manuscript's headline empirical claim is not yet supported with the current level of evaluation detail. The authors should be asked to provide per-sequence BD-rate data, full VTM configuration, RD point tables, and either code or sufficiently detailed protocols to make the comparison reproducible. The overbroad claims in the abstract and introduction should be tempered. If the authors address these points, the paper could be a useful contribution to INR-based video coding."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: the paper's real contribution is the architecture combination and the representation improvements, not the VTM claim. The VTM comparison is real but under-supported; if you need the headline, ask for code before believing it.\n\nWhat's actually new: they fuse temporal windows with learned weights, GoP-level grids, multi-resolution supervision, high-frequency boosting, and cross-depth fusion into one INR codec. No single component is revolutionary—the window fusion and GoP grids do useful work, and the ablation study backs them up. Table 1 shows consistent PSNR gains over HiNeRV and HNeRV-boost at matched model sizes across ClassB and UVG. That part I trust. The inpainting/interpolation appendix is a nice bonus and supports the claim that the learned grids generalize.\n\nWhere it gets soft: the headline is a 4.5% BD-rate saving over VTM-23.7 RA on ClassB, from five sequences, with no per-sequence BDBR, no error bars, no QP ladder details, and no code. The UVG PSNR row is +7.8% worse than VTM, so the contribution bullet 'superior compression performance to VTM on average' is overbroad for the evidence. Equation (5) is also mathematically odd: it adds the high-passed absolute residual to the ground truth and uses that as the loss target, which looks like it rewards matching a version of the video that contains artifacts of the current reconstruction error. The ablation doesn't isolate what that term actually does. These are not fatal—the architecture's gains over HiNeRV are likely real—but the VTM comparison needs either code, per-sequence numbers, or a verified protocol before it should be cited as a VTM-beating result.\n\nBottom line: worth a serious referee, not a desk reject. It's a competent incremental paper with an overreaching headline. If you review it, push for code and per-sequence BDBR.","headline":"Solid incremental INR codec with believable representation gains, but the VTM-beating headline is under-evidenced and needs code or per-sequence data.","tokens_in":15226,"tokens_out":1573,"would_cite":true,"duration_ms":14989,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"MSNeRV outperforms the VTM random-access codec on dynamic 1080p video under both PSNR and MS-SSIM.","keywords":["implicit neural representation","video compression","multi-scale feature fusion","temporal window","GoP-level grid","multi-resolution supervision","high-frequency boosting","rate-distortion optimization"],"falsifier":"Rerun the comparison with per-sequence BD-rate tables and confidence intervals across the five ClassB sequences under at least two distinct VTM-23.7 random-access configurations; if the average PSNR BD-rate advantage over VTM fails to reproduce, or inverts on any single sequence, the central claim is not robust.","tokens_in":14161,"feed_emoji":"🎥","tokens_out":6914,"duration_ms":65223,"temperature":0.7,"pith_summary":"This paper argues that implicit neural representations can be made competitive with, and in dynamic scenes better than, a state-of-the-art conventional video codec, if the network is designed around the multi-scale structure of video. It introduces MSNeRV, an INR codec whose encoder fuses frame-level, window-level, and GoP-level grids for temporal consistency, and whose decoder supervises intermediate resolutions and boosts high-frequency residuals. The reported result is a 4.5% BD-rate saving under PSNR and 30.5% under MS-SSIM over VTM-23.7 Random Access on HEVC ClassB, together with 42% and 28% bitrate savings versus HiNeRV on ClassB and UVG. The reason to care: if this holds, a parameter-based codec that needs no training prior can edge past a mature block-based standard on content where motion and detail are hardest.","feed_headline":"Neural video codec edges past VVC reference on 1080p motion","feed_subtitle":"On HEVC ClassB, MSNeRV saves 4.5% bitrate over VTM under PSNR and 30.5% under MS-SSIM.","key_machinery":"The central object is the multi-scale fused grid: a per-frame learnable base grid $\\gamma_{\\mathrm{base}}(t)$, a sliding temporal-window fusion producing $X_{\\mathrm{temporal}}$, and a GoP-level background grid $\\gamma_{\\mathrm{GoP}}(k)$, summed into $X_{\\mathrm{fused}}$. This feeds a decoder built from Multi-Scale Feature blocks, each combining bilinear and pixel-shuffle upsampling with a hierarchical local-grid encoding, fusing depth-wise convolutions of different kernel sizes, and concatenating features across depths. The machinery works by making the network reuse low- and mid-resolution features through multi-resolution supervision, high-frequency boosting, and a scale-adaptive loss, so that representation capacity rises without increasing model size.","core_discovery":"MSNeRV claims to be an INR-based codec that, on average, surpasses VTM-23.7 with Random Access configuration on HEVC ClassB, a dataset dominated by dynamic and detailed 1080p content. The central discovery is that intermediate features of the upsampling decoder and multi-scale temporal context are under-used resources: capturing them yields better rate-distortion performance at the same model size. The paper's mechanism combines temporal windows with learnable weights, GoP-level background grids, multi-resolution supervision, high-frequency boosting, and multi-scale feature blocks with cross-depth fusion. In video representation experiments, MSNeRV reports higher PSNR than HiNeRV and HNeRV-Boost at comparable model sizes, and in compression experiments it reports the BD-rate savings summarized above.","pith_inferences":["Editorial inference: the temporal-window and GoP-grid fusion mechanism is not specific to video regression; the same idea could be dropped into any coordinate-based representation of dynamic scenes to improve temporal consistency.","Editorial inference: because the paper reports bitrate savings only under one VTM-23.7 random-access configuration, the 4.5% figure should be read as the result of that protocol; whether it transfers to other VTM configurations or content classes is untested.","Editorial inference: the 30-epoch QAT schedule and channel-based bitrate control suggest two cheap experiments the paper does not run—longer QAT or a stronger entropy model—that could either widen or shrink the reported gap."],"forward_implications":["If the reported results hold, INR-based codecs can surpass a mature block-based codec on dynamic 1080p content under PSNR, not only under perceptual metrics.","Because the gains come from fusing existing internal features rather than enlarging the network, the approach points to a way of increasing representation capacity without increasing bitrate.","Multi-resolution supervision and high-frequency boosting imply that INR training can use intermediate decoder features as explicit targets, which may improve training efficiency and stability.","On UVG, MSNeRV is within 7.8% of VTM under PSNR and better than VTM by 26.2% under MS-SSIM, suggesting the codec is competitive across content types while being perceptually strong.","Compared with HiNeRV under the same coding scheme, MSNeRV saves 42% bitrate on ClassB and 28% on UVG, a direct corollary of its representation experiments."],"supporting_citations":[{"why":"Supplies the VTM test-model configuration used as the conventional-codec anchor in the headline BD-rate comparison.","marker":"[50]"},{"why":"HiNeRV is the main INR baseline whose bitrate savings MSNeRV reports on both datasets.","marker":"[4]"},{"why":"HNeRV-Boost is the conditional-decoder baseline compared in both representation and compression experiments.","marker":"[23]"},{"why":"FFNeRV provides the grid-based positional encoder that MSNeRV adopts and extends with temporal fusion.","marker":"[15]"},{"why":"DCVC-DC is the learning-based codec included in the compression comparison table.","marker":"[7]"},{"why":"DCVC-FM is the other learning-based codec used to position MSNeRV against learned video compression.","marker":"[8]"},{"why":"Provides the entropy coding applied to generate the actual bitstreams for INR model parameters.","marker":"[48]"},{"why":"Defines the HEVC ClassB dataset on which the headline claim of beating VTM is made.","marker":"[16]"},{"why":"Defines the UVG dataset used for the second set of representation and compression evaluations.","marker":"[17]"}],"fun_headline_variants":["MSNeRV multi-scale fusion beats VTM on dynamic 1080p","Neural codec MSNeRV outdoes VVC reference on high motion","Multi-scale feature fusion pushes neural video codec past VTM","MSNeRV surpasses VVC reference on dynamic 1080p with multi-scale fusion","Neural video codec beats VVC on fast-moving 1080p with multi-scale fusion"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The headline rests on averaging BD-rate over five HEVC ClassB sequences under one fixed VTM-23.7 random-access configuration and one fixed training and QAT schedule; if that protocol is not representative, the reported 4.5% PSNR advantage could invert.","fun_headline_variants_meta":{"raw":{"variants":["MSNeRV multi-scale fusion beats VTM on dynamic 1080p","Neural codec MSNeRV outdoes VVC reference on high motion","Multi-scale feature fusion pushes neural video codec past VTM","MSNeRV surpasses VVC reference on dynamic 1080p with multi-scale fusion","Neural video codec beats VVC on fast-moving 1080p with multi-scale fusion"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000821,"raw_usage":{"total_tokens":3586,"prompt_tokens":933,"completion_tokens":2653,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":549,"completion_tokens_details":{"reasoning_tokens":2548}},"tokens_in":549,"tokens_out":2653,"duration_ms":19396,"temperature":1.0,"reasoning_tokens":2548,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T19:38:01.731177+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Rerun the comparison with per-sequence BD-rate tables and confidence intervals across the five ClassB sequences under at least two distinct VTM-23.7 random-access configurations; if the average PSNR BD-rate advantage over VTM fails to reproduce, or inverts on any single sequence, the central claim is not robust.","supporting_citations":[{"cited_title":"Algorithm description for versatile video coding and test model 17 (vtm 17),","cited_arxiv_id":null,"evidence_quote":"Supplies the VTM test-model configuration used as the conventional-codec anchor in the headline BD-rate comparison."},{"cited_title":"Hin- erv: Video compression with hierarchical encoding-based neural representation,","cited_arxiv_id":null,"evidence_quote":"HiNeRV is the main INR baseline whose bitrate savings MSNeRV reports on both datasets."},{"cited_title":"Boosting neural representations for videos with a conditional decoder,","cited_arxiv_id":null,"evidence_quote":"HNeRV-Boost is the conditional-decoder baseline compared in both representation and compression experiments."},{"cited_title":"Ffnerv: Flow-guided frame-wise neural representations for videos,","cited_arxiv_id":null,"evidence_quote":"FFNeRV provides the grid-based positional encoder that MSNeRV adopts and extends with temporal fusion."},{"cited_title":"Neural video compression with diverse contexts,","cited_arxiv_id":null,"evidence_quote":"DCVC-DC is the learning-based codec included in the compression comparison table."},{"cited_title":"Neural video compression with feature modulation,","cited_arxiv_id":null,"evidence_quote":"DCVC-FM is the other learning-based codec used to position MSNeRV against learned video compression."},{"cited_title":"Practical full resolution learned lossless image compression,","cited_arxiv_id":null,"evidence_quote":"Provides the entropy coding applied to generate the actual bitstreams for INR model parameters."},{"cited_title":"Overview of the high efficiency video coding (hevc) stan- dard,","cited_arxiv_id":null,"evidence_quote":"Defines the HEVC ClassB dataset on which the headline claim of beating VTM is made."},{"cited_title":"Uvg dataset: 50/120fps 4k sequences for video codec analysis and develop- ment,","cited_arxiv_id":null,"evidence_quote":"Defines the UVG dataset used for the second set of representation and compression evaluations."}],"review_version":2}