{"id":"dc6ab5aa-b8d0-4142-973b-cb031b115c5b","arxiv_id":"2509.03922","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A fully learned multiview video codec that keeps single-view decodability and random access, cutting bitrate about 25% to 31% versus the MV-HEVC reference by conditioning dependent views on the independent view's decoded features.","lead":"This paper builds a learned compression system for videos captured by several cameras at once. It saves about 25% of bitrate versus the standard multiview codec by letting the extra camera streams reuse motion and texture information already stored in the first camera's stream.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Training on synthetic warps leaves wide-baseline/parallax generalization of the content-side inter-view modules unverified; the central claim should remain conditional pending a wide-baseline test.","rationale":"The central claim is plausible and internally consistent: the independent view is exactly DCVC-B, giving backward compatibility; hierarchical B provides random access; the ablation attributes the gains to the four proposed modules; and the runtime overhead is modest. I credit the use of standard MV-HEVC CTC test sequences and the explicit limitation statement in Sec. IV.D. The weakest load-bearing premise is not an internal inconsistency but the domain over which the inter-view conditioning is valid. Because the only training signal for inter-view geometry is affine/homography warps, and the content modules are deliberately disparity-free, the quantitative contribution of IVCP/IVCEM is only directly validated for the tested narrow-baseline regime. This does not refute the stated claim on the standard test sequences, but it keeps the verdict at CONDITIONAL pending a wide-baseline/parallax check and, ideally, release of code/weights for independent reproduction. I do not elevate to reject because no internal error is identified; the reported numbers, if reproduced, support the claim for the evaluated content.","tokens_in":18360,"tokens_out":12961,"duration_ms":138910,"concrete_test":"Run LMVC, M0 (DCVC-B), and HTM-16.3 on a held-out real wide-baseline multiview set, e.g., Ballet and Breakdancers (8 cameras, large baseline, occlusions), using the same intra-period and lambda settings as Tables II/III, and compute (i) Ours-vs-DCVC-B BD-rate and (ii) the incremental BD-rate of M4 over M3. If the IVCP+IVCEM increment disappears or reverses relative to the 6.2-pt contribution in Table V, or the Ours-vs-DCVC-B gain falls below the motion-side ~5.3-pt contribution, the content-side modules are overfit to the synthetic affine/homography training distribution and the claim should be restricted to narrow-baseline content.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The content-side gains (IVCP + IVCEM = 6.2 of the 11.5-pt ablation gain, Table V) rest on the assumption in Sec. III.D that an H/8 decoded independent-view feature is a useful inter-view context without explicit disparity compensation. This assumption is only trained on affine/homography warps of single-view videos (Sec. III.F), which cannot generate true parallax, occlusions, or wide-baseline disparities. On real camera arrays with depth variation, a spatially concatenated feature at H/8 may be misaligned by more than the effective receptive field; the network may then fail to use it or use it incorrectly. The measured 3-view/2-view gains come from the MV-HEVC CTC sequences (Table I), which are predominantly narrow-baseline natural content; they do not establish transfer to wide-baseline/occluded multiview content. The authors' own Sec. IV.D limitation shows the codec is sensitive to content shift on Shark/Undo Dancer, so robustness to geometric distribution shift is genuinely unverified. If this premise fails on real wide-baseline content, the content-side advantage over DCVC-B disappears, lowering the headline advantage from about 10-12 pts to roughly the motion-side ~5 pts.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes LMVC, an end-to-end learned multiview video coding framework built on DCVC-B. The independent view is compressed with a frozen DCVC-B codec, and the dependent views extend DCVC-B with four modules: inter-view motion vector prediction (IVMVP), an inter-view motion entropy model (IVMEM), a disparity-free inter-view context prediction module (IVCP), and an inter-view contextual entropy model (IVCEM). Training relies on synthesized multiview sequences obtained by affine and homography warping of single-view videos. Evaluation uses MV-HEVC common test condition sequences in 2-view and 3-view configurations. The paper reports average BD-rate reductions of 24.9% over HTM-16.3 for 3-view and 31.2% for 2-view, with gains over a DCVC-B simulcast anchor. An ablation study attributes an 11.5% BD-rate improvement to the proposed modules, and a per-stream bitrate analysis shows reductions in both motion and content bitrate. The framework preserves backward compatibility and random access.","tokens_in":18702,"tokens_out":6295,"duration_ms":64291,"significance":"If the reported results hold, this is a notable advance: it is one of the first fully end-to-end learned multiview codecs with random access and backward compatibility, and it outperforms both a traditional MV-HEVC reference and a learned simulcast anchor on standard test content. The evaluation is grounded in standard CTC sequences, and the comparison against DCVC-B uses the same frozen independent-view codec, which isolates the contributions of the dependent-view modules. The ablation study and per-stream bitrate analysis are informative. However, the training data are synthetic warps of single-view videos, and the test set consists mostly of narrow-baseline natural content. The generality of the content-side inter-view modules to real wide-baseline camera arrays, with true parallax and occlusions, is not established. The paper's own negative results on animated sequences suggest sensitivity to domain shift. The headline claims should therefore be regarded as conditional on the evaluated content distribution.","major_comments":[{"comment":"The dependent-view codec is trained only on affine/homography warps of single-view videos. Such synthetic data cannot generate true parallax, occlusion, or wide-baseline disparity. The content-side modules (IVCP and IVCEM) rely on the premise in Sec. III.D that an H/8 independent-view feature remains useful as inter-view context without explicit disparity compensation. The MV-HEVC CTC test set is predominantly narrow-baseline natural content; the only content-shift stress reported is animation (Shark, Undo Dancer), not geometric baseline. To support the general claim of a multiview codec, please add a wide-baseline or real-camera-array evaluation, or explicitly scope the claims to narrow-baseline content, and report whether the IVCP/IVCEM gains persist under such conditions.","section":"Section III.F, Section IV.D, Tables II-III"},{"comment":"The ablation table does not state the anchor codec, the view configuration (2-view vs 3-view), the number of test sequences, or the rate points used to compute BD-rate. Since the abstract and Sec. IV.B.1 attribute the overall gain to the four modules, these conditions are necessary for reproducibility and for interpreting the 11.5% figure. Please specify the protocol (e.g., same 3-view set and HTM anchor as Table II) and report per-model BD-rate values on the same sequences.","section":"Section IV.C, Table V"},{"comment":"All BD-rate values are point estimates from a single trained model. Given the large spread across sequences (e.g., 3-view values range from -56.8% to +39.7%), the averages may be sensitive to particular sequences. Please provide error bars, multiple training runs, or at least per-rate-point average RD curves for the headline comparisons, so that the claimed gains over DCVC-B can be assessed with some measure of variability.","section":"Tables II and III"}],"minor_comments":[{"comment":"The manuscript header and full-text title use 'LMVC', while the arXiv metadata and the provided abstract use 'DCVC-MV'. Please unify the title and abstract across all versions.","section":"Title/Abstract"},{"comment":"There are several typos and acronym inconsistencies, e.g., 'IVMP' vs 'IVMVP', 'Ballons' for 'Balloons', 'GT F ly' with a stray space. A careful proofreading pass is needed.","section":"Throughout"},{"comment":"The text says 'We use AdamW optimizer' but reference [65] is cited as 'Adam: A method for stochastic optimization'. Please update the reference to the AdamW source or correct the optimizer name.","section":"Section IV.A.2"},{"comment":"The phrase 'more than 10% BD-rate reduction' relative to DCVC-B could be misread as a relative improvement. Please clarify whether this is a percentage-point difference in BD-rate (24.9 vs 14.7) or a relative rate reduction.","section":"Section IV.B.1"},{"comment":"The choice of H/8 resolution for the inter-view context is motivated by a citation to prior work, but no ablation or quantitative justification is given. Please add a short analysis or ablation on context resolution, or at least discuss the receptive field / alignment limits.","section":"Section III.D"},{"comment":"The related work discusses [38] as a hybrid INR-based multiview codec, but no experimental comparison is made. A brief sentence explaining why direct comparison is omitted (e.g., different input modalities or evaluation protocols) would help.","section":"Section II.B"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is technically sound within the evaluated narrow-baseline regime, and the use of a frozen DCVC-B anchor is a clean experimental design. The main risk is the mismatch between the broad 'multiview video coding' claim and the synthetic training data plus narrow-baseline test set. If the authors add a wide-baseline evaluation or explicitly limit the scope, I would be willing to accept a revised version. Also, please resolve the title inconsistency between the arXiv listing and the manuscript before publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe short take: this is the first fully end-to-end learned multiview video codec (2-view and 3-view) that keeps random access and backward compatibility, and the reported gains over HTM-16.3 (24.9% on 3-view, 31.2% on 2-view) and roughly 10 points over a DCVC-B simulcast are credible on the evidence shown. The main caveat is that the content-side inter-view modules are trained only on synthetic affine/homography warps of single-view video, so how well they transfer to real wide-baseline arrays with parallax and occlusion is unverified.\n\nWhat the paper does well: the evaluation is fair and reasonably careful. The independent view uses the same frozen DCVC-B in the baseline and the proposed system, so the gain isolates the dependent-view modules. The test set is the MV-HEVC common test conditions. The ablation (Table V) and the per-stream bitrate analysis (Section IV.C.2) consistently attribute the gain to the four new modules. The authors also state a limitation about animation content, which is more honesty than many papers in this area show.\n\nWhere it is soft: (1) The wide-baseline concern from the stress-test note is legitimate. The disparity-free context at H/8 may misalign on real camera arrays with depth variation, and the paper does not test on such content. The poor results on Shark and Undo Dancer show sensitivity to distribution shift; the fact that GT Fly, also a CG sequence, shows gains makes the stated 'animation sequences' limitation under-specified. (2) No comparison to the closest learned baselines: [38] is described but not run, and no stereo codec is included on the 2-view case. (3) All numbers are single-run, with no error bars. These are addressable rather than fatal.\n\nThe stress-test concern is not a refutation of the measured claim, but it does mean the central claim should be read as: works on narrow-baseline natural content, transfer unverified. That is the right scope until a wide-baseline experiment is added.\n\nBottom line: this is a serious, useful contribution to learned multiview coding, with genuine novelty and a mostly sound evaluation. I would cite it, bring it to a reading group, and send it to capable referees.","headline":"First fully learned multiview video codec with random access and backward compatibility; measured gains on standard CTC are credible, but content-side transfer to wide-baseline/parallax content is unverified.","tokens_in":19223,"tokens_out":4756,"would_cite":true,"duration_ms":42545,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A fully learned multiview codec beats the MV-HEVC reference software by an average of 24.9% on 3-view natural-content test sequences, while keeping the primary view independently decodable.","keywords":["learned multiview video coding","inter-view prediction","deep contextual video compression","entropy model","backward compatibility","random access","BD-rate","MV-HEVC"],"falsifier":"Run the trained LMVC codec on a real camera-array multiview sequence with wide baseline, occlusions, and no animation, never seen in training, and compare BD-rate against HTM-16.3 and DCVC-B; if it fails to beat DCVC-B simulcast by a meaningful margin, or fails to beat HTM on natural content, the central compression claim is falsified. Equivalently, ablation of the inter-view content modules on such content should show whether their reported contribution persists outside the synthetic training distribution.","tokens_in":18224,"feed_emoji":"🎥","tokens_out":6156,"duration_ms":59988,"temperature":0.7,"pith_summary":"The paper aims to establish that a fully end-to-end learned multiview video codec can outperform the traditional MV-HEVC reference and a learned simulcast baseline, while still offering the two properties practical systems require: the primary view's bitstream decodes with an ordinary single-view decoder, and any viewpoint can be accessed randomly. To do this, it extends the single-view learned B-frame codec DCVC-B with four inter-view modules that transfer motion and content information from the independently coded view into the dependent views. If the reported BD-rate numbers are right, learned multiview coding becomes a viable alternative to standardized hybrid codecs for 3D, VR, and free-viewpoint content, at a modest computational cost.","feed_headline":"Learned multiview codec beats MV-HEVC by 24.9 percent","feed_subtitle":"End-to-end dependent-view coding borrows inter-view motion and content priors while keeping the primary view independently decodable.","key_machinery":"The load-bearing mechanism is the use of the already-decoded independent view as a conditional prior for dependent-view compression, without explicit disparity estimation. Decoded motion-difference features and motion vectors condition the dependent motion encoder and its entropy model; a downsampled decoded content feature is passed through a modified context-mining feature extractor to predict an inter-view context; and both motion and content latent distributions are estimated by a quadtree-partition spatial context model that integrates temporal, inter-view, and hyper priors. This moves inter-view redundancy removal into the feature and latent domain rather than the pixel-disparity domai","core_discovery":"The paper proposes an end-to-end learned multiview video coding framework, named LMVC in the text and DCVC-MV in the title, built on the single-view B-frame codec DCVC-B. It adds four inter-view modules: feature-based inter-view motion vector prediction, an inter-view motion entropy model, a disparity-free inter-view context prediction module, and an inter-view contextual entropy model. The independent view is compressed exactly as in DCVC-B and can be decoded without any other view, while each dependent view is conditioned on decoded motion features, motion vectors, low-resolution content features, and latent representations from the independent view. Reported results on the MV-HEVC common","pith_inferences":["Because training uses only synthetic multiview sequences made by affine- and homography-warping single-view videos, the architecture may need retraining or fine-tuning on real camera-array captures before its gains transfer to wide-baseline, occluded, or parallax-heavy scenes; the paper's own negative results on the Shark and Undo Dancer animation sequences point to this distribution-shift sensiti","The disparity-free context transfer is not limited to one dependent view in principle: a nearest already-coded view could serve as the independent prior for additional views, though the paper only reports 2-view and 3-view configurations.","Adding a depth coding layer, which the paper lists as future work, could let inter-view priors condition on warped depth instead of raw downsampled features, potentially recovering performance on occluded regions.","A testable extension is to add animation-style or synthetic-rendered content to the training set and measure whether the 39.7% BD-rate penalty on Shark shrinks, which would confirm that the failure is purely a training-distribution effect."],"forward_implications":["If the reported numbers hold, fully learned multiview codecs can replace MV-HEVC for natural-content 3D and free-viewpoint video, saving roughly 25% of bit rate on average in the 3-view setting.","The independent view's bitstream remains backward-compatible because it uses unmodified DCVC-B: a single-view decoder can decode it, and dependent-view data never affects its reconstruction.","Random access is preserved through the hierarchical B-frame structure, so viewers can jump between perspectives without decoding the entire multiview stream.","The gains come from real bit-rate reductions in dependent-view motion and content coding, not from a quality trade-off: on Balloons, motion bit rate falls to between 14.3% and 30.6% of DCVC-B's rate and content bit rate to between 58.0% and 75.4%, depending on the operating point.","The computational overhead over DCVC-B is modest: encoding time rises from 0.50s to 0.57s per frame, decoding from 0.38s to 0.44s, MACs/pixel from 3004K to 3457K, and model size from 21.40M to 29.13M parameters."],"supporting_citations":[{"why":"Supplies the base single-view B-frame codec (DCVC-B) that the framework extends, the hierarchical reference structure, and the primary baseline it must beat.","marker":"[40]"},{"why":"Supplies the pre-trained optical-flow network (SpyNet) used to estimate the bi-directional motion vectors for dependent views.","marker":"[59]"},{"why":"Supplies the temporal context-mining feature extractor that is modified to predict inter-view contexts, and the downsampling-mitigates-misalignment principle used by the disparity-free design.","marker":"[60]"},{"why":"Supplies the quadtree-partition spatial context model used in the motion and content entropy models, and the shared I-frame codec.","marker":"[61]"},{"why":"Supports the premise that downsampling reduces inter-frame or inter-view misalignment, which the inter-view context prediction relies on.","marker":"[62]"},{"why":"Defines the standardized multiview test sequences and coding configurations used for evaluation and comparison against HTM.","marker":"[63]"},{"why":"Supplies the single-view training videos (Vimeo-90k) that are warped into synthetic multiview sequences for training the dependent-view codec.","marker":"[64]"}],"fun_headline_variants":["DCVC-MV: Learned codec cuts multiview video bitrate by 24.9%","Multiview video compression gets 24.9% better with deep context","New learned codec slashes multiview video data by a quarter","Deep codec for multiview video saves 24.9% bandwidth","Inter-view priors boost multiview video compression by 24.9%"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The framework's gains rest on training only on synthetic multiview sequences produced by warping single-view videos, plus the premise that downsampling removes inter-view misalignment without explicit disparity estimation; if real camera arrays present wider baselines, parallax, occlusion, or animated content, the inter-view modules' contribution to the compression gain may not transfer.","fun_headline_variants_meta":{"raw":{"variants":["DCVC-MV: Learned codec cuts multiview video bitrate by 24.9%","Multiview video compression gets 24.9% better with deep context","New learned codec slashes multiview video data by a quarter","Deep codec for multiview video saves 24.9% bandwidth","Inter-view priors boost multiview video compression by 24.9%"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000629,"raw_usage":{"total_tokens":2767,"prompt_tokens":787,"completion_tokens":1980,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":531,"completion_tokens_details":{"reasoning_tokens":1877}},"tokens_in":531,"tokens_out":1980,"duration_ms":14152,"temperature":1.0,"reasoning_tokens":1877,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T10:32:44.756204+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the trained LMVC codec on a real camera-array multiview sequence with wide baseline, occlusions, and no animation, never seen in training, and compare BD-rate against HTM-16.3 and DCVC-B; if it fails to beat DCVC-B simulcast by a meaningful margin, or fails to beat HTM on natural content, the central compression claim is falsified. Equivalently, ablation of the inter-view content modules on such content should show whether their reported contribution persists outside the synthetic training distribution.","supporting_citations":[{"cited_title":"Bi-directional deep contextual video compression,","cited_arxiv_id":null,"evidence_quote":"Supplies the base single-view B-frame codec (DCVC-B) that the framework extends, the hierarchical reference structure, and the primary baseline it must beat."},{"cited_title":"Optical flow estimation using a spatial pyramid network,","cited_arxiv_id":null,"evidence_quote":"Supplies the pre-trained optical-flow network (SpyNet) used to estimate the bi-directional motion vectors for dependent views."},{"cited_title":"Temporal context min- ing for learned video compression,","cited_arxiv_id":null,"evidence_quote":"Supplies the temporal context-mining feature extractor that is modified to predict inter-view contexts, and the downsampling-mitigates-misalignment principle used by the disparity-free design."},{"cited_title":"Neural video compression with diverse contexts,","cited_arxiv_id":null,"evidence_quote":"Supplies the quadtree-partition spatial context model used in the motion and content entropy models, and the shared I-frame codec."},{"cited_title":"Spatial decomposition and temporal fusion based inter prediction for learned video compression,","cited_arxiv_id":null,"evidence_quote":"Supports the premise that downsampling reduces inter-frame or inter-view misalignment, which the inter-view context prediction relies on."},{"cited_title":"Common test conditions of 3DV core experiments, document JCT3V-E1100,","cited_arxiv_id":null,"evidence_quote":"Defines the standardized multiview test sequences and coding configurations used for evaluation and comparison against HTM."},{"cited_title":"Video enhance- ment with task-oriented flow,","cited_arxiv_id":null,"evidence_quote":"Supplies the single-view training videos (Vimeo-90k) that are warped into synthetic multiview sequences for training the dependent-view codec."}],"review_version":1}