{"id":"67a8f4d3-bd22-496f-bff8-9711e20a1de6","arxiv_id":"2606.13110","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"JOMP jointly optimizes mixed-precision quantization parameters and bit widths across neural video coding frameworks, achieving rate-distortion performance comparable to DCVC-FM while cutting bit operations by 87.6%.","lead":"JOMP is a framework that treats quantization bit widths as learnable variables during training of neural video codecs to enable efficient integer implementations. Smart generalists might read it to see how AI-based video compression can move from floating-point research models to practical low-power hardware.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"Learnable bit widths may fail to yield stable/generalizable assignments without post-training adjustments or per-framework retuning","rationale":"The reader's weakest_assumption exactly isolates the load-bearing condition for the headline claim. Full-text details on training schedule, regularization, or ablation on bit-width variance would be needed to refute it; absent those, the concern stands and keeps the verdict at UNVERDICTED.","tokens_in":1798,"tokens_out":322,"duration_ms":13043,"concrete_test":"Re-train the best-performing model with JOMP on one framework (e.g., the temporal-buffering variant), freeze the learned bit widths, and evaluate zero-shot on a second framework without any retuning or fine-tuning; if BD-rate increases by >5% or bit-op savings drop below 80% relative to the reported numbers, the stability/generalizability assumption does not hold.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim requires that jointly optimizing quantization parameters and bit widths as learnable variables during end-to-end training directly produces usable mixed-precision integer codecs with the stated RD parity to DCVC-FM and 87.6% bit-op reduction. In quantization-aware training this is non-trivial: bit-width variables are typically relaxed via straight-through estimators or Gumbel-softmax and can collapse or require auxiliary losses/annealing to remain stable. The abstract and claimed systematic study across frameworks and buffering strategies provide no evidence that the learned widths transfer without retuning or that training dynamics were monitored for collapse or variance across random seeds.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper claims to introduce JOMP, the first mixed-precision quantization framework for neural video codecs, in which both quantization parameters and bit widths are treated as learnable variables during end-to-end training. It validates effectiveness via a systematic investigation across coding frameworks and temporal buffering strategies, develops a complete integerization pipeline for deterministic decoding, and reports that application to the best-performing model yields RD performance comparable to DCVC-FM while reducing bit operations by 87.6%.","tokens_in":1928,"tokens_out":425,"duration_ms":22223,"significance":"If the central claims hold under rigorous validation, the work would meaningfully advance practical deployment of neural video codecs by addressing high computational complexity through mixed-precision integer arithmetic. The systematic cross-framework and cross-buffering study could provide a unified perspective on practicality considerations, and the integerization pipeline represents a concrete contribution toward reproducible integer implementations.","major_comments":[{"comment":"Abstract: the central claim of RD performance comparable to DCVC-FM with an 87.6% bit-operation reduction is presented without experimental details, baselines, variance across seeds, or ablation results on the joint optimization; this directly prevents assessment of whether the learnable bit-width procedure produces stable and generalizable assignments as required by the weakest assumption.","section":"Abstract"},{"comment":"Training procedure (assumed §3): no description is given of the relaxation technique for bit-width variables (e.g., straight-through estimator or Gumbel-softmax), auxiliary losses, or monitoring for collapse/variance, which is load-bearing for the assertion that end-to-end training directly yields usable mixed-precision integer codecs without post-training retuning.","section":"§3"}],"minor_comments":[{"comment":"The novelty statement that JOMP is the first mixed-precision framework for neural video codecs would benefit from an explicit comparison table against prior mixed-precision methods applied to other neural codecs or vision models.","section":"Introduction"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive and detailed review. We address each major comment below and indicate the planned revisions.","responses":[{"response":"We acknowledge that the abstract presents the central claim at a high level without sufficient supporting details. In the revised manuscript, we will expand the abstract to include brief references to the experimental setup (including the DCVC-FM baseline), the sections reporting variance across seeds, and the ablation studies on joint optimization. This will enable readers to more readily assess the stability and generalizability of the learnable bit-width assignments.","revision_made":"yes","referee_comment":"[Abstract] Abstract: the central claim of RD performance comparable to DCVC-FM with an 87.6% bit-operation reduction is presented without experimental details, baselines, variance across seeds, or ablation results on the joint optimization; this directly prevents assessment of whether the learnable bit-width procedure produces stable and generalizable assignments as required by the weakest assumption."},{"response":"We agree that the manuscript lacks an explicit description of the relaxation technique and related training details for the bit-width variables. In the revised version, we will expand the training procedure section to describe the specific relaxation method used, any auxiliary losses, and the monitoring procedures employed to detect collapse or excessive variance. This addition will strengthen the support for the end-to-end training claim.","revision_made":"yes","referee_comment":"[§3] Training procedure (assumed §3): no description is given of the relaxation technique for bit-width variables (e.g., straight-through estimator or Gumbel-softmax), auxiliary losses, or monitoring for collapse/variance, which is load-bearing for the assertion that end-to-end training directly yields usable mixed-precision integer codecs without post-training retuning."}],"tokens_in":1442,"tokens_out":389,"duration_ms":14615,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The paper's core move is to make both quantization parameters and per-module bit widths trainable end-to-end inside variational neural video codecs. They apply this across several coding frameworks and buffering strategies, add an integerization pipeline for deterministic decoding, and report that the best model matches DCVC-FM rate-distortion while cutting bit operations by 87.6%. That is the concrete claim worth checking.\n\nWhat stands out is the attempt to treat precision assignment as part of the joint optimization rather than a post-hoc step. The systematic sweep over frameworks and temporal buffers is also new; most prior quantization work on neural codecs has been narrower. If the learned widths really transfer without heavy retuning, that would be useful for people trying to ship these models.\n\nThe main weakness is that none of the supporting evidence appears in the abstract. There are no reported training curves for the bit-width variables, no mention of straight-through estimator behavior or collapse, no ablation on the joint loss, no variance across seeds, and no comparison against standard mixed-precision baselines. The stress-test concern about stability and generalization therefore lands: learnable bit widths often need auxiliary tricks to stay usable, and without those details it is impossible to tell whether the reported numbers required per-framework retuning or post-training fixes.\n\nThe work is aimed at researchers who already build neural video codecs and now face deployment constraints. A reader who needs integer-only implementations might pick up the overall pipeline idea, but anyone wanting to reproduce or extend the results will have to wait for the full experimental section.\n\nIt is worth sending to referees. The topic is practical and the framing is clear enough that a review can test the central claim directly.","headline":"JOMP treats bit widths as learnable variables for mixed-precision integer neural video codecs and claims 87% bit-op cuts with DCVC-FM level RD, but the abstract supplies no training details, ablations, or stability checks.","tokens_in":2412,"tokens_out":429,"would_cite":false,"duration_ms":12667,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"JOMP makes bit widths learnable variables so neural video codecs can train end-to-end in mixed-precision integer arithmetic.","keywords":["mixed-precision quantization","neural video coding","integer neural codecs","rate-distortion-complexity trade-off","variational autoencoder","temporal buffering strategies","end-to-end optimization","deterministic decoding"],"falsifier":"A controlled test in which the learned bit-width assignments from JOMP require extensive per-framework retraining or post-processing to reach the reported rate-distortion performance.","tokens_in":2732,"feed_emoji":"🔢","tokens_out":667,"duration_ms":13784,"temperature":0.7,"pith_summary":"The paper introduces JOMP to solve the gap between high-performing floating-point neural video codecs and practical integer deployments. It treats quantization parameters and bit widths as jointly learnable during training, so modules inside a codec can run at different precisions while the rate-distortion-complexity trade-off is optimized directly. Experiments apply the method across multiple coding frameworks and temporal buffering strategies, and include a full integerization pipeline that produces deterministic decoding. When used on the strongest model, the resulting integer codec matches the rate-distortion performance of the floating-point state-of-the-art DCVC-FM while cutting bit operations by 87.6 percent.","feed_headline":"Mixed-precision learning cuts neural video codec bit ops 87.6%","feed_subtitle":"JOMP trains bit widths as variables so integer codecs match floating-point rate-distortion performance across frameworks.","key_machinery":"The JOMP framework, in which quantization parameters and bit widths are optimized jointly as learnable variables during training.","core_discovery":"By treating both quantization parameters and bit widths as learnable variables, JOMP performs end-to-end mixed-precision optimization for neural video codecs. This produces integer implementations whose rate-distortion performance is comparable to DCVC-FM while reducing bit operations by 87.6 percent. The same framework also supplies a complete integerization pipeline that guarantees deterministic decoding.","pith_inferences":["The same joint-optimization idea could be tested on other neural compression domains such as image or point-cloud coding.","Hardware designers could use the learned precision maps to allocate specialized low-precision arithmetic units inside video codecs.","The method may reduce the need for separate quantization-aware training pipelines when new buffering strategies are introduced.","If the learned bit widths prove stable across datasets, future codec standards could publish precision maps instead of full floating-point weights."],"forward_implications":["Different codec modules can run at different precision levels while the overall rate-distortion-complexity optimum is found automatically.","A single training procedure works across multiple neural video coding frameworks and temporal buffering strategies.","Integer neural video codecs become feasible with deterministic decoding and no floating-point arithmetic at inference time.","Bit-operation count can be reduced by 87.6 percent while rate-distortion performance stays comparable to the strongest floating-point baseline."],"fun_headline_variants":["JOMP jointly optimizes quantization and bit widths in video codecs","Integer video codecs match DCVC-FM with 87.6% fewer bit ops","End-to-end mixed precision for deterministic integer neural video coding","JOMP learns bit widths to reduce bit operations 87.6% in neural codecs"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"Making bit widths learnable during training produces stable mixed-precision assignments that generalize across frameworks without post-training retuning.","fun_headline_variants_meta":{"raw":{"variants":["JOMP jointly optimizes quantization and bit widths in video codecs","Integer video codecs match DCVC-FM with 87.6% fewer bit ops","End-to-end mixed precision for deterministic integer neural video coding","JOMP learns bit widths to reduce bit operations 87.6% in neural codecs"]},"model":"grok-4.3","cost_usd":0.006215,"raw_usage":{"total_tokens":2963,"prompt_tokens":739,"num_sources_used":0,"completion_tokens":68,"cost_in_usd_ticks":62149500,"prompt_tokens_details":{"text_tokens":739,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":2156,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":739,"tokens_out":68,"duration_ms":16840,"temperature":1.0,"reasoning_tokens":2156,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-27T05:43:59.356965+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A controlled test in which the learned bit-width assignments from JOMP require extensive per-framework retraining or post-processing to reach the reported rate-distortion performance.","supporting_citations":[],"review_version":1}