{"id":"3acc1719-006c-4221-abaa-447fb5de2820","arxiv_id":"2607.04371","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.5,"correctness_risk":"low","formal_verification":"none","parameter_count":6,"one_line_summary":"Iterative Puzzle plus KD, RL, quantization, and MTP compresses Nemotron-3-Super to 75B total / 9B active parameters with ~2× interactive throughput and 8× 1M-context concurrency while retaining most parent accuracy.","lead":"NVIDIA compresses its 120B hybrid MoE Nemotron-3-Super into a 75B/9B-active model that roughly doubles interactive server throughput and raises 1M-token concurrency from 1 to 8 on one H100. The result is a practical recipe for making frontier hybrid MoE models cheaper to serve without collapsing accuracy.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection identified to the central throughput/accuracy claim; residual risk is external generality, not internal inconsistency.","rationale":"I agree with CONDITIONAL and HIGH confidence on the reported Super-matched results. The reader correctly flags Iterative Puzzle’s additivity assumption and incomplete appendix items, but that assumption is not required for the strongest claim as stated (measured ~2× server throughput and 1→8 concurrency with retained accuracy). The claim is systems-empirical, not a proof of search optimality. Table 6 already shows iterative beats single-shot modestly; Fig. 2 documents non-uniform capacity; Table 7 and Fig. 6 give the Pareto numbers under matched quantization. The main remaining risk is stack- and traffic-specificity (production mix, independent kernels, external baselines), which the reader already notes—hence CONDITIONAL remains appropriate without moving to REJECT or ACCEPT. No internal inconsistency or measurement contradiction is evident in the manuscript.","tokens_in":24588,"tokens_out":567,"duration_ms":6352,"concrete_test":"Re-run the 8×B200 8K/64K Pareto sweep of §3.3 at UT=100 with the public BF16/NVFP4 checkpoints under an independent serving stack (e.g., vLLM or TensorRT-LLM with the same TP/EP/batch grid), and recompute suite-average accuracy on the Table 3 set; if TPS boost falls below ~1.6× or average accuracy drops >2 points vs Super under matched quantization, the headline claim weakens outside the authors’ stack.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper’s strongest claim is empirical and well-supported inside its own stack: matched-NVFP4 Pareto throughput on 8×B200 (Table 7, Fig. 6) and 1M-context concurrency on H100, with suite accuracy retained relative to Super (Table 3). The reader’s weakest assumption—that Iterative Puzzle’s local scores become sufficiently additive after intermediate KD—is a real methodological soft spot for claiming near-optimality of the heterogeneous allocation, but it is not load-bearing for the reported ~2× result. That result only requires that the final architecture (Table 1, Fig. 2) plus recovery actually delivers the measured TPS/UT and accuracy, which the ablations (Table 6: +0.57 vs single-shot) and training progression (Fig. 4) support. Residual higher-order interactions would at most leave headroom for a better architecture; they do not falsify the measured gains. Incomplete appendix MTP TODOs and limited external baselines affect how far the claim generalizes, not whether the reported Super→Puzzle comparison holds.","agreement_with_reader":"partial"},"referee_report":{"model":"grok-4.5","summary":"The manuscript presents Nemotron-Labs-3-Puzzle-75B-A9B, a deployment-oriented compression of Nemotron-3-Super from 120.7B total / 12.8B active parameters to 75.3B total / 9.3B active. Compression uses Iterative Puzzle (sequential hardware-aware NAS over heterogeneous MoE intermediate size and top-k, plus uniform Mamba SSM state pruning 128→96), interleaved with short KD, then longer-context KD, SWE-focused RL, PTQ (FP8/NVFP4), and a transferred/continued shared MTP head. On a single 8×B200 node at matched NVFP4, the model reports roughly 1.6–2.14× Pareto-optimal total throughput versus Super at fixed user-throughput thresholds (UT≥100/125/150) in 50K/2K and 8K/64K regimes (Table 7, Fig. 6); on a single H100 at 1M context it raises concurrency from 1 to 8. Suite accuracy is largely retained relative to Super across reasoning, coding, long-context, multilingual, and agentic benchmarks (Table 3), with supporting ablations on iterative vs single-shot Puzzle (Table 6), recovery progression (Fig. 4), MTP acceptance (Table 5), and disaggregated prefill (Fig. 5).","tokens_in":25078,"tokens_out":1488,"duration_ms":30209,"significance":"If the reported Super→Puzzle comparisons hold under independent reproduction, this is a strong systems contribution: it shows that hybrid Mamba–Attention–MoE models can be non-uniformly compressed for interactive serving and ultra-long-context memory limits while keeping most parent capability. Strengths include matched-quantization Pareto sweeps over TP/EP/batch size, explicit UT-constrained serving metrics (not only peak TPS), public Hugging Face checkpoints, and concrete ablations (iterative Puzzle +0.57 avg, training-stage recovery, MTP acceptance lengths, prefill-only disaggregation). The heterogeneous active MoE capacity map (Fig. 2, Table 1) and contribution-based SSM channel selection (Appendix A) are useful methodological artifacts for the community. Significance is primarily empirical and stack-specific rather than a new general theory of MoE compression.","major_comments":[{"comment":"Appendix C still contains an explicit unfinished placeholder: “TODO: drop in the headline MTP boost numbers (Super-Turbo-MTP-best vs. Super-Turbo-no-MTP at UT≥100, both scenarios) from the latest report.” The abstract and §3.1/Table 4 advertise large MTP-compounded gains, but the main interactive Pareto analysis in §3.3/Table 7 is single-step only, and the MTP overlay (Fig. 8) is incomplete. For the deployment claim to be fully load-bearing, the manuscript needs finalized, matched-quantization Pareto numbers for Puzzle+MTP vs Super+MTP at the same UT thresholds used in Table 7, with draft length and acceptance-length assumptions stated.","section":null},{"comment":"External baselines beyond the Nemotron Super/Nano family are essentially absent from the accuracy–efficiency comparison (Fig. 1, Table 4). The central Super→Puzzle 2× claim is internally well supported, but the broader conclusion that “large hybrid MoE models can be substantially optimized for deployment efficiency while maintaining strong downstream capability” would be much stronger with at least one independent open model of similar active-parameter or throughput class (or a published compressed MoE baseline) under the same UT-constrained Pareto protocol. Without that, generality remains an open risk even if the parent comparison is correct.","section":null},{"comment":"§2.1.2 and Table 6: Iterative Puzzle is presented as the core methodological advance, yet the only architecture-search ablation is a +0.57 unweighted average over a modest benchmark subset versus single-shot Puzzle at the same final budget. That gain supports preferring the iterative procedure, but does not establish near-optimality of the heterogeneous ρ_l allocation (Fig. 2) under residual higher-order layer interactions. A load-bearing clarification is needed: either (i) state clearly that the paper claims measured efficiency of the final architecture, not global optimality of the search, or (ii) add a stronger control (e.g., uniform capacity at matched active params, or a second iterative path with different stage budgets) so readers can separate search quality from recovery quality.","section":null}],"minor_comments":[{"comment":"Abstract “approximately 2×” should be qualified by regime: Table 7 shows ~1.60–1.79× on prefill-heavy 50K/2K and ~2.03–2.14× on decode-heavy 8K/64K at the listed UT points.","section":null},{"comment":"Table 4 “TPS×Super” and “Rel. Req./min” columns are easy to misread against the Super MTP row (3.04× vs text mentioning 3.57× elsewhere). Align all Super+MTP multipliers with one consistent measurement protocol.","section":null},{"comment":"Fig. 6 legend uses “Turbo Pareto” while the model is named Puzzle-75B-A9B throughout; unify naming to avoid implying a different product.","section":null},{"comment":"§2.2 RL: “the impact of RL training in our experiments was small” (Fig. 4) is important; consider moving a one-sentence statement into the main recovery narrative so readers do not over-attribute SWE recovery to RL.","section":null},{"comment":"Appendix A/B/D are valuable but dense; a short pointer in §2.1 to which pruning axes entered the final MIP (and which were explored then dropped, e.g. latent dimension) would improve navigability.","section":null},{"comment":"Typos/consistency: “Nemotron-3-Puzzle” vs “Nemotron-Labs-3-Puzzle”; “Super Turbo” vs “Puzzle-75B-A9B”; ensure Table 3/9 benchmark names and tool-use settings match the parent Super paper’s protocol citations.","section":null}],"recommendation":"minor_revision","confidential_remarks":"This is a high-quality industrial systems paper with public weights and careful serving methodology. Fit is good for a systems/ML systems venue; for a more theory-oriented journal the Iterative Puzzle novelty is incremental over Puzzle and the main value is the end-to-end hybrid-MoE deployment result. I would not block on external baselines alone if MTP appendix numbers are completed and the optimality language is toned to match the evidence. No integrity concerns from the text; residual risk is stack-specificity (self-teacher Super/Nano, proprietary serving stack) rather than internal inconsistency."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The punchline is practical, not theoretical: they ship a 75.3B/9.3B hybrid MoE (from Super 120.7B/12.8B) that, at matched NVFP4 on one 8×B200, roughly doubles Pareto server throughput at fixed user TPS (e.g. ~2.0× single-step on 8K/64K at UT≥100), and on one H100 lifts 1M-token concurrency from 1 to 8 by freeing weight memory while keeping the same attention layout. Suite accuracy stays close to Super across reasoning, coding, RULER, multilingual, and agentic sets (Table 3); NVFP4 tracks BF16 well.\n\nWhat is new is the concrete architecture and the measured serving result, not a new paradigm. Iterative Puzzle is a sequential wrap of their prior Puzzle NAS: moderate MoE intermediate/top-k and uniform Mamba SSM 128→96 pruning, short KD between rounds, then longer KD + SWE-focused RL, PTQ, and shared-head MTP. Heterogeneous routed capacity (Fig. 2) is the interesting design choice; the iterative-vs-1-step ablation (+0.57 avg, Table 6) and recovery curves (Fig. 4) support that the pipeline works. Pareto sweeps over TP/EP/batch at matched quantization (Table 7, Fig. 6), MTP acceptance lengths, and the disaggregated-prefill side study are the right kind of evidence for a systems claim. Models are on Hugging Face.\n\nSoft spots are real but secondary. The additivity assumption after intermediate KD is a soft spot for claiming the allocation is near-optimal—not for the reported Super→Puzzle gap, which only needs the final model to deliver the measured TPS and accuracy. RL recovery is small by their own account. External baselines beyond Super/Nano are thin, so generality outside their stack is open. Appendix still has an MTP TODO and incomplete headline numbers. Citation pattern is heavy on their own Super/Nano/Puzzle line, which is fine for an internal compression paper but limits the comparative frame.\n\nThis is for people who care about hybrid MoE serving economics and long-context concurrency, not for pure theory. Math is light (contribution scores, MIP search); data and ablations are the load-bearing part and look solid enough for referee time. I would engage: read the Pareto and architecture sections, try the HF checkpoints if you serve similar workloads. Send to peer review; fix TODOs and add uncertainty/external baselines in revision.","headline":"Solid systems paper: ~2× interactive throughput and 1→8 1M-context concurrency on one H100, with accuracy mostly held vs Super; Iterative Puzzle is incremental but the measurements are real.","tokens_in":26035,"tokens_out":634,"would_cite":true,"duration_ms":7967,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"A hybrid MoE parent model can be cut to 75B total / 9B active parameters and still serve about twice as many tokens per second under tight per-user latency targets.","keywords":["hybrid MoE compression","Iterative Puzzle","knowledge distillation","Mamba pruning","interactive serving","multi-token prediction","NVFP4 quantization","long-context concurrency"],"falsifier":"Run the same three-stage compression budget as a single-shot Puzzle search (no intermediate recovery) and check whether the final model still matches the reported suite-average accuracy and the ~2× throughput gains at fixed user throughput on the 8×B200 node; a large drop would falsify the value of the iterative loop.","tokens_in":25511,"feed_emoji":"⚡","tokens_out":656,"duration_ms":7288,"temperature":0.7,"pith_summary":"This paper argues that large hybrid mixture-of-experts language models can be made much cheaper to serve without giving up most of their reasoning, coding, long-context, and agentic skill. Starting from a 120B-total / 13B-active parent, the authors build a 75B-total / 9B-active student by repeatedly pruning MoE experts and Mamba state, healing each intermediate model with short knowledge distillation, then finishing with longer distillation, reinforcement learning, quantization, and multi-token prediction. On a single eight-GPU node the compressed model roughly doubles total server throughput at the same per-user token rate; on one GPU it raises the number of simultaneous million-token requests from one to eight. The claim is practical: production systems can keep near-parent quality while meeting interactive and ultra-long-context budgets that the parent cannot.","feed_headline":"Compressed hybrid MoE serves ~2× tokens at same user rate","feed_subtitle":"75B/9B active model keeps near-parent accuracy while lifting 1M-context concurrency from 1 to 8","key_machinery":"Iterative Puzzle: a sequential neural-architecture search that alternates moderate hardware-aware pruning of MoE and Mamba layers with short knowledge-distillation recovery so that replacement scores are recomputed on the current compressed model rather than only on the original parent.","core_discovery":"Heterogeneous, iterative structural compression of a hybrid MoE-Mamba model—jointly choosing per-layer expert width, number of active experts, and Mamba state size, then recovering with distillation and reinforcement learning—yields a student that delivers approximately 2× Pareto-optimal server throughput versus the parent at matched user-throughput targets and multiplies 1M-context concurrency eight-fold while retaining strong accuracy across a broad evaluation suite.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["Hybrid MoE compression doubles server throughput at matched user rate","Puzzle 75B/9B multiplies 1M-context concurrency from 1 to 8","Iterative MoE-Mamba pruning yields ~2× Pareto-optimal throughput","Joint expert and state compression keeps accuracy at 2× tokens","75B hybrid MoE compresses to 9B active for higher serving efficiency"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"The method assumes that after each short healing step the quality of individual layer replacements remains roughly additive, so the optimizer can still pick a near-optimal capacity layout for the final target.","fun_headline_variants_meta":{"raw":{"variants":["Hybrid MoE compression doubles server throughput at matched user rate","Puzzle 75B/9B multiplies 1M-context concurrency from 1 to 8","Iterative MoE-Mamba pruning yields ~2× Pareto-optimal throughput","Joint expert and state compression keeps accuracy at 2× tokens","75B hybrid MoE compresses to 9B active for higher serving efficiency"]},"model":"grok-4.5","effort":"low","cost_usd":0.00615,"raw_usage":{"total_tokens":1645,"prompt_tokens":834,"num_sources_used":0,"completion_tokens":104,"cost_in_usd_ticks":61500000,"prompt_tokens_details":{"text_tokens":834,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":707,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":834,"tokens_out":104,"duration_ms":6854,"temperature":1.0,"reasoning_tokens":707,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-11T19:40:34.724821+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Run the same three-stage compression budget as a single-shot Puzzle search (no intermediate recovery) and check whether the final model still matches the reported suite-average accuracy and the ~2× throughput gains at fixed user throughput on the 8×B200 node; a large drop would falsify the value of the iterative loop.","supporting_citations":[],"review_version":1}