{"id":"a0be3ab3-38ca-4030-9aa6-3feeaa01e2f1","arxiv_id":"2608.08910","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Tying the two scales of a ternary weight decomposition at ratio 3 yields a uniform nine-level quantizer whose folded 4-bit format matches a 4.5-bit baseline on measured fidelity while decoding faster.","lead":"This paper ties the two per-group scales of a ternary weight decomposition to a fixed ratio, turning it into one uniform nine-level quantizer that stores each weight in about four bits. The authors stream the experts of a 284B-parameter model from a laptop SSD and report no detected quality difference versus a standard 4.5-bit format, with faster decoding and smaller files.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Fidelity parity rests on two protocol-sensitive cells; the q4_k depth cell swung 4/4 to 2/4 under the canonical protocol with unexplained root cause, so the 12/14-vs-11/14 margin is not stable evidence for replacement.","rationale":"The paper is unusually candid: the tie identity is not claimed as novel, the persistent-bytes invariant is the narrow novel conjunction, and every unstable cell is disclosed. The math in §2.1 is sound, and the mxfp4 anchor arm is a real methodological improvement. The load-bearing weakness is not the novelty or the math; it is the empirical parity evidence. The central deployment claim—tq20fx4 can replace q4_k without detected fidelity loss—rests on two margins (5/5 vs 4/5 and 12/14 vs 11/14) that are only cells wide. The step-0 margin is the documented knife-edge short; the depth margin is affected by a second documented condition-sensitive cell in the q4_k arm, whose 4/4-to-2/4 swing has an unexplained root cause. Because the paper itself says no formal equivalence margin was prespecified and the evaluation is five prompts with fourteen dependent steps, the evidence supports 'no detected difference' only in the weak, non-rejection sense. That is an honest claim, but it is not enough to support a replacement recommendation without resolving the protocol-dependent cell or enlarging the fixture set. The reader identified the knife-edge and small-set issue; this review extends it to the second q4_k depth cell, hence partial agreement. The verdict stays CONDITIONAL: conditional on a protocol-stress replication and/or a larger pre-registered fixture set. No change to the reader's verdict is needed.","tokens_in":13828,"tokens_out":8868,"duration_ms":85965,"concrete_test":"Run a protocol-stress matrix on the two disclosed condition-sensitive fixtures (short code completion, long memory archive) for both q4_k and tq20fx4: 50 fresh OS processes per arm under the canonical one-process-per-fixture protocol, 50 under the earlier full-process protocol, and 50 on a recompile with a deliberate schedule or ulp perturbation (e.g., different -O ordering or vectorization), recording step-0 and matched depth each time. If q4_k's long-memory depth distribution contains a 4/4 mode in any protocol condition while tq20fx4's does not, the canonical 11/14-vs-12/14 margin is protocol-dependent and the parity claim must be re-baselined on the stable 4/4 and 11/13 cells with an explicit equivalence bound; if all q4_k replicates are 2/4 and only the knife-edge short flips, the disclosed canonical numbers are reproducible and the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The core claim in §4.1 and Table 1—tied ternary shows no detected fidelity difference from q4_k at 4.06 bits—rests on step-0 5/5 vs 4/5 and matched continuation depth 12/14 vs 11/14, with every arm difference disclosed as the knife-edge short code-completion cell. But §6 and Appendix A disclose a second condition-sensitive cell: q4_k's long memory archive depth was 4/4 in an earlier full-process run on a pre-merge build and 2/4, deterministically, under the canonical one-process-per-fixture protocol; the paper states root causes are unexplained. If the earlier q4_k value is the more faithful measurement, q4_k's matched depth becomes 13/14, above tied's 12/14, reversing the reported margin. The canonical protocol is a reasonable choice, but an unexplained two-step protocol swing in one arm means the depth comparison is not a stable measurement of arm quality. Separately, the step-0 margin is a single cell that flips with process history and ulp-level build changes; the paper's exclusion computation (identical 4/4 and 11/13 without it) is honest, but it means the headline numbers are not evidence of advantage, only of non-rejection. Together, the fidelity-parity claim is a non-rejection on a five-prompt, fourteen-step set with two unresolved condition-sensitive cells, not a demonstrated equivalence; the deployment-replacement recommendation should be conditioned on resolving these cells.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper imposes a ratio-3 scale tie on PTQTP's two ternary planes, collapsing the decomposition into a uniform nine-level quantizer whose two trit planes fold losslessly into a single 4-bit code plane that serves as the persistent representation across disk, cache, and kernel. The construction is applied to the routed experts of DeepSeek-V4-Flash-0731, quantized one-shot from the released MXFP4 weights, and evaluated against a q4_k baseline and an expert-lossless anchor arm using official-API fixtures, a 100-item MMLU subset, WikiText-2 perplexity, and decode throughput on two machines. The headline result is that the tied quantizer shows no detected fidelity difference from q4_k at 4.06 bits/weight, with smaller files and faster decode, with every fixture-level difference traced to a single measured near-tie cell. The paper is unusually candid, explicitly labeling the result a non-rejection at small evaluation sizes and disclosing two protocol-sensitive cells whose root causes remain unexplained.","tokens_in":14129,"tokens_out":4939,"duration_ms":51362,"significance":"If the empirical claims hold, the contribution is practically significant: it demonstrates that a uniform nine-level quantizer can match a 4.5-bit K-quant baseline in behavioral fidelity at 4.0625 bits/weight, while the persistent folded format simplifies SSD-streamed MoE serving by keeping disk bytes, cache bytes, and kernel input identical. The paper also contributes a measured dissociation between proxy metrics (reconstruction error, perplexity) and reference fidelity, and a cumulative trunk-ternarization ladder. The manuscript is exceptionally transparent: it releases code and artifacts, pins the two ISA kernel arms bitwise-identical, uses an expert-lossless anchor arm, and discloses the exact cells that are unstable or within noise. The core tie identity is elementary arithmetic, so the contribution is in the application, measurement, and serving-system design rather than in a new mathematical result. The significance is conditional: the central empirical claim currently rests on a small, non-random fixture set with two unresolved condition-sensitive cells, and the speed claim lacks formal uncertainty and causal attribution.","major_comments":[{"comment":"The matched-depth margin favoring tied ternary (12/14 vs. 11/14) is not stable evidence. Appendix A discloses that q4_k's long memory archive depth measured 4/4 in an earlier full-process run on a pre-merge build but 2/4 deterministically under the canonical one-process-per-fixture protocol, with root causes unexplained. If the earlier value is the more faithful measurement, q4_k's depth becomes 13/14, which exceeds tied ternary's 12/14 and reverses the direction of the only continuation-depth difference. Since the abstract and §4.1 present the tied format as a practical replacement for q4_k on this deployment, this unresolved protocol swing is load-bearing. Please report both protocol values for all affected arms, add a sensitivity table that excludes both condition-sensitive cells, and either resolve the root cause or explicitly condition the replacement recommendation on the canonical protocol.","section":"§4.1/Table 1, Appendix A/Table 4"},{"comment":"The claim of 'no detected fidelity difference' is a non-rejection, not a demonstration of equivalence. The step-0 margin is a single knife-edge cell that flips with process history and ulp-level build changes; excluding it, the two arms are identical (4/4 at step 0, 11/13 on matched depth). With n=5 prompts, 14 dependent continuation steps, an MMLU comparison whose paired test gives p=0.6875, no prespecified equivalence margin, and a hosted reference API that can change over time, the evidence does not establish that the tied quantizer can replace q4_k without fidelity loss. The paper's wording in §6 is appropriately cautious, but the abstract and §4.1 should either state a prespecified equivalence margin with a power analysis or explicitly downgrade the headline to a descriptive non-rejection and remove any replacement implication.","section":"§4.1 and §6"},{"comment":"The +6.7% decode-throughput advantage is headline material but lacks formal uncertainty and causal attribution. The paper itself notes that because the formats generate different continuations, their routed-expert workloads differ, and the speed comparison does not causally separate format, kernel, cache, and workload effects. The byte-read and pin-count decomposition is a strong mechanical check and is invariant across all rounds, but the throughput headline rests on one prompt at 32 generated tokens, six gated rounds, and includes an unresolved q4_k round at 1.95 tok/s with degraded expert-path timing. Please report confidence intervals or bootstrap uncertainty for the throughput comparison, and ideally add a workload-matched control (for example, forced decoding of identical continuation tokens) so the format-level claim is not confounded by the different continuations.","section":"§4.4/Table 3"},{"comment":"The free-scale solve-log distribution is not released, and the paper's inversion finding (free scales improve perplexity while behavioral evidence is inconclusive) depends on the quality of those free-scale fits. The manuscript flags this gap, but because the inversion is a central empirical observation about the tie's cost, the absence of the solve logs and the per-expert error distribution prevents an independent check of whether the free-fit perplexity advantage is a solver artifact or a genuine property of the constraint. Please release the solve-log distribution or a representative subset, and state explicitly which of the inversion claims can be verified from the released artifacts.","section":"§2.1 and §4.2"}],"minor_comments":[{"comment":"The abstract contains a typographical error: 'uniformnine-level' should read 'uniform nine-level'; also, 'q4 k' is inconsistently spaced throughout the manuscript and should be normalized to 'q4_k'.","section":"Abstract"},{"comment":"The quiet-substrate acceptance gate (first attempt whose per-arm three-round spread is ≤5%) is a selection procedure that could bias the reported speed comparison; the paper states the rule was fixed in advance but not publicly registered. Please consider pre-registering the rule in the repository and reporting the number of rejected attempts and their values.","section":"§3, Speed protocol"},{"comment":"The caption notes that R3 was measured under the earlier protocol, but the table cell for the long memory archive (0/4) is particularly wide; it would help to mark R3's row visually with a footnote symbol in the table body rather than only in the caption.","section":"Table 2, R3 row"},{"comment":"The cross-ISA perplexity drift (4.44 vs. 4.39 at 512 tokens) is reported as a 1.1% relative difference; restating it as an absolute NLL-space difference of 0.0113 nats/token in the main text would make the magnitude clearer and align with the paper's own guidance to state NLL-space figures when comparing small differences.","section":"§4.5"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is unusually transparent, to the point of disclosing the exact measurements that undermine its headline. The core problem is that the two condition-sensitive cells identified in Appendix A are precisely the cells that carry the headline margins: the q4_k long-fixture depth swing and the knife-edge short-fixture step-0 flip. This is not a circularity or fabrication concern; it is an evaluation-stability concern. I do not recommend rejection, because the central derivation is sound, the artifacts are public, and the framing can be corrected. I would ask the authors to resolve or fully bound the q4_k depth-cell discrepancy, present a sensitivity analysis excluding both unstable cells, soften the replacement implication to match the non-rejection character of the evidence, and report the speed comparison with formal uncertainty or an explicit descriptive-only label. The novelty disclosure appears fair: the prior instances of the ratio-3 identity and of folded four-bit codes are acknowledged, and the narrow conjunction claim is stated carefully."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The construction is real and the writeup is refreshingly honest. Tying PTQTP's two scales at ratio three to get a uniform nine-level quantizer is a known identity, and the paper says so. What is new is the conjunction: imposing the tie inside the solver and then making the folded 4-bit code the persistent serving bytes, identical on disk, in the cache slab, and at the kernel input, with a CPU-SIMD SSD-streaming MoE implementation. That is a useful engineering result, and the open-source artifacts and the expert-lossless anchor arm make it reproducible. The paper earns real credit for disclosing exactly where its own margins collapse.\n\nThe soft spots are empirical, not mathematical. The entire fidelity comparison reduces to one knife-edge code-completion fixture: excluding it, tied and q4_k are identical on step-0 and depth. That is a non-rejection, and the paper mostly says so. But there is a second unstable cell that cuts worse: q4_k's long memory archive depth was 4/4 in an earlier full-process run and 2/4 deterministically under the canonical protocol, with root causes unexplained. If the earlier value is the more faithful one, q4_k's depth becomes 13/14, above tied's 12/14, reversing the reported edge. The canonical protocol is defensible, but an unexplained two-step swing in one arm means the depth comparison is not a stable measurement of arm quality. The speed result also lacks formal uncertainty and a causal decomposition, and the two formats generate different continuations, so workload differs. MMLU is a 100-item descriptive comparison, within noise. None of this is hidden, but the deployment-replacement recommendation outruns the evidence.\n\nOn the other side, the math is arithmetic and sound, the free scale per block is the only parameter, and the evaluation is anchored to the official API rather than to the paper's own outputs. There is no circularity. The paper is a clear-thinking, honest report that simply overreaches slightly on one empirical claim.\n\nThis deserves a serious referee. The construction, the persistent folded format, and the candid methodology are valuable, and peer review could push for a larger preregistered fixture set, stable-arm uncertainty, and a causal speed decomposition. I would not desk-reject it, and I would probably cite the format even while staying skeptical of the parity claim.","headline":"Honest, open, and mechanically sound, but the fidelity-parity headline is a well-disclosed non-rejection on five prompts with two unstable cells, so treat it as a promising systems result rather than proof that tied ternary replaces q4_k.","tokens_in":14656,"tokens_out":1753,"would_cite":true,"duration_ms":19686,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Tying PTQTP's two ternary scales at ratio 3 yields a uniform nine-level quantizer that shows no detected fidelity difference from a 4.5-bit baseline in a disk-streamed mixture-of-experts deployment.","keywords":["ternary quantization","nine-level quantizer","post-training quantization","mixture-of-experts","persistent serving format","disk streaming","reference fidelity","PTQTP"],"falsifier":"Take a larger set of, say, 500 prompts with greedy continuations captured from the official serving API, run both quantizers in one fresh process per prompt, and count step-0 agreement and continuation prefix depth against a prespecified equivalence margin; if the tied format's agreement rate falls below the conventional 4.5-bit baseline by more than that margin, the paper's no-detected-difference claim does not hold beyond its small evaluation set.","tokens_in":13620,"feed_emoji":"⚙️","tokens_out":13080,"duration_ms":117684,"temperature":0.7,"pith_summary":"PTQTP decomposes each weight group into two ternary planes with two free scales. This paper constrains those scales to a fixed ratio of three, which collapses the decomposition into a single uniform nine-level quantizer at 4.0625 bits per weight, and it makes that folded code the persistent serving representation for disk-streamed mixture-of-experts inference: disk bytes, cache bytes, and kernel input are identical, consumed in one integer dot pass. Applied to the routed experts of a 284B-parameter MoE, the tied quantizer matched the official serving API on 5/5 step-0 fixtures and 12/14 continuation steps, versus 4/5 and 11/14 for a conventional 4.5-bit quantizer, scored 86 vs 84 on a 100-item MMLU subset, decoded 6.7% faster, and shipped 9% smaller files. Because every fixture-level difference traces to one measured near-tie cell, the paper claims no detected fidelity difference, not superiority. The same measurements show higher weight-reconstruction error and worse perplexity for the tied fit, a dissociation the paper documents between proxy metrics and behavioral reference fidelity.","feed_headline":"Nine-level tied quantizer matches a 4.5-bit baseline in deployment","feed_subtitle":"Streams experts from disk with no detected fidelity loss, 7% faster decode, 9% smaller files.","key_machinery":"The central mechanism is the ratio-3 scale tie $\\alpha=(3s,s)$ inside PTQTP's alternating solver. With this constraint the composite code $c=3t_1+t_2$ becomes a uniformly spaced nine-level grid $\\{-4,\\ldots,4\\}$ with one scale $s$, equivalent to fitting a uniform nine-level quantizer directly. The two ternary planes then fold losslessly into one 4-bit code (two codes per byte, four f16 column scales per 520-byte block), and because the folded code is the persistent representation, expert misses are single contiguous reads, cache tiers can use file-identical bytes, and the kernel does one integer dot pass with arithmetic decoding on NEON and AVX2/AVX-VNNI that is pinned bitwise-identical across ISAs.","core_discovery":"On the paper's own terms, the discovery is that the tie $\\alpha_1 = 3s,\\ \\alpha_2 = s$ turns PTQTP's free two-scale ternary decomposition into a single uniform nine-level quantizer $\\hat{W}=s\\,c$ with $c=3t_1+t_2\\in\\{-4,\\ldots,4\\}$, and that this code can be folded into a persistent 4-bit plane (4.0625 bits/weight) that is the exact bytes on disk, in the expert-cache slab, and at the integer dot-product kernel's input. In the measured deployment, that format matched the official serving API on 5/5 step-0 fixtures and 12/14 continuation steps, versus 4/5 and 11/14 for the conventional 4.5-bit baseline, scored 86 versus 84 on a 100-item MMLU subset, decoded 6.7% faster, and shipped 9% smaller files; all fixture-level differences between the arms collapse to a single knife-edge cell, so the claim is no detected fidelity difference at these evaluation sizes, not superiority. The tied fit has higher weight-reconstruction error and worse WikiText-2 perplexity than the baseline, which the paper reads as evidence that proxy objectives and reference fidelity can disagree.","pith_inferences":["Editorial inference: if the no-detected-difference result extends to larger, stratified reference captures, the ratio-3 tie could become a default constraint for post-training quantization of MoE experts, not a special case.","Editorial inference: the paper's finding that perplexity can rank arms opposite to reference agreement suggests that near-baseline quantizer comparisons should report small behavioral fixtures and task agreement alongside perplexity; this could be tested by applying the same protocol to other models and quantizer families.","Editorial inference: the knife-edge fixture's flip under process history and tiny binary changes implies that some argmax comparisons on tiny prompt sets are measurement noise rather than quality signals; a natural extension is to repeat near-tie fixtures many times in isolated processes and report flip probabilities."],"forward_implications":["The tied nine-level code can replace a conventional 4.5-bit quantizer for routed MoE experts without detected behavioral change on the tested fixtures, at a lower bit rate.","Persistent folded bytes make an expert miss a single contiguous read, and cached bytes are identical to disk bytes, so outputs are independent of where an expert was cached.","The measured dissociation between perplexity or reconstruction error and reference fidelity implies that a quantizer can look worse on proxy metrics yet match reference behavior on tested prompts, so deployment choices need behavioral anchoring.","The cumulative ternarization ladder shows that read-side attention projections can be ternarized without reducing fixture agreement, localizing full-model sensitivity to other components.","Nine percent smaller expert files at equal fidelity translate to less disk traffic and more resident experts under a fixed RAM budget."],"supporting_citations":[{"why":"Introduces the dual trit-plane decomposition and alternating solver that the tie constrains.","marker":"[1]"},{"why":"Supplies the official-API captured fixtures, conversion lineage, and quality-harness methodology the behavioral comparison rests on.","marker":"[26]"},{"why":"Identifies the released model artifact whose routed-expert weights are quantized and served.","marker":"[27]"},{"why":"Defines the conventional 4.5-bit quantizer and file-format ecosystem used as the comparison baseline.","marker":"[28]"},{"why":"Specifies the MXFP4 expert weight format that the quantizer reads and exactly dequantizes.","marker":"[32]"}],"fun_headline_variants":["Tied trit planes fold to one 9-level code: 4-bit disk format, 7% faster","Uniform 9-level quantizer from tied trits: matches 4.5-bit baseline, 9% smaller","Tie collapses two trit planes into one 9-level code, matches API on 5/5 fixtures","Nine-level tie from PTQTP: disk-streamed experts, no detected fidelity gap","Nine-level tie: higher reconstruction error, yet matches API"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the few captured greedy responses from the official serving API are a stable and meaningful reference, and that the one fixture whose first token flips under process history and tiny binary differences is measurement noise rather than a real difference between the two quantizers.","fun_headline_variants_meta":{"raw":{"variants":["Tied trit planes fold to one 9-level code: 4-bit disk format, 7% faster","Uniform 9-level quantizer from tied trits: matches 4.5-bit baseline, 9% smaller","Tie collapses two trit planes into one 9-level code, matches API on 5/5 fixtures","Nine-level tie from PTQTP: disk-streamed experts, no detected fidelity gap","Nine-level tie: higher reconstruction error, yet matches API"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001472,"raw_usage":{"total_tokens":6067,"prompt_tokens":1240,"completion_tokens":4827,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":856,"completion_tokens_details":{"reasoning_tokens":4705}},"tokens_in":856,"tokens_out":4827,"duration_ms":33601,"temperature":1.0,"reasoning_tokens":4705,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T04:20:23.128251+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a larger set of, say, 500 prompts with greedy continuations captured from the official serving API, run both quantizers in one fresh process per prompt, and count step-0 agreement and continuation prefix depth against a prespecified equivalence margin; if the tied format's agreement rate falls below the conventional 4.5-bit baseline by more than that margin, the paper's no-detected-difference claim does not hold beyond its small evaluation set.","supporting_citations":[{"cited_title":"Sanfilippo","cited_arxiv_id":null,"evidence_quote":"Supplies the official-API captured fixtures, conversion lineage, and quality-harness methodology the behavioral comparison rests on."},{"cited_title":"DeepSeek-V4-Flash-0731 (revision 7872f01b)","cited_arxiv_id":null,"evidence_quote":"Identifies the released model artifact whose routed-expert weights are quantized and served."},{"cited_title":"Gerganov and contributors","cited_arxiv_id":null,"evidence_quote":"Defines the conventional 4.5-bit quantizer and file-format ecosystem used as the comparison baseline."},{"cited_title":"OCP Microscaling Formats (MX) Specification, Version 1.0","cited_arxiv_id":null,"evidence_quote":"Specifies the MXFP4 expert weight format that the quantizer reads and exactly dequantizes."}],"review_version":1}