{"id":"325ad0f4-ce66-42b9-a5a7-38949c814c76","arxiv_id":"2603.04956","paper_version":2,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"WaterSIC allocates per-column quantization rates via waterfilling and is proven within 0.255 bits of the IT optimum for linear-layer quantization, beating GPTQ and setting SOTA on Llama/Qwen.","lead":"WaterSIC quantizes neural-network linear layers by giving different bit budgets to different weight columns, matching classical waterfilling. It claims a uniform 0.255-bit gap to the information-theoretic optimum and new state-of-the-art LLM compression from 1–4 bits.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"Readable body is missing; the uniform 0.255-bit gap claim cannot be checked for finite-rate validity or exact distortion, so the central IT guarantee remains unverified.","rationale":"The reader correctly refused to ACCEPT/CONDITIONAL/REJECT because the CACHEABLE full manuscript is corrupted and unreadable; only the abstract is usable. That abstract states a coherent program (IT analysis, GPTQ’s arbitrary gap, WaterSIC’s uniform 0.255-bit waterfilling gap, SOTA on Llama/Qwen, public code), but the load-bearing steps—exact distortion, finite-rate vs high-rate status of the 0.255 constant, and transfer from second-order activation model to end-to-end quality—live in the body. My concern is the same weakest link the reader named, sharpened only by noting that 0.255 is the classic high-rate scalar gap, which raises a concrete finite-rate risk. No independent verification (proofs, tables, or clean source) is available in the provided materials, so the verdict must stay UNVERDICTED with low confidence. If a clean manuscript later shows a non-asymptotic uniform proof and matching layer-MSE experiments, the claim would become checkable; until then no upgrade is justified.","tokens_in":3205,"tokens_out":698,"duration_ms":24181,"concrete_test":"Obtain a clean PDF of 2603.04956; locate the theorem that states the 0.255-bit uniform gap; record whether its proof is high-rate/asymptotic or holds for finite rate for arbitrary PSD Σ. Then, for one real Llama/Qwen linear layer, compute WaterSIC’s empirical output MSE at 2 bits versus the IT lower bound from the measured activation covariance; if the implied excess rate exceeds ~0.3 bits on a badly conditioned Σ, the practical uniformity claim weakens.","verdict_should_be":"UNVERDICTED","load_bearing_attack":"The paper’s theoretical centerpiece is that WaterSIC stays within 0.255 bits of the IT rate–distortion limit uniformly over every input-activation covariance, while GPTQ can be arbitrarily far. That constant matches the classical high-rate scalar-vs-Gaussian gap (½ log₂(πe/6) ≈ 0.2546), so the argument almost certainly reduces layer output error to a quadratic form in the activation covariance Σ, then waterfills column rates and invokes a scalar-quantization excess. For the claim to underwrite both the theory and the 1–4-bit SOTA, three things must hold: (1) the distortion in the IT analysis is exactly the layer output discrepancy used in practice, (2) the 0.255 gap is non-asymptotic (or at least tight at 1–4 bits) for every PSD Σ, including rank-deficient or highly skewed ones, and (3) practical codebook/rounding choices do not open a larger gap. The supplied full-text payload is encoding garbage (and even carries a mismatched arXiv id), so none of the theorems, finite-blocklength assumptions, or tables can be inspected. Until those are readable, the uniform finite-rate guarantee and its link to end-to-end LLM quality are uncheckable load-bearing steps.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper studies post-training conversion of a dense linear layer to low precision under an information-theoretic (rate–distortion) lens on the tradeoff between compressed length and layer output discrepancy. It claims that GPTQ can be arbitrarily far from the IT optimum, and introduces WaterSIC, which allocates different quantization rates to different weight columns (in-features) in a waterfilling style and is asserted to stay within a uniform 0.255-bit rate gap of the IT limit for every input-activation covariance. Empirically, WaterSIC is reported to set new SOTA on Llama and Qwen models for 1–4 bit quantization; code is linked.","tokens_in":3533,"tokens_out":896,"duration_ms":14884,"significance":"If the uniform finite-rate gap and the GPTQ separation are correctly proved, and if they translate into the reported end-to-end LLM gains, the work would be a meaningful bridge between classical rate–distortion/waterfilling and practical weight-only quantization. The explicit IT benchmark, the claimed uniformity over all covariances, and the public code are strengths that would make the contribution checkable and reusable. Those claims cannot currently be confirmed from the supplied manuscript body.","major_comments":[{"comment":"The supplied full manuscript body is encoding-corrupted (mojibake) and even carries a mismatched arXiv id/category fragment (2603.04958 / cs.CV). Theorems, proofs, finite-blocklength assumptions, and tables are unreadable. The central claims—an arbitrary GPTQ gap to the IT limit, a uniform 0.255-bit WaterSIC gap for every PSD covariance, and 1–4 bit SOTA—therefore cannot be verified from equations or experiments in this package.","section":null},{"comment":"Abstract claim of a uniform 0.255-bit rate gap: this constant matches the classical high-rate scalar-vs-Gaussian excess ½ log₂(πe/6) ≈ 0.2546. For the claim to underwrite both theory and 1–4 bit practice, the manuscript must show (i) the IT distortion is exactly the layer output discrepancy used in experiments, (ii) the gap is non-asymptotic (or tight at 1–4 bits) for every PSD Σ including rank-deficient/skewed cases, and (iii) practical codebooks/rounding do not open a larger gap. None of (i)–(iii) can be checked until the body is readable.","section":null},{"comment":"Abstract claim that GPTQ may have an arbitrarily large gap: this is load-bearing for motivating WaterSIC. Without the construction (e.g., a sequence of covariances and rates where GPTQ’s excess diverges), the separation remains an uncheckable assertion rather than a proved limitation of equal-rate column quantization.","section":null},{"comment":"Empirical SOTA on Llama/Qwen for 1–4 bits: end-to-end quality depends on calibration data, codebook design, and rounding. The abstract does not state the exact distortion or calibration protocol; until tables and ablations are readable, it is unclear whether the IT gap controls LLM metrics or whether unstated choices drive the reported gains.","section":null}],"minor_comments":[{"comment":"Abstract is clear and self-contained; the waterfilling intuition and code link are helpful once the body is restored.","section":null},{"comment":"When the PDF is fixed, please ensure the arXiv id, primary category, and theorem numbering match the abstract’s claims so the 0.255-bit constant and GPTQ separation can be cited precisely.","section":null}],"recommendation":"uncertain","confidential_remarks":"The review package is not usable for a technical assessment: the full-text blob is garbage and even points at a different arXiv entry. I cannot responsibly accept, revise, or reject on scientific grounds until a clean PDF is provided. Recommendation is therefore uncertain pending a readable manuscript; if the authors resubmit a clean version, the load-bearing items above (uniform finite-rate gap, GPTQ separation construction, distortion match, and SOTA tables) should be the first checks."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"Punchline: they cast dense-layer quantization as a rate–distortion problem under activation covariance, show GPTQ can sit arbitrarily far from the IT limit, and propose WaterSIC—per-column rate allocation that mimics waterfilling—with a claimed uniform 0.255-bit gap and reported SOTA on Llama/Qwen at 1–4 bits. Code is public.\n\nWhat is new is not waterfilling itself, but treating columns (in-features) as channels with different rates, proving a uniform gap over all input covariances, and the negative result on GPTQ. That 0.255 figure is exactly the classical high-rate scalar-vs-Gaussian excess (½ log₂(πe/6)), so the theory almost certainly reduces output error to a quadratic form in Σ, waterfills, and invokes a scalar-quantization gap. If the body makes that rigorous and non-asymptotic (or tight at 1–4 bits) for every PSD Σ, including rank-deficient ones, that is a real contribution to the compression literature, not just another quantizer recipe.\n\nSoft spots, in proportion: the full text we have is encoding garbage (wrong arXiv fragment even), so I cannot check theorems, finite-blocklength assumptions, the exact distortion, or the tables. The load-bearing steps—whether the IT distortion matches the practical layer discrepancy, whether the gap stays ~0.255 at low rates for every Σ, and whether codebook/rounding choices reopen a larger gap—are therefore unverified. Empirical SOTA also depends on calibration protocol we cannot inspect. Those are real caveats, not nitpicks; they are also the normal things a referee would demand.\n\nWho it is for: people who care about principled LLM weight quantization and about when GPTQ-style methods fail. A serious editor should send this to peer review; the program is coherent, the classical tools are used in the right place, and the claims are sharp enough to deserve checking. I would not cite it yet until the body is readable and the uniform gap is confirmed. Bring it to reading group only after we have a clean PDF—abstract alone is not enough to discuss the math.","headline":"Clean IT framing of layer quantization with a waterfilling-style fix for GPTQ’s gap; body is unreadable here so the 0.255-bit claim and SOTA stay unchecked.","tokens_in":4146,"tokens_out":548,"would_cite":false,"duration_ms":8779,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"WaterSIC compresses dense linear layers to within 0.255 bits of the information-theoretic limit by allocating different bit rates to different input features, and it sets new best results on large language models at 1–4 bits.","keywords":["linear layer quantization","information theory","waterfilling","rate-distortion","LLM compression","low-precision weights","WaterSIC","GPTQ"],"falsifier":"For a layer whose measured input covariance is highly ill-conditioned, compute the information-theoretic lower bound on rate for a target output mean-square error; if WaterSIC’s realized rate exceeds that bound by more than 0.255 bits, or if at matched total rate WaterSIC fails to improve layer output error and downstream perplexity relative to equal-rate GPTQ, the central claim is false.","tokens_in":4086,"feed_emoji":"💧","tokens_out":925,"duration_ms":17593,"temperature":0.7,"pith_summary":"Compressing a dense linear layer to few bits always trades storage against how much the layer’s outputs change. This paper studies that tradeoff with information theory and proves that a popular equal-treatment method (GPTQ) can sit arbitrarily far above the best possible rate for a given output error. WaterSIC nearly closes the gap: for every possible second-order model of the inputs, its rate is at most 0.255 bits above the information-theoretic minimum. It does so by giving more bits to important input features and fewer to weak ones, exactly as classical waterfilling would prescribe. On the Llama and Qwen model families the method improves accuracy at every rate from 1 to 4 bits per weight.","feed_headline":"WaterSIC quantizes layers within 0.255 bits of the IT limit","feed_subtitle":"Column-wise waterfilling beats GPTQ and sets SOTA on Llama and Qwen at 1–4 bits.","key_machinery":"WaterSIC: a column-wise rate allocation that mimics classical waterfilling on the input-activation covariance. Different in-features receive different bit budgets, which is the mechanism that keeps the rate gap to the information-theoretic limit bounded by 0.255 bits for every covariance.","core_discovery":"A fixed dense weight matrix can be quantized so that the excess rate above the information-theoretic minimum needed for any prescribed output discrepancy stays at most 0.255 bits, and this bound holds uniformly for every input covariance. The same analysis shows that GPTQ’s gap to that minimum can be made arbitrarily large. The algorithm that achieves the near-optimal rates is WaterSIC, which assigns unequal quantization rates to the columns of the weight matrix.","pith_inferences":["End-to-end gains should be largest in layers whose activation covariances are highly anisotropic, where equal-bit methods waste bits on low-energy directions.","Under the second-order model the remaining 0.255-bit gap leaves limited headroom for more elaborate vector codes, so further practical wins may come mainly from better finite-blocklength rounding rather than better rate allocation.","The same column-wise waterfilling idea could be applied to other linear maps (attention projections, convolutions) once an analogous output-discrepancy distortion is defined."],"forward_implications":["At any fixed bit budget between 1 and 4 bits, linear layers can be quantized with strictly smaller output discrepancy than prior popular methods.","Because the 0.255-bit gap is uniform, the same algorithm can be applied without redesigning rate allocation for each new covariance geometry.","Equal-bit or GPTQ-style schemes leave unused rate whenever input features have very unequal energy; waterfilling recovers that rate.","New state-of-the-art accuracy is obtained for the Llama and Qwen families across the entire 1–4 bit range."],"fun_headline_variants":["WaterSIC keeps layer quant within 0.255 bits of IT limit","Column waterfilling nears IT optimum for linear layers","WaterSIC bounds rate gap at 0.255 bits for any input cov","Per-column rates beat GPTQ gap that can grow arbitrarily","WaterSIC sets Llama Qwen SOTA at 1-4 bit layer quant"],"cache_read_input_tokens":128,"weakest_assumption_plain":"The right figure of merit is an information-theoretic rate–distortion tradeoff under a fixed second-order model of the activations, so a 0.255-bit gap in that model still controls practical end-to-end model quality after real codebooks and rounding.","fun_headline_variants_meta":{"raw":{"variants":["WaterSIC keeps layer quant within 0.255 bits of IT limit","Column waterfilling nears IT optimum for linear layers","WaterSIC bounds rate gap at 0.255 bits for any input cov","Per-column rates beat GPTQ gap that can grow arbitrarily","WaterSIC sets Llama Qwen SOTA at 1-4 bit layer quant"]},"model":"grok-4.5","effort":"low","cost_usd":0.00546,"raw_usage":{"total_tokens":1462,"prompt_tokens":732,"num_sources_used":0,"completion_tokens":95,"cost_in_usd_ticks":54600000,"prompt_tokens_details":{"text_tokens":732,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":635,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":732,"tokens_out":95,"duration_ms":5612,"temperature":1.0,"reasoning_tokens":635,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-15T14:53:11.519338+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"For a layer whose measured input covariance is highly ill-conditioned, compute the information-theoretic lower bound on rate for a target output mean-square error; if WaterSIC’s realized rate exceeds that bound by more than 0.255 bits, or if at matched total rate WaterSIC fails to improve layer output error and downstream perplexity relative to equal-rate GPTQ, the central claim is false.","supporting_citations":[],"review_version":1}