{"id":"ec99eb69-a2c8-4a3e-91a7-31186136a2ab","arxiv_id":"2607.14247","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"IR-HARQ can request exactly the needed retransmission size by predicting it from SNR or first-transmission reliability values, approaching the undetected-error floor with up to 60% smaller retransmissions.","lead":"This paper designs smarter retransmission feedback for wireless HARQ: instead of a simple yes/no ACK, the receiver predicts how many extra coded bits are needed and requests only those. It uses channel statistics and machine learning to size retransmissions, reporting up to 60% smaller retransmissions and 96% prediction accuracy in simulations.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Theorem 1's proof relies on an unproved and false assertion that I(M;Y^{n+1}|Y^n)>0 whenever uncertainty remains; this gap undercuts the achievability claim used to justify finite-Δn2 LUT fitting.","rationale":"The paper's Lemma 1 is correct, and the empirical demonstrations (polar-code savings, LLR predictor accuracy) are plausible and of interest. However, the proof of Theorem 1—the only theoretical support for the claim that the uBLER floor is asymptotically achievable—contains a demonstrably false step: the assertion that I(M;Y^{n+1}|Y^n) is strictly positive whenever uncertainty remains. A simple systematic-plus-parity IR-HARQ scheme provides a concrete case where this mutual information is zero despite H(M|Y^n)>0. The proof also fails to condition the entropy on E1^d. These are not merely stylistic gaps; they invalidate the logical chain that leads to Pr[E|E1^d]→0 and hence to the design principle used to build the LUT. Since the theorem's conclusion might still be true under weaker or different assumptions, the appropriate action is not rejection but a request for a corrected proof. The reader's verdict of CONDITIONAL already captures this, so no change is needed. The proposed test settles the question by either exhibiting the false assertion or showing that the conclusion holds despite it, which would still require a rewrite of the proof.","tokens_in":8665,"tokens_out":14985,"duration_ms":155877,"concrete_test":"Construct a concrete IR-HARQ counterexample: binary symmetric channel with crossover ε, k=2 info bits, first transmission sends bits (M1,M2), second transmission sends M1⊕M2. Choose a first-round noise realization so that the posterior is uniform over {00,01} (e.g., M2 known to be 0, M1 ambiguous). Analytically compute I(M;Y_3|Y^2) and verify it equals 0 while H(M|Y^2)=1 bit. This directly falsifies the proof's central assertion. Then extend the retransmission to length N by appending random parity bits and compute Pr[E|E1^d] as N grows; if it decays to 0, the theorem's conclusion may still hold but requires a corrected proof; if it does not decay, the achievability claim itself fails, and the design principle must be re-derived without relying on Theorem 1.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central theoretical result (Theorem 1) is intended to justify the design principle of choosing the smallest retransmission length Δn2 such that Pr[E] ≈ β Pr[E1^(u)]. Its proof reduces the problem to showing H(M|Y^{nTmax})→0, and the only step ruling out a strict sub-ceiling L<k is the claim: 'For a non-degenerate channel and a reasonable IR-HARQ scheme, I(M;Y^{n+1}|Y^n) is strictly positive whenever uncertainty about M remains.' This assertion is neither proved nor generally true. For a memoryless channel, I(M;Y_{n+1}|Y^n) = H(Y_{n+1}|Y^n) − H(Y_{n+1}|M). If the posterior after n observations is supported on messages that share the same next transmitted symbol X_{n+1}, then Y_{n+1} is conditionally independent of M given Y^n and the mutual information is exactly zero. Example: k=2, first transmission sends the two info bits, second transmission sends their XOR; if the posterior is uniform over {00,01} (first bit ambiguous, second bit known), then X_{n+1}=0 for both messages and I=0 while H(M|Y^n)=1 bit. Thus the proof's key step is false for a perfectly reasonable IR-HARQ scheme. Moreover, even if each incremental bit had strictly positive mutual information, positivity alone does not imply the sum diverges (positive summable sequences exist), so the limit argument would still be incomplete. A further issue is that the entropy bound should condition on E1^(d), but the proof uses unconditional entropy. Consequently, Theorem 1's achievability claim is not established as written, and the paper's reliance on it to justify the simulation-based LUT (Section III-B) is a genuine weak point. The empirical results may still hold, but they lose the theoretical anchor claimed in the abstract.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper considers incremental redundancy (IR) HARQ systems and proposes two feedback mechanisms. The first is an SNR-driven redundancy-allocation policy: after deriving a lower bound (Lemma 1) stating that any IR-HARQ scheme's overall BLER cannot fall below the first-transmission undetected BLER, the authors claim asymptotic achievability of this bound (Theorem 1) and use the resulting design principle to construct an SNR-dependent second-transmission length Δn₂(SNR) for a polar-coded system. Numerical results in §V-A report retransmission savings up to 60% relative to equal-size retransmissions. The second contribution is a per-codeword early-feedback predictor that uses first-transmission LLRs (or SNR) to decide, without full decoding, whether the codeword is decodable, how many additional redundancy versions are needed, or whether retransmission is futile; link-level simulations for 5G NR LDPC codes report about 96% prediction accuracy.","tokens_in":9179,"tokens_out":4984,"duration_ms":54753,"significance":"If the claims held as stated, the paper would make a useful engineering contribution: the uBLER floor is a simple but easily overlooked constraint in HARQ resource allocation, and predicting per-codeword redundancy needs from first-round LLRs is a plausible way to save resources and latency. The machine-code demonstrations are not shipped, but the experimental setup for the LDPC study is described in enough detail to be reconstructed. The main weakness is that the central theoretical result, Theorem 1, rests on an unproved and false assertion about conditional mutual information, so the asymptotic achievability of the uBLER lower bound is not established. In addition, the headline 60% saving is obtained by simulation-based optimization of the very Δn₂(SNR) values being reported, rather than by a rule derived from the theory, which limits the generality of that claim. The LLR-predictor results, while encouraging, lack the dataset and statistical details needed to substantiate the 96% accuracy claim. These are load-bearing issues, but the underlying ideas are plausible and could be repaired in a major revision.","major_comments":[{"comment":"The proof's key step is the assertion that I(M;Y^{n+1}|Y^n)>0 whenever uncertainty about M remains. This is unproved and is not generally true. Example: k=2, first transmission sends the two information bits, second transmission sends their XOR; when the posterior is uniform over {00,01}, H(M|Y^n)=1 bit but the next transmitted symbol is identical for both hypotheses, so I(M;Y^{n+1}|Y^n)=0. Thus the argument does not rule out a strict sub-ceiling L<k. Even if every increment had strictly positive conditional mutual information, positivity alone does not imply that the sum diverges; one would need a uniform lower bound or a summability argument. This step is essential for the achievability claim that Pr[E|E1^d]→0 as Δn→∞.","section":"§III-B, Theorem 1 proof (Eq. (10))"},{"comment":"The MAP error bound is written with the unconditional posterior P_{M|Y^{nTmax}} and unconditional entropy H(M|Y^{nTmax}), but the target is Pr[E|E1^d], i.e., the error probability conditioned on the detected-error event after the first transmission. The distribution of the received signal given E1^d differs from the unconditional distribution, and H(M|Y^n)→0 does not by itself imply H(M|Y^n,E1^d)→0. The proof must either condition throughout on E1^d or show that the event E1^d has vanishing influence on the posterior in the large-Δn limit. Without this, the conclusion Pr[E|E1^d]→0 is not justified.","section":"§III-B, Eqs. (6)–(9)"},{"comment":"The reported redundancy vector Δn₂(SNR)=[544,444,256,200,144,100] is obtained by 'optimizing Δn₂ as a function of SNR ... via simulations' (Section V-A). Consequently, the 60% saving is a fitted outcome on the same SNR grid used for optimization, not a prediction of Theorem 1 or of a principled resource-allocation rule. The claim that the uBLER bound 'can be closely approached' is demonstrated only at the fitting points; no held-out SNR values, interpolation rule, or uncertainty quantification is given. The abstract's 'savings up to 60%' should be presented as a simulation-optimized example, not as a general consequence of the theory.","section":"§V-A, Fig. 2 and Δn₂(SNR) vector"},{"comment":"The 96% accuracy figure is reported for a single 80/10/10 split, but the paper does not state the dataset size, the class counts after rejection sampling, the actual SNR distribution of test samples, or confidence intervals for the accuracy. The comparison with the SNR-based classifier is made only under 'precise channel-quality knowledge'; the claimed degradation under estimation error or mobility is not quantified. The 'dominant error mode' of over-allocation by one RV is described but not measured. To support the abstract claim that 'both predictors achieve high accuracy (about 96%)', the paper needs a more complete evaluation, including baselines and per-class statistics.","section":"§V-B, LLR-based predictor"}],"minor_comments":[{"comment":"The abstract says 'both predictors achieve high accuracy'; the SNR-based classifier reaches 96% only under precise channel-quality knowledge. Please state this qualification in the abstract as well as in §V-B1.","section":"Abstract"},{"comment":"The five classes are described in Section IV as 'already decodable, one/two/three additional RVs, undecodable', while the dataset labels in §V-B are 'rounds 1–4 mapped to classes 1–4'. Clarify the correspondence between 'additional RVs' and 'rounds', since RV 0 is the first round and subsequent RVs are added in the sequence 0→2→3→1.","section":"§IV, class definitions"},{"comment":"For low SNR, the total transmitted length exceeds 512 bits (e.g., 256+544=800); the statement that the shorter code is obtained from the longer one via puncturing deserves a clearer explanation, since the second transmission at low SNR is longer than the first.","section":"§V-A, footnote 3"},{"comment":"Reference [14] is the authors' own patent application and is cited as the source of the LUT-construction method. It would strengthen the paper to describe the LUT construction in the manuscript itself, so that the method is self-contained and not dependent on an unpublished application.","section":"References"},{"comment":"No code or data availability statement is provided. Given the empirical nature of Section V, a release of the dataset-generation scripts and the trained-predictor evaluation details would improve reproducibility.","section":"Overall"}],"recommendation":"major_revision","confidential_remarks":"The paper is a plausible industry-oriented engineering contribution, but the current version should not be published as is. The Theorem 1 proof contains a false assertion that invalidates the achievability claim as written; the 60% saving is a simulation-optimized curve rather than a derived result; and the ML evaluation lacks statistical detail. The ideas are worth pursuing, and a revised version with a corrected theorem (e.g., imposing a uniform positive-information condition or restricting to good rateless code sequences) and a more rigorous empirical evaluation could be suitable. I would not reject the work outright because the lower-bound observation and the per-codeword prediction approach are useful and likely repairable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe genuinely new thing here is the second contribution: a CNN that takes first-transmission LLRs and predicts not just 'decodable or not' but how many additional RVs are needed, or whether to give up and request an MCS change. That is a real extension of the binary decodability predictors in [10]-[12], and the link-level results with 5G LDPC codes look plausible. The uBLER lower bound (Lemma 1) is elementary but correct, and it is a useful design lens for two-shot HARQ.\n\nThe soft spots are concentrated in the theory. Theorem 1 claims the bound is achievable as the retransmission length grows. The proof hinges on the assertion that I(M;Y^{n+1}|Y^n) is strictly positive whenever uncertainty remains, and that is simply false. A simple IR scheme that sends the XOR of two information bits after a first transmission that leaves two messages equally likely is a counterexample: both messages produce the same next symbol, so the conditional mutual information is zero even though H(M|Y^n) is one bit. Even if strict positivity held, it would not imply the limit reaches H(M); a sum of positive numbers can converge below the ceiling. There is also a conditioning mismatch: the proof bounds the unconditional error but needs the conditional error given a detected first-round error. So the achievability claim is not established.\n\nThe rest of the paper depends on this theorem mainly as motivation. The SNR-based LUT for Δn2 is obtained by simulation optimization, so the 60% saving is a fitted outcome, not a predicted one. That is fine as a proof-of-concept, but it should be described that way. The ML accuracy of 96% is reported on a class-balanced test set, which hides the real class imbalance; the confusion matrix shows the dominant error is over-allocation by one RV, which is benign. No code or data are released, so the numbers cannot be checked.\n\nBottom line: this is a plausible engineering contribution with a load-bearing proof gap in the theoretical section. It deserves a serious referee, not a desk reject, because the empirical idea is useful and the lower bound is a nice design lens. But the authors should be asked to either prove a much weaker version of Theorem 1 (e.g., showing the floor can be approached for an explicit family of codes) or explicitly drop the achievability claim and present the LUT as a simulation-based design.\n\nFor peer review: yes, engage with it, but expect a revision.","headline":"Useful engineering idea backed by a flawed proof: the per-codeword redundancy predictor is worth a look, but Theorem 1 should not be trusted as written.","tokens_in":9636,"tokens_out":4544,"would_cite":false,"duration_ms":41567,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["94A15","94A24"],"pacs":[],"model":"deepseek-v4-flash","headline":"Incremental-redundancy HARQ systems cannot beat the first transmission's undetected-error rate, and the paper shows how to allocate retransmission bits to approach that floor and predict per-codeword needs from early reliability information","keywords":["incremental redundancy","HARQ","undetected block error rate","resource allocation","early feedback","LLR-based prediction","5G NR LDPC","polar codes"],"falsifier":"Construct or simulate an IR-HARQ scheme whose incremental bits are deterministic repetitions of already-transmitted bits, so I(M;Y^{n+1}|Y^n)=0 while residual message uncertainty is positive; for such a scheme, Pr[E|E1^(d)] will not vanish and the overall BLER will sit above the uBLER floor, directly contradicting Theorem 1. Alternatively, run the paper's polar-coded setup at fixed SNR with the optimized Δn2 and check-by exact MAP enumeration or very long Monte Carlo simulation whether overall BLER actually reaches β·Pr[E1^(u)]; a systematic gap would show the asymptotic result does not carry","tokens_in":8575,"feed_emoji":"📡","tokens_out":3782,"duration_ms":42835,"temperature":0.7,"pith_summary":"This paper tries to establish that in incremental-redundancy hybrid ARQ (IR-HARQ), the overall block error probability is lower-bounded by the probability of an undetected error after the first transmission, and that this lower bound is asymptotically achievable with enough redundancy under optimal decoding. A sympathetic reader cares because this reframes retransmission design as a resource-allocation problem: choose the smallest second-transmission redundancy that just reaches the undetected-error floor instead of conservatively over-provisioning. The paper develops two mechanisms to do this: an SNR-based look-up table for one- or two-shot resource allocation, and an early-feedback predictor using first-transmission log-likelihood ratios to decide, per codeword, whether decoding is already possible or how many extra redundancy versions are needed. Link-level results with polar codes show that approaching the floor saves up to 60% of retransmission bits at high SNR, and a CNN-based predictor for 5G NR LDPC codes reaches about 96% accuracy in predicting the minimum number of redundancy versions needed. The finite-blocklength usefulness of the asymptotic achievability claim depends on an unproved entropy step in the proof.","feed_headline":"Retransmissions can't beat first-round undetected errors","feed_subtitle":"Sizing retransmission bits to the uBLER floor saves up to 60% redundancy; an LLR-based predictor reaches 96% accuracy.","key_machinery":"The key object is the undetected-error floor: the event E1^(u), an erroneous first-round decoding that passes the error-detection check, is a subset of the overall error event E, forcing Pr[E] ≥ Pr[E1^(u)]. The achievability argument hinges on the entropy chain linking MAP decoding success probability to conditional entropy: 1 − Pr[E] ≥ 2^{-H(M|Y)}, plus the unproved assertion that the conditional mutual information I(M;Y^{n+1}|Y^n) is strictly positive whenever message uncertainty remains, so the posterior entropy collapses as redundancy grows. The design machinery is a SNR-to-redundancy look-up table, built by simulation so that overall BLER stays near the uBLER floor, and an early-feedbac","core_discovery":"The paper's central claim is Lemma 1: for any IR-HARQ scheme, the overall block error probability Pr[E] is at least Pr[E1^(u)], the undetected block error rate of the first transmission, because an undetected first-round error already terminates retransmissions and becomes an overall error. Theorem 1 claims that equality is achievable as the total retransmission redundancy grows without bound under a MAP decoder, using an entropy argument: the conditional probability of correct decoding is at least 2^{-H(M|Y)}, and if the conditional mutual information contributed by each new incremental bit is strictly positive while message uncertainty remains, then H(M|Y) tends to zero and the second-roun","pith_inferences":["The uBLER-floor argument should apply to any rateless or incremental scheme whose first chunk is protected by a finite error-detection code; comparing competing HARQ feedback designs should therefore be done against uBLER, not raw BLER.","Because the main error mode is one-RV over-allocation, a testable refinement would be to schedule the predicted number of RVs minus one and keep a conventional ACK/NACK fallback, trading a small residual failure probability for larger resource savings.","The asymptotic achievability, if the entropy step is rigorized, would give a non-asymptotic bound on the required second-transmission length as a function of the residual conditional entropy—a natural finite-blocklength extension that the paper leaves implicit.","The savings from the proposed allocation should grow as the CRC gets shorter, since the uBLER floor rises; this suggests the scheme is especially attractive for polar codes with short CRCs, as the paper notes in its motivation."],"forward_implications":["Retransmission sizes in IR-HARQ should be sized against the first-round undetected-error rate, not against the target BLER alone, yielding large redundancy savings at high SNR without sacrificing reliability.","With strict two-transmission latency, choosing the second-transmission length to just reach the uBLER floor is enough; extra bits beyond that floor waste resources without reducing the error rate.","Early LLR-based feedback can decide the retransmission budget without running the full decoder, allowing successful decoding within at most two transmission occasions and reducing expected latency.","The dominant one-RV over-allocation error mode is resource-inefficient but not reliability-critical, so the predictor can be used as a conservative scheduler without risking additional decoding failures."],"fun_headline_variants":["First-round undetected errors set HARQ's error floor","Sizing retransmissions to the error floor saves up to 60%","96%-accurate predictor decides retransmission need early","Undetected first-transmission errors are HARQ's hard floor"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing assumption is the unproved step in Theorem 1 that every additional incremental bit carries strictly positive conditional mutual information about the message whenever posterior uncertainty remains; if the retransmission stream can stall this information flow, the claimed approach to the uBLER floor with finite redundancy is not established.","fun_headline_variants_meta":{"raw":{"variants":["First-round undetected errors set HARQ's error floor","Sizing retransmissions to the error floor saves up to 60%","96%-accurate predictor decides retransmission need early","Undetected first-transmission errors are HARQ's hard floor"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000431,"raw_usage":{"total_tokens":2061,"prompt_tokens":795,"completion_tokens":1266,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":539,"completion_tokens_details":{"reasoning_tokens":1205}},"tokens_in":539,"tokens_out":1266,"duration_ms":12694,"temperature":1.0,"reasoning_tokens":1205,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T02:38:39.448531+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Construct or simulate an IR-HARQ scheme whose incremental bits are deterministic repetitions of already-transmitted bits, so I(M;Y^{n+1}|Y^n)=0 while residual message uncertainty is positive; for such a scheme, Pr[E|E1^(d)] will not vanish and the overall BLER will sit above the uBLER floor, directly contradicting Theorem 1. Alternatively, run the paper's polar-coded setup at fixed SNR with the optimized Δn2 and check-by exact MAP enumeration or very long Monte Carlo simulation whether overall BLER actually reaches β·Pr[E1^(u)]; a systematic gap would show the asymptotic result does not carry","supporting_citations":[],"review_version":1}