REVIEW 3 major objections 4 minor 29 references
Closed-Loop Bayesian Bandit Encoder with GRAND Receiver for a Bursty Interference Channel
T0 review · 3 major / 4 minor · reviewed 2026-08-01 · deepseek-v4-flash
Pith's one-line read A receiver that learns the interference pattern makes non-interleaved transmission beat interleaving, with 5–10x lower block error rate in the evaluated configuration.
desk verdict A genuinely new closed-loop GRAND/bandit architecture with a plausible but under-tested preference reversal that rests on an asymmetric decoder comparison. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is GRAND (Guessing Random Additive Noise Decoding), whose codebook test and noise model are modular: the code determines only which candidate noise patterns are valid, while the noise model determines the order in which they are tested. The learned noise model builds a hidden Markov model over interferer on/off states from the estimated parameters, computes per-bit posterior flip probabilities via forward–backward smoothing, and orders GRAND queries with an SGRAND-style max-heap under a factorized posterior approximation. A discounted Thompson sampler at the transmitter selects between the two arms using a goodput-minus-latency reward, and a model-informed variant injects
What would settle it
Run the closed-loop system with the interleaved arm also equipped with a learned noise model that accounts for the interleaver permutation; if the interleaved arm's post-adaptation block error rate then matches or beats the non-interleaved arm's, the claimed reversal is an artifact of asymmetric adaptation.
Extended reading notes
Core claim
In the evaluated closed-loop configuration, the preferred transmission mode changes as a result of the receiver's own learning. Before channel estimation converges, the interleaved arm earns higher reward; after the receiver switches to the learned HMM-ordered GRAND decoder, the non-interleaved arm becomes preferable, with block error rates 5–10x lower than the interleaved arm across five simulation seeds. The reversal is endogenous: it is the decoder adaptation, not a change in the channel, that alters the reward distribution of the non-interleaved arm. In the reference configuration the learned noise model reduces block error rate by approximately one order of magnitude relative to ORBGRAN
Load-bearing premise
Only the non-interleaved arm receives the learned HMM decoder; the interleaved arm continues to use ORBGRAND, so the central preference reversal is measured against an unadapted baseline rather than two equally adapted options.
Editorial extensions
If this is right
- If the receiver can learn the interference structure, interleaving's reliability advantage disappears while its delay cost remains, so non-interleaved transmission becomes the preferred mode.
- Adapting the GRAND noise model to the learned channel reduces block error rate by roughly an order of magnitude compared with an untuned ORBGRAND decoder on the same channel.
- Using partial channel estimates before full convergence reduces pre-convergence block error rate by up to 4.5x, so early adaptation is beneficial even before the estimator is confident.
- Confidence-weighted model predictions accelerate the preference reversal by 23–38% and reduce post-transition selection of the lower-reward arm by about 65%.
- Under a 100 MHz 5G NR-like symbol rate, the learning transient corresponds to a few milliseconds of occupied symbol time, suggesting the mechanism can operate within existing control-loop timescales.
Reading between the lines
- The measured preference reversal compares an adapted non-interleaved arm against a fixed ORBGRAND interleaved arm; if the interleaved arm also received a learned, interleaver-aware noise model, the post-adaptation gap could shrink or vanish. The paper explicitly leaves this out of scope, but the headline claim depends on this asymmetry.
- The reward's latency weight c_l is a free design parameter; the boundary between the two arms likely shifts with c_l, so the claimed reversal is a statement about this specific configuration, not a universal criterion for disabling interleaving.
- The same closed loop could be applied to a richer family of transmission modes (different interleaver depths, code rates, or MCS entries), with the bandit selecting among them while the receiver adapts its noise model for each; whether the reversal persists across code parameters is a testable extension.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a closed-loop system for packet-level selection between a non-interleaved random linear code and the same code with cross-codeword interleaving over an AWGN channel with unknown on/off interferers. The receiver uses GRAND with a replaceable noise model; aggregate channel statistics are fed back to a Bayesian learner at the transmitter, which estimates interference amplitudes/duty cycles and, once confident, enables an HMM-based noise model that computes posterior per-bit flip probabilities and orders GRAND queries. A discounted Thompson sampler chooses the transmission mode using a goodput-minus-latency reward. Simulations in five phases show that the interleaved arm is preferred before the receiver adapts and the non-interleaved arm becomes preferred after the learned HMM decoder is activated, with 5–10× lower BLER across five seeds; model-informed pseudo-observations accelerate the switch and reduce post-transition selection of the inferior arm.
Significance. If the central claim holds, the paper makes a useful conceptual contribution: receiver-side decoder adaptation can alter the relative value of interleaved versus non-interleaved transmission, so interleaving policy, channel learning, and decoder design should be treated as coupled. The paper is careful in its structure: V0 validates GRAND against an analytic abandonment bound, V1 provides oracle-parameter decoder comparisons, and the limitation statements in Sections V-E, VI-C, and XI are unusually explicit. The five-seed consistency is suggestive. However, the headline preference reversal is evaluated against an asymmetric decoder baseline, and the quantitative support is thin—five seeds, one operating point, point estimates, and no code/data—so the result is plausible but not yet established at the strength claimed in the abstract.
major comments (3)
- [Section VI-D and Table VI] The post-adaptation comparison is asymmetric. Section VI-D states that 'only arm A receives the learned model. Arm B continues to use ORBGRAND.' Table VI's 5–10× BLER advantage of arm A over arm B therefore bundles the arm choice with the decoder choice: the non-interleaved arm gets the learned HMM ordering, while the interleaved arm is stuck with a fixed generic decoder. The paper asserts the interleaver weakens temporal dependence but does not quantify this. Since the receiver already collects statistics in channel order for arm B (Section V-A), an interleaver-aware learned decoder for arm B is a natural matched baseline and is explicitly deferred as out of scope. Without that baseline, the preference reversal may be partly a decoder-capability artifact rather than a property of non-interleaved transmission. Please either add the matched baseline or reframe the claim as 'non-interleave
- [Section IX-E, Table VI] The statistical support for the headline claim needs strengthening. Table VI reports BLER over the full run, but arm B is played predominantly before the confidence switch and arm A predominantly after it, so the full-run 5–10× BLER ratio conflates the pre- and post-adaptation regimes; the per-arm separation is shown only for seed 0 in Fig. 5(b). The claim that non-interleaved transmission is preferable after decoder activation needs per-arm BLER measured over the post-switch interval for all five seeds, with error counts and Wilson intervals. Additionally, all numerical results rest on five seeds at a single Eb/N0, and the manuscript provides no code or data availability statement, which limits reproducibility of the simulation-based claims.
- [Eq. (7) and Section XI] The 'reliability–latency tradeoff reversal' depends on how latency enters the reward. The reward in Eq. (7) collapses interleaving latency to a constant penalty c_l D with c_l chosen so the penalty is about 10% of maximum goodput; delayed acknowledgments, retransmission timing, and scheduling delays are not modeled, as Section XI acknowledges. Since no sensitivity sweep over c_l is provided, the reward-based preference reversal is conditional on this specific design weight. The paper should either provide a c_l sensitivity analysis or state explicitly that the post-adaptation advantage is primarily a BLER effect, which would make the latency modeling less consequential for the central claim.
minor comments (4)
- [Abstract and Table VII] The abstract's 'approximately 65%' reduction in post-transition selection of the inferior arm is based on three seeds in V3b and the percentage is not derived from the tabulated shares. Please state the calculation or soften the claim.
- [Section X] The airtime conversion treats the BPSK stream as occupying 100 Msym/s after deriving roughly 92 Msym/s from the 5G NR numerology. Clarify that 100 Msym/s is an idealization used for the top axes.
- [References] Reference [3] contains stray spaces in the author initials ('V . Y . Wei'); please fix formatting. There are also a few inconsistent spaces in the bibliography that copyediting should catch.
- [Section V-B] The claim that EM failure 'suggests that the difficulty is statistical rather than solely an initialization failure' is reasonable but appears to be based on a small number of fits. A sentence describing the number of initializations and convergence criterion would be helpful.
Circularity Check
No significant circularity: the headline preference reversal is a simulation result driven by real ACK/NACK rewards, not by construction from the paper's assumptions.
full rationale
The paper's central claims are empirical simulation outcomes rather than derivations that reduce to their inputs. The post-adaptation preference for the non-interleaved arm is evaluated on actual simulated ACK/NACK rewards; the HMM noise model is built from estimated channel parameters (amplitudes, duty cycles, ON lengths) and the GRAND query ordering is a stated factorized approximation (Section VI-C), not a quantity fitted to the arm-preference outcome. Oracle comparisons use true channel parameters as independent benchmarks (Section IX-B). The model-informed Thompson sampler injects Monte Carlo predicted utilities computed from the same channel estimate (Section VII-A), but the reported benefit is measured against real rewards and is not a renamed fit: the pseudo-observations are confidence-weighted predictions, not fitted constants reproducing the claimed improvement. The only self-citation is reference [29] in a future-work enumeration of bandit algorithms; it is not load-bearing for any result. The paper explicitly acknowledges that only arm A receives the learned HMM decoder while arm B continues to use ORBGRAND (Section VI-D), and Section XI lists this and other limitations; this asymmetry is a threat to external validity or fairness of comparison, not circularity, because it does not make the observed preference equal to the model assumption by construction. No step was found where a prediction is definitionally identical to a fitted input or where the argument relies on an unverified self-citation chain.
Assumptions & free parameters
free parameters (11)
- latency weight c_l =
0.005
- interleaver depth D =
16
- run threshold tau =
1.5 * sigma_hat
- min run length =
3
- query budget Q =
1e6 (V0), 1e5 (loop)
- feedback period =
200 packets
- TS discount gamma =
0.99
- pseudo-count m0 =
30
- confidence thresholds =
S>100, c>=2, w>=0.15, Hausdorff <0.15
- operating Eb/N0 =
6 dB
- code parameters =
(128,105)
assumptions (7)
- standard math GRAND first-hit is ML if query order matches channel statistics
- domain assumption Interference model: k on/off interferers, exponential OFF, fixed ON length and amplitude, BPSK/AWGN
- domain assumption HMM state space assumes at most one active interferer per symbol
- ad hoc to paper GRAND ordering uses factorized posterior approximation, not full joint HMM posterior
- domain assumption Arm B (interleaved) is not given a learned model; stays ORBGRAND
- domain assumption Feedback link error-free, rate-limited, with no airtime cost
- domain assumption Statistics are collected only from successfully decoded packets
Cite this review
Pith. "Pith review of Closed-Loop Bayesian Bandit Encoder with GRAND Receiver for a Bursty Interference Channel." pith.science (2026). https://pith.science/paper/XWU6QESS
@misc{pith2026260715404,
author = {Pith},
title = {Pith review of: Closed-Loop Bayesian Bandit Encoder with GRAND Receiver for a Bursty Interference Channel},
year = {2026},
howpublished = {\url{https://pith.science/paper/XWU6QESS}},
note = {Machine review of arXiv:2607.15404}
}
abstract
Interleaving mitigates burst errors but introduces decoding delay and removes temporal error structure that a channel-aware decoder could exploit. We consider packet-level selection between a random linear code and the same code used with cross-codeword interleaving, over a channel with an unknown number of on/off interferers. The receiver uses Guessing Random Additive Noise Decoding (GRAND) with a replaceable noise model and feeds aggregate channel statistics back to a Bayesian estimator at the transmitter. Once the interference amplitudes and timing parameters are estimated, the receiver's noise model is replaced: it computes hidden-Markov-model posterior bit-flip probabilities and uses them to order GRAND queries. A discounted Thompson sampler selects between the two transmission modes using a goodput-minus-latency reward whose distribution is endogenously nonstationary: receiver adaptation, rather than channel change, alters the value of each mode. Across five simulation seeds, the interleaved mode is preferred before channel estimation converges. After the learned decoder is activated, the non-interleaved mode becomes preferable because it achieves lower block error rate without interleaving delay. In the reference configuration, the learned noise model reduces block error rate by approximately one order of magnitude relative to ORBGRAND. Using partial channel estimates before full convergence reduces pre-convergence block error rate by up to $4.5\times$. Adding model-predicted utilities as confidence-weighted pseudo-observations reduces post-transition selection of the inferior arm by approximately $65\%$. Under an idealized airtime conversion at a 100~MHz 5G~NR-like symbol rate, the learning transient corresponds to a few milliseconds of occupied symbol time.
Figures
Reference graph
Works this paper leans on
-
[1]
Capacity-achieving guessing random additive noise decoding,
K. R. Duffy, J. Li, and M. M ´edard, “Capacity-achieving guessing random additive noise decoding,”IEEE Transactions on Information Theory, vol. 65, no. 7, pp. 4023–4040, 2019
2019
-
[2]
Ordered reliability bits guessing random additive noise decoding,
K. R. Duffy, W. An, and M. M ´edard, “Ordered reliability bits guessing random additive noise decoding,”IEEE Transactions on Signal Process- ing, vol. 70, pp. 4528–4542, 2022
2022
-
[3]
ORBGRAND is almost capacity-achieving,
M. Liu, Y . Wei, Z. Chen, and W. Zhang, “ORBGRAND is almost capacity-achieving,”IEEE Transactions on Information Theory, vol. 69, no. 5, pp. 2830–2840, 2023
2023
-
[4]
High-throughput and energy-efficient VLSI architecture for ordered reliability bits GRAND,
S. M. Abbas, T. Tonnellier, F. Ercan, M. Jalaleddine, and W. J. Gross, “High-throughput and energy-efficient VLSI architecture for ordered reliability bits GRAND,”IEEE Transactions on Very Large Scale Integration (VLSI) Systems, vol. 30, no. 6, pp. 681–693, 2022
2022
-
[5]
A sub-0.8-pj/bit 16.3-gbps/mm 2 universal soft-detection decoder using ORBGRAND in 40-nm CMOS,
A. Riaz, A. Yasar, F. Ercan, W. An, J. Ngo, K. Galligan, M. M ´edard, K. R. Duffy, and R. T. Yazicigil, “A sub-0.8-pj/bit 16.3-gbps/mm 2 universal soft-detection decoder using ORBGRAND in 40-nm CMOS,” inProceedings of the IEEE International Solid-State Circuits Conference (ISSCC), 2023
2023
-
[6]
Soft maximum likelihood decoding using GRAND,
A. Solomon, K. R. Duffy, and M. M ´edard, “Soft maximum likelihood decoding using GRAND,” inProceedings of the IEEE International Conference on Communications (ICC), 2020
2020
-
[7]
Keep the bursts and ditch the interleavers,
W. An, M. M ´edard, and K. R. Duffy, “Keep the bursts and ditch the interleavers,”IEEE Transactions on Communications, vol. 70, no. 6, pp. 3655–3667, 2022
2022
-
[8]
Guessing random additive noise decoding with symbol reliability information (SRGRAND),
K. R. Duffy, M. M ´edard, and W. An, “Guessing random additive noise decoding with symbol reliability information (SRGRAND),”IEEE Transactions on Communications, vol. 70, no. 1, pp. 3–18, 2022
2022
Show all 29 references
-
[9]
GRAND for Gaussian intersymbol interference channels,
Z. Li and W. Zhang, “GRAND for Gaussian intersymbol interference channels,”arXiv preprint arXiv:2603.08325, 2026
2026
-
[10]
Decoding in the presence of ISI without interleaving – ORBGRAND-AI,
K. R. Duffy, M. Grundei, J. A. Millward, M. Rangaswamy, and M. M ´edard, “Decoding in the presence of ISI without interleaving – ORBGRAND-AI,”arXiv preprint arXiv:2510.14939, 2025
2025 arXiv
-
[11]
Universal decoding over finite-state additive channels via noise guessing,
H. K. Miyamoto and S. Yang, “Universal decoding over finite-state additive channels via noise guessing,”arXiv preprint arXiv:2501.12971, 2025
2025 arXiv
-
[12]
Joint error correction and fading channel estimation enhancement leveraging GRAND,
C. Wiame, K. R. Duffy, and M. M ´edard, “Joint error correction and fading channel estimation enhancement leveraging GRAND,”arXiv preprint arXiv:2506.14756, 2025
2025 arXiv
-
[13]
Capacity of a burst-noise channel,
E. N. Gilbert, “Capacity of a burst-noise channel,”Bell System Technical Journal, vol. 39, no. 5, pp. 1253–1265, 1960
1960
-
[14]
Estimates of error rates for codes on burst-noise channels,
E. O. Elliott, “Estimates of error rates for codes on burst-noise channels,” Bell System Technical Journal, vol. 42, no. 5, pp. 1977–1997, 1963
1977
-
[15]
Models for channels with memory and their applications to error control,
L. N. Kanal and A. R. K. Sastry, “Models for channels with memory and their applications to error control,”Proceedings of the IEEE, vol. 66, no. 7, pp. 724–744, 1978
1978
-
[16]
Cognitive radio: Brain-empowered wireless communica- tions,
S. Haykin, “Cognitive radio: Brain-empowered wireless communica- tions,”IEEE Journal on Selected Areas in Communications, vol. 23, no. 2, pp. 201–220, 2005
2005
-
[17]
Cognitive medium access: Exploration, exploitation, and competition,
L. Lai, H. E. Gamal, H. Jiang, and H. V . Poor, “Cognitive medium access: Exploration, exploitation, and competition,”IEEE Transactions on Mobile Computing, vol. 10, no. 2, pp. 239–253, 2011
2011
-
[18]
Optimal rate sampling in 802.11 systems: Theory, design, and implementation,
R. Combes, J. Ok, A. Prouti `ere, D. Yun, and Y . Yi, “Optimal rate sampling in 802.11 systems: Theory, design, and implementation,”IEEE Transactions on Mobile Computing, vol. 18, no. 5, pp. 1145–1158, 2019
2019
-
[19]
Rate adaptation with Thompson sampling in 802.11ac WLAN,
H. Qi, Z. Hu, X. Wen, and Z. Lu, “Rate adaptation with Thompson sampling in 802.11ac WLAN,”IEEE Communications Letters, vol. 23, no. 10, pp. 1888–1892, 2019
2019
-
[20]
Bayesian link adaptation under a BLER target,
V . Saxena and J. Jald´en, “Bayesian link adaptation under a BLER target,” inProceedings of the IEEE 21st International Workshop on Signal Processing Advances in Wireless Communications (SPAWC), 2020
2020
-
[21]
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples,
W. R. Thompson, “On the likelihood that one unknown probability exceeds another in view of the evidence of two samples,”Biometrika, vol. 25, no. 3–4, pp. 285–294, 1933
1933
-
[22]
A tutorial on Thompson sampling,
D. Russo, B. V . Roy, A. Kazerouni, I. Osband, and Z. Wen, “A tutorial on Thompson sampling,”Foundations and Trends in Machine Learning, vol. 11, no. 1, pp. 1–96, 2018
2018
-
[23]
On upper-confidence bound policies for switching bandit problems,
A. Garivier and E. Moulines, “On upper-confidence bound policies for switching bandit problems,” inProceedings of the International Conference on Algorithmic Learning Theory (ALT), 2011
2011
-
[24]
Link adaptation using Joint-Thompson sampling,
V . Vinjam, M. Kolavennu, M. Vajha, and K. P. Narayanaprasad, “Link adaptation using Joint-Thompson sampling,”arXiv preprint arXiv:2607.11075, 2026
2026 arXiv
-
[25]
Soft-output (SO) GRAND and iterative decoding to outperform LDPC codes,
P. Yuan, M. M ´edard, K. Galligan, and K. R. Duffy, “Soft-output (SO) GRAND and iterative decoding to outperform LDPC codes,”IEEE Transactions on Wireless Communications, 2025
2025
-
[26]
Finite-time analysis of the multiarmed bandit problem,
P. Auer, N. Cesa-Bianchi, and P. Fischer, “Finite-time analysis of the multiarmed bandit problem,”Machine Learning, vol. 47, no. 2–3, pp. 235–256, 2002
2002
-
[27]
Deterministic sequencing of exploration and exploitation for multi-armed bandit problems,
S. Vakili, K. Liu, and Q. Zhao, “Deterministic sequencing of exploration and exploitation for multi-armed bandit problems,”IEEE Journal of Selected Topics in Signal Processing, vol. 7, no. 5, pp. 759–767, 2013
2013
-
[28]
The non- stochastic multiarmed bandit problem,
P. Auer, N. Cesa-Bianchi, Y . Freund, and R. E. Schapire, “The non- stochastic multiarmed bandit problem,”SIAM Journal on Computing, vol. 32, no. 1, pp. 48–77, 2002
2002
-
[29]
Combinatorial network op- timization with unknown variables: Multi-armed bandits with linear rewards and individual observations,
Y . Gai, B. Krishnamachari, and R. Jain, “Combinatorial network op- timization with unknown variables: Multi-armed bandits with linear rewards and individual observations,”IEEE/ACM Transactions on Net- working, vol. 20, no. 5, pp. 1466–1478, 2012
2012
Reviewed August 1, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.