REVIEW 3 major objections 4 minor 1 cited by
Link Adaptation Using Joint-Thompson Sampling
T0 review · 3 major / 4 minor · reviewed 2026-07-14 · grok-4.5
Pith's one-line read Joint Thompson sampling that respects ordered MCS success rates delivers competitive link throughput in every tested regime.
desk verdict Correct Gibbs sampler for ordered Beta priors in MCS selection; useful robustness data, but N=1000 mixing for K=29 is unproven. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Multivariate ordered Beta (MOB) prior together with a Gibbs sampler that draws from the correct restricted-Beta conditionals; each draw automatically lies inside the ordered simplex and the posterior stays MOB after a Bernoulli observation.
What would settle it
Re-run the same 5G-NR Doppler simulations with N reduced to a few tens of Gibbs iterations (or with an exact MOB sampler) and check whether the reported throughput gaps versus UTS/LTS reverse or disappear.
Extended reading notes
Core claim
By sampling the vector of MCS success probabilities jointly from a multivariate ordered Beta prior (instead of independent Betas), Thompson sampling automatically respects the physical ordering of the arms; the posterior remains ordered after every ACK/NACK update, and the resulting online policy yields competitive spectral efficiency in every static and time-varying setting examined.
Load-bearing premise
That a fixed number of Gibbs steps (N=1000) produces samples close enough to the true ordered posterior for every online MCS decision to stay near-optimal.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper formulates MCS selection for link adaptation as a multi-armed bandit and proposes Joint-TS, which places a multivariate ordered Beta (MOB) prior on the vector of success probabilities so that samples always obey the natural monotonicity θ1 ≥ ··· ≥ θK. Conjugacy of the MOB under Bernoulli observations is shown, and sampling is performed by Gibbs iteration over restricted-Beta conditionals (Algorithm 2, Lemma 1). Empirical comparisons against OLLA, MTS, UTS and LTS are reported for static SINR and for three Doppler values, both with perfect CQI and without CQI, using the 29 MCS indices of 5G NR. The central claim is that Joint-TS yields competitive throughput and is more robust across these regimes than the listed baselines.
Significance. If the sampler is faithful and the empirical gains hold under more realistic channel models, the work supplies a clean, table-free Bayesian alternative to OLLA and to independent-Beta Thompson sampling that correctly exploits monotonicity. The conjugacy argument and the explicit contrast with the non-joint sequential sampler of CoTS [18] are technically sound contributions. The multi-scenario simulation suite (static, Doppler, CQI/no-CQI) is a useful practical stress test for the 5G/6G community. The absence of regret bounds is acceptable for an applications-oriented paper, but the lack of any mixing diagnostic for the K=29 case that underpins every reported number currently limits the strength of the attribution of robustness to the joint prior.
major comments (3)
- Sec. IV-C, Lemma 2 and Algorithm 2: Lemma 2 only guarantees total-variation convergence of the Gibbs chain as N o∞. The sole empirical check is a visual 2-D histogram for K=2 (Fig. 1c). All throughput results, however, use K=29 MCS indices with a hard-coded N=1000 restarts from ordered uniforms at every TTI. No effective-sample-size, autocorrelation, or decision-stability diagnostic is supplied for the parameter regimes that arise after many updates (large or unbalanced α,β). Without such evidence the claim that Joint-TS realises the true MOB posterior—and that this is the source of its robustness—remains unverified for the experimental setting that matters.
- Sec. V and Table II: The robustness narrative rests on averages over only 20 independent trials and a single-path channel model. For the high-Doppler (111 Hz) no-CQI cell the absolute throughputs of UTS, LTS and Joint-TS are close (1.34 / 1.29 / 1.32 bps/Hz); it is unclear whether the ranking is statistically stable. A modest increase in trial count, multi-path models, and confidence intervals would be needed before the “consistent performance in all scenarios” claim can be regarded as established.
- Sec. V (CQI-less paragraph): The exponential forgetting window w=50 is fixed by hand and the text defers adaptive choice of w to future work. Because the same w is used for the smoothed MTS/UTS/Joint-TS variants, any sensitivity of the ranking to w is unexplored; a short ablation would strengthen the comparison.
minor comments (4)
- Fig. 2 captions and Table II: units and averaging horizon (1000 TTIs) should be stated uniformly in both places; the Oracle curves in panels (b)–(c) are not defined in the caption.
- Remark 1 and Algorithm 2: the special case of identical (α,β) that reduces to sorting i.i.d. Betas is useful; a one-line note that the general Gibbs still applies when some α i,βi become large would help implementers.
- Typographical: “pyitpp” / “py-itpp” inconsistency; missing space in “ai .iith.ac.in”; “BLER(i,c)” sometimes written without spaces.
- Related-work paragraph: a brief quantitative comparison of wall-clock cost of N=1000 Gibbs sweeps versus independent Beta sampling would clarify real-time feasibility at the MAC layer.
Circularity Check
No circularity: Joint-TS derivation and empirical claims rest on standard MOB prior, Gibbs sampling, and external simulator benchmarks, not on self-fitted predictions or load-bearing self-citations.
full rationale
The paper defines Joint-TS by replacing independent Beta priors with the multivariate ordered Beta (MOB) distribution of Al-Saidi et al. [16], proves that the posterior remains MOB under Bernoulli updates (Eq. 3), and samples via a standard Gibbs procedure whose conditional distributions are restricted Betas (Lemma 1, Algorithm 2). Lemma 2 only asserts asymptotic total-variation convergence of the Markov chain; the finite-N=1000 choice is an implementation detail, not a fitted quantity later re-labeled as a prediction. Throughput results are obtained by running the algorithm against the independent py-itpp channel simulator and published baselines (OLLA, MTS, UTS, LTS) under static and Doppler conditions; no parameter is estimated from the same data that is later claimed as a forecast. The critique of CoTS [18] is a technical distinction about sampling correctness, not a self-citation that forces the present claims. Consequently the derivation chain contains no self-definitional loop, no fitted-input-as-prediction step, and no uniqueness theorem imported from overlapping authors. Score 0 is therefore the correct assessment.
Assumptions & free parameters
free parameters (3)
- initial Beta hyperparameters αi, βi =
1
- Gibbs iterations N =
1000
- smoothing window w (CQI-less) =
50
assumptions (4)
- domain assumption MCS success probabilities satisfy θ1 ≥ θ2 ≥ ⋯ ≥ θK for any fixed channel state.
- domain assumption Transmission outcomes are i.i.d. Bernoulli given the chosen MCS and the (unknown) success probability.
- standard math The Gibbs chain on the ordered simplex is positive Harris-recurrent and converges in total variation to the MOB (Lemma 2).
- domain assumption Spectral efficiencies s1 < s2 < ⋯ < sK are known constants.
invented entities (1)
-
Joint-TS algorithm (MOB prior + Gibbs sampling for MCS selection)
Cite this review
Pith. "Pith review of Link Adaptation Using Joint-Thompson Sampling." pith.science (2026). https://pith.science/paper/S54HW5TE
@misc{pith2026260711075,
author = {Pith},
title = {Pith review of: Link Adaptation Using Joint-Thompson Sampling},
year = {2026},
howpublished = {\url{https://pith.science/paper/S54HW5TE}},
note = {Machine review of arXiv:2607.11075}
}
read the original abstract
The choice of Modulation and Coding (MCS) type for a particular channel condition is made through link adaptation (LA) algorithms that operate at the MAC layer. These algorithms rely on the ACK/NACK statistics and the channel quality index (CQI) feedback. Several existing works model LA as a multi-armed bandit (MAB) problem across cellular and Wi-Fi links. In the MAB formulation, each available MCS is a Bernoulli arm parameterized by its transmission success probability, and the goal is to design a selection strategy that accrues maximum reward. Several popular MAB algorithms, such as upper confidence bound (UCB) and Thompson Sampling (TS), have been proposed in the literature. Using the fact that MCS success probabilities are ordered, we propose the Joint-Thompson Sampling (Joint-TS) algorithm. Unlike classical TS, which assumes independent Beta distributions for each arm, Joint-TS utilizes a multivariate ordered Beta distribution as the prior to preserve the inherent monotonicity of success probabilities. Our simulation results show that while existing MAB algorithms fail in specific scenarios, Joint-TS delivers competitive throughput with robust, consistent performance in all scenarios.
Figures
Forward citations
Cited by 1 Pith paper
-
Closed-Loop Bayesian Bandit Encoder with GRAND Receiver for a Bursty Interference Channel
In simulation, a receiver that learns bursty interference makes non-interleaved transmission preferable to interleaved, reversing the preference once the learned decoder is active.
Reference graph
Works this paper leans on
-
[18]
Link Rate Selection using Constrained Thompson Sampling,
H. Gupta, A. Eryilmaz, and R. Srikant, “Link Rate Selection using Constrained Thompson Sampling,” inIEEE INFOCOM 2019 - IEEE Conference on Computer Communications, 2019, pp. 739–747
2019
-
[1]
Link Performance Models for System Level Simulations of Broadband Radio Access Systems,
K. Brueninghaus, D. Astely, T. Salzer, S. Visuri, A. Alexiou, S. Karger, and G.-A. Seraji, “Link Performance Models for System Level Simulations of Broadband Radio Access Systems,” inIEEE 16th International Symposium on Personal, Indoor and Mobile Radio Communications (PIMRC), vol. 4, 2005, pp. 2306–2311 V ol. 4
2005
-
[2]
Coding Versus ARQ in Fading Channels: How Reliable Should the PHY Be?
P. Wu and N. Jindal, “Coding Versus ARQ in Fading Channels: How Reliable Should the PHY Be?” inIEEE Global Communications Conference (GLOBECOM), 2009, pp. 1–6
2009
-
[3]
Optimizing the Target Error Rate for Link Adaptation,
S. Park, R. C. Daniels, and R. W. Heath, “Optimizing the Target Error Rate for Link Adaptation,” inIEEE Global Communications Conference (GLOBECOM), 2015, pp. 1–6
2015
-
[4]
NR; Physical layer procedures for data,
3rd Generation Partnership Project (3GPP), “NR; Physical layer procedures for data,” 3GPP, Technical Specification TS 38.214, 2018, release 15
2018
-
[5]
A Method for Outer Loop Rate Control in High Data Rate Wireless Networks,
D. Paranchych and M. Yavuz, “A Method for Outer Loop Rate Control in High Data Rate Wireless Networks,” inProceedings IEEE 56th Vehicular Technology Conference (VTC), vol. 3, 2002, pp. 1701–1705 vol.3
2002
-
[6]
Frequency Domain Scheduling for OFDMA with Limited and Noisy Channel Feedback,
K. I. Pedersen, G. Monghal, I. Z. Kovacs, T. E. Kolding, A. Pokhariyal, F. Frederiksen, and P. Mogensen, “Frequency Domain Scheduling for OFDMA with Limited and Noisy Channel Feedback,” inIEEE 66th Vehicular Technology Conference (VTC), 2007, pp. 1792–1796
2007
-
[7]
Analysis of Throughput Performance Statistics for Benchmarking LTE Networks,
V . Buenestado, J. M. Ruiz-Avil ´es, M. Toril, S. Luna-Ram ´ırez, and A. Mendo, “Analysis of Throughput Performance Statistics for Benchmarking LTE Networks,”IEEE Communications Letters, vol. 18, no. 9, pp. 1607–1610, 2014
2014
Show all 22 references
-
[8]
Self-Optimization Algorithm for Outer Loop Link Adaptation in LTE,
A. Dur ´an, M. Toril, F. Ruiz, and A. Mendo, “Self-Optimization Algorithm for Outer Loop Link Adaptation in LTE,”IEEE Communications Letters, vol. 19, no. 11, pp. 2005–2008, 2015
2005
-
[9]
Optimal Rate Sampling in 802.11 Systems: Theory, Design, and Implementation,
R. Combes, J. Ok, A. Proutiere, D. Yun, and Y . Yi, “Optimal Rate Sampling in 802.11 Systems: Theory, Design, and Implementation,”IEEE Transactions on Mobile Computing, vol. 18, no. 5, pp. 1145–1158, 2019
2019
-
[10]
Contextual Multi-Armed Bandits for Link Adaptation in Cellular Networks,
V . Saxena, J. Jald ´en, J. E. Gonzalez, M. Bengtsson, H. Tullberg, and I. Stoica, “Contextual Multi-Armed Bandits for Link Adaptation in Cellular Networks,” inProceedings of the 2019 Workshop on Network Meets AI & ML, ser. NetAI’19. New York, NY , USA: Association for Computi...
2019
-
[11]
Low-Complexity, Low-Regret Link Rate Selection in Rapidly-Varying Wireless Channels,
H. Gupta, A. Eryilmaz, and R. Srikant, “Low-Complexity, Low-Regret Link Rate Selection in Rapidly-Varying Wireless Channels,” inIEEE INFOCOM 2018 - IEEE Conference on Computer Communications, 2018, pp. 540–548
2018
-
[12]
Bayesian Link Adaptation under a BLER Target,
V . Saxena and J. Jald ´en, “Bayesian Link Adaptation under a BLER Target,” inIEEE 21st International Workshop on Signal Processing Advances in Wireless Communications (SPAWC), 2020, pp. 1–5
2020
-
[13]
Unimodal thompson sampling for graph-structured arms,
S. Paladino, F. Trov `o, M. Restelli, and N. Gatti, “Unimodal thompson sampling for graph-structured arms,” inProceedings of the Thirty-First AAAI Conference on Artificial Intelligence, ser. AAAI’17. AAAI Press, 2017, p. 2457–2463
2017
-
[14]
Reinforcement Learning for Efficient and Tuning-Free Link Adaptation,
V . Saxena, H. Tullberg, and J. Jald ´en, “Reinforcement Learning for Efficient and Tuning-Free Link Adaptation,”IEEE Transactions on Wireless Communications, vol. 21, no. 2, pp. 768–780, 2022
2022
-
[15]
Rate Control With Spatial Reuse for Wi-Fi 6 Dense Deployments,
A. Krotov, A. Kiryanov, and E. Khorov, “Rate Control With Spatial Reuse for Wi-Fi 6 Dense Deployments,”IEEE Access, vol. 8, pp. 168 898–168 909, 2020
2020
-
[16]
On Ordered Beta Distribution and the Generalized Incomplete Beta Function,
M. Al-Saidi, A. Kuznetsov, and M. Nediak, “On Ordered Beta Distribution and the Generalized Incomplete Beta Function,”Method. Comput. Appl. Prob., vol. 27, no. 1, Dec. 2024
2024
-
[17]
py-itpp,
V . Saxena, “py-itpp,” https://github.com/vidits-kth/py-itpp, accessed: 2025-09-13
2025
-
[19]
On the Likelihood That One Unknown Probability Exceeds Another in View of the Evidence of Two Samples,
W. R. Thompson, “On the Likelihood That One Unknown Probability Exceeds Another in View of the Evidence of Two Samples,”Biometrika, vol. 25, no. 3/4, pp. 285–294, 1933
1933
-
[20]
Stochastic Relaxation, Gibbs Distributions, and the Bayesian Restoration of Images,
S. Geman and D. Geman, “Stochastic Relaxation, Gibbs Distributions, and the Bayesian Restoration of Images,”IEEE Transactions on pattern analysis and machine intelligence, no. 6, pp. 721–741, 1984
1984
-
[21]
Markov Chains for Exploring Posterior Distributions,
L. Tierney, “Markov Chains for Exploring Posterior Distributions,”the Annals of Statistics, pp. 1701–1728, 1994
1994
-
[22]
S. P. Meyn and R. L. Tweedie,Markov Chains and Stochastic Stability. Springer Science & Business Media, 2012
2012
Reviewed July 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.