REVIEW 4 major objections 6 minor 34 references
Semantic tokens can be scheduled like radio resources by treating token similarity as a measurable form of interference.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-12 01:54 UTC pith:UBISOJWY
load-bearing objection Solid SemCom systems paper: token-level multi-access under an explicit SSINR model is new enough to matter, but the gains rest on synthetic embeddings and fixed α coefficients. the 4 major comments →
ATS-ToDMA: Adaptive Token Selection and Token-Domain Multiple Access for Cross-Modal Semantic Communications
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
ATS-ToDMA establishes that jointly performing adaptive token selection, transformer-based interference-aware slot assignment, and semantic-aware power allocation under an SSINR reliability constraint yields substantially higher semantic throughput and decoding accuracy while reducing aggregate semantic interference and transmit power relative to OMA, Semantic NOMA, Random-TS, and Greedy ATS.
What carries the argument
Semantic Signal-to-Interference-plus-Noise Ratio (SSINR): each token’s useful power is divided by the sum of thermal noise and pairwise semantic-interference terms built from cosine similarity, modality-dependent coefficients, and channel-aware gates; this single metric drives both occupancy bounds and closed-form power allocation.
Load-bearing premise
The paper assumes that decoding ambiguity between tokens is well captured by their pairwise cosine similarity, scaled by fixed offline-calibrated coefficients into an equivalent interference power that can sit next to ordinary channel noise inside one reliability number.
What would settle it
Replace the synthetic embeddings with real foundation-model tokens (BERT/ViT/Wav2Vec-class) and a real multi-modal task decoder; if measured end-task accuracy no longer tracks the proposed SSINR (or the closed-form occupancy and power rules fail to keep accuracy above the design target), the central claim does not transfer.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes ATS-ToDMA, a cross-modal multi-user semantic communication framework that treats semantic tokens as the basic resource units. It jointly performs adaptive token selection (ATS), transformer-based interference-aware assignment of tokens to ToDMA slots, and a closed-form semantic-aware power allocation. Semantic interference is modeled via pairwise cosine similarity of unit embeddings, mapped into an equivalent interference power through offline-calibrated coefficients α_intra and α_cross, and folded into a Semantic SINR (SSINR) reliability metric. Analytical upper bounds on aggregate interference and feasible slot occupancy (Theorems 1–2) and a first-order Neumann approximation for power (Theorem 3) are derived. Simulations with synthetic cosine statistics report substantial gains in semantic throughput and decoding accuracy, and reductions in aggregate interference and transmit power, relative to OMA, Semantic NOMA, Random-TS, and Greedy ATS (Table IV).
Significance. If the SSINR model and the reported gains transfer beyond the synthetic setup, the work would be a meaningful systems contribution: it is among the first frameworks to treat semantic similarity as a schedulable interference resource, to couple token selection with multi-user token-domain multiple access, and to supply closed-form occupancy and power expressions under that model. The transformer scheduler, hybrid soft/hard constraint enforcement, and the explicit intra- vs cross-modal distinction are useful design ideas for multi-modal SemCom. The analytical bounds (standard but clean) and the approximate power formula give interpretable design guidelines. The main value is therefore architectural and modeling rather than a deep new theorem; significance hinges on whether SSINR remains a faithful reliability metric for real foundation-model embeddings and decoders.
major comments (4)
- [§III-B6 Eqs. (29)–(30a); §VI-B Eq. (50); Table IV] Semantic throughput is defined inconsistently in a load-bearing way. In §III-B6 Eq. (29) and the optimization (30a), R_s = |T|/T_slots (token count per slots). In the simulation section Eq. (50), R_s = ∑ x_i s_i log2(1+SSINR_i). These are inequivalent objectives. It is unclear which quantity the scheduler and power allocation actually optimize, and which quantity Table IV and Fig. 3(a) report. Align the problem statement, the training objective, and the reported metric, or explicitly justify and separate them.
- [Theorem 2; Appendix B; Eq. (27)] The proof of Theorem 2 is incomplete relative to the SSINR definition used. Appendix B bounds I_i by ∑ α_ij P ξ²_ij ≤ α_max δ² P(M−1), omitting the ‖g_j‖² factors present in Eq. (27), then inserts an unexplained factor d in the final M_max expression. If d is intended as the bound ‖g_j‖² ≤ d for g_j ∈ (0,1)^d, that step must be stated explicitly and the intermediate inequality corrected. As written, the occupancy bound does not follow rigorously from (27).
- [§III-B4; §VI (esp. Table IV, Figs. 3–6); §VI-I] All quantitative claims that support the central contribution—Table IV gains, Figs. 3–6, and the practical usefulness of Theorems 1–3—rest on synthetic cosine-similarity statistics and fixed hand-set coefficients α_intra = 0.8, α_cross = 0.4 (calibration §III-B4). No real BERT/ViT/Wav2Vec (or joint multi-modal) encoder–decoder is used, and no sensitivity sweep over α or over non-synthetic similarity distributions is reported. Because SSINR is the reliability metric that drives scheduling, occupancy, and power, the manuscript needs either experiments with actual foundation-model embeddings and a concrete task decoder, or a systematic sensitivity study plus a much sharper separation between model-driven analysis and empirical system gains. The brief Limitations paragraph (§VI-I) is not sufficient given how central these numbers are.
- [Fig. 3 caption vs. §VI-C; Fig. 5] Figure 3 is internally inconsistent: the body text (§VI-C) describes Fig. 3(b) as semantic decoding accuracy versus SNR, while the figure caption states that (b) validates Theorem 2 (occupancy). The actual Theorem 2 validation appears as Fig. 5(b). This is not cosmetic; it confuses which empirical result supports which claim. Correct captions and cross-references so that accuracy, throughput, and theorem validations are unambiguously identified.
minor comments (6)
- [Table III; Theorem 2; Appendix B] Notation collision: d is the embedding dimension and also appears as the occupancy-bound factor. Introduce a distinct symbol (e.g., d_g or g_max) for the ‖g‖² bound if that is the intended meaning.
- [Index Terms] Index terms list “LSTM” prominently while the scheduler and main architecture are transformer-based; align keywords with the actual technical core.
- [§IV-2] The relationship between soft training penalties (λ1, λ2, λ3 on ˆM_k, ˆI_k) and the frequency/severity of hard post-processing removals at inference is not quantified. A short ablation on how often capacity/interference pruning fires would strengthen the hybrid-scheduler claim.
- [Eq. (23)] Eq. (23) writes cross-modal sums with mixed index variables (e.g., ξ(τ_i, ψ_j) under a k sum); tidy the indices for readability.
- [References] Several references appear twice or with near-duplicate entries (e.g., Dosovitskiy et al. listed more than once). Clean the bibliography.
- [§III-A4; Eqs. (9)–(10), (13)] State explicitly whether embeddings remain unit-norm after channel-aware gating g_i ⊙ Z_i; if not, cosine similarity and the SSINR derivation need a brief note on renormalization.
Circularity Check
SSINR reliability is partly by construction: α_ij fitted from accuracy drop D_ij is re-inserted into the same SSINR used to enforce/predict reliability and drive gains.
specific steps
-
fitted input called prediction
[Section III-B4 (eq. 25) together with III-B3/III-B5 (eqs. 24, 27) and constraint (30f)]
"αij = Dij / Dref , where Dref denotes the accuracy degradation observed under the reference interference power Pref. Consequently, αij acts as a semantic-to-physical conversion factor, enabling semantic interference and thermal noise to be represented in a common SSINR framework. … SSINRi = Pi‖gi‖^{2} / (∑j eq i αij ho ij ξ^{2}ij ‖gj‖^{2} + N0) … Constraint (30f) ensures reliable semantic decoding by maintaining the SSINR of each transmitted token above the target threshold Γ."
α_ij is obtained by measuring the very accuracy drop D_ij that the metric is later claimed to protect; once inserted into SSINR, the constraint SSINR≥Γ and the subsequent claims of improved semantic decoding accuracy become true by the calibration construction (at least under the reference conditions used to fit α) rather than an independent first-principles prediction.
full rationale
The algebraic bounds (Theorems 1–2) and Neumann power approximation (Theorem 3) are non-circular consequences of the definitions of I_ij and SSINR under the stated max-similarity and equal-power assumptions; they do not reduce to fitted data. The sole mild circularity is the offline calibration that defines the load-bearing conversion factor α_ij from measured task-accuracy degradation and then re-uses that factor inside the SSINR reliability metric and optimization constraints. Simulations operate entirely inside this synthetic model (fixed α_intra=0.8/α_cross=0.4, cosine statistics matched to literature embeddings) so reported throughput/accuracy/power gains are model-internal rather than independent external predictions. No load-bearing self-citation chain or uniqueness import exists; prior author papers appear only in related-work tables. Score remains low because the circular step is confined to metric interpretation, not the derivation of the closed-form results themselves.
Axiom & Free-Parameter Ledger
free parameters (8)
- α_intra =
0.8
- α_cross =
0.4
- γ (ATS/similarity threshold) =
0.5
- δ (max semantic similarity) =
0.9
- Γ (target SSINR) =
2
- τ_ATS
- λ1, λ2, λ3 (scheduler penalty weights)
- P_ref / N0 normalization =
N0=1
axioms (5)
- domain assumption Pairwise cosine similarity of unit-norm embeddings, possibly thresholded by γ, quantifies semantic interference relevant to task decoding.
- ad hoc to paper Cross-modal interactions induce strictly less effective semantic distortion than intra-modal ones (α_cross < α_intra).
- ad hoc to paper Semantic accuracy loss under an interfering token can be linearly mapped to an equivalent interference power via α_ij = D_ij/D_ref.
- standard math Equal-power and max-similarity bounds plus first-order Neumann expansion (I−F)⁻¹≈I+F yield useful design formulas when coupling is weak.
- domain assumption Rayleigh fading plus AWGN and synthetic BERT/ViT/Wav2Vec-like cosine statistics adequately stress-test the framework.
invented entities (4)
-
SSINR (Semantic Signal-to-Interference-plus-Noise Ratio)
no independent evidence
-
ToDMA (Token-Domain Multiple Access) slots
no independent evidence
-
Modality-dependent semantic distortion coefficients α_intra / α_cross
no independent evidence
-
Channel-aware semantic protection gate g_i
no independent evidence
read the original abstract
Adaptive token processing has emerged as a promising approach for improving the efficiency of semantic communication systems. However, existing semantic communication frameworks largely overlook token-level multiple access and the impact of semantic interference among simultaneously transmitted semantic tokens. In this paper, we propose Adaptive Token Selection and Token-Domain Multiple Access (ATS-ToDMA), a novel cross-modal semantic communication framework that jointly performs semantic token selection, interference-aware scheduling, and semantic-aware power allocation. The proposed framework introduces a Semantic Signal-to-Interference-plus-Noise Ratio (SSINR) metric that captures the combined effects of channel impairments and semantic interference arising from token similarity. A transformer-based scheduler is developed to allocate selected semantic tokens across token-domain transmission slots while mitigating both intra-modal and cross-modal semantic interference. To characterize the behavior of the proposed system, analytical bounds on semantic interference and feasible token occupancy are derived, together with a closed-form approximation for semantic-aware power allocation. Simulation results demonstrate significant gains in semantic throughput and semantic decoding accuracy while reducing aggregate semantic interference and transmit power compared with OMA, Semantic NOMA, Random-TS, and Greedy ATS benchmarks.
Figures
Reference graph
Works this paper leans on
-
[1]
A mathematical theory of communication,
C. E. Shannon, “A mathematical theory of communication,”Bell System Technical Journal, vol. 27, pp. 379–423, 1948
1948
-
[2]
Deep learning enabled semantic communication systems,
H. Xie, Z. Qin, G. Y . Li, and B.-H. Juang, “Deep learning enabled semantic communication systems,”IEEE transactions on signal pro- cessing, vol. 69, pp. 2663–2675, 2021
2021
-
[3]
6G: The next frontier: From holographic messaging to artificial intelligence using subterahertz and visible light communication,
E. C. Strinati, S. Barbarossa, J. L. Gonzalez-Jimenez, D. Ktenas, N. Cassiau, L. Maret, and C. Dehos, “6G: The next frontier: From holographic messaging to artificial intelligence using subterahertz and visible light communication,”IEEE Vehicular Technology Magazine, vol. 14, no. 3, pp. 42–50, 2019
2019
-
[4]
Semantic communication: A survey of its theoretical development,
G. Xin, P. Fan, and K. B. Letaief, “Semantic communication: A survey of its theoretical development,”Entropy, vol. 26, no. 2, p. 102, 2024
2024
-
[5]
Less data, more knowledge: Building next-generation semantic communication networks,
C. Chaccour, W. Saad, M. Debbah, Z. Han, and H. V . Poor, “Less data, more knowledge: Building next-generation semantic communication networks,”IEEE Communications Surveys & Tutorials, vol. 27, no. 1, pp. 37–76, 2024
2024
-
[6]
Deep Joint Source- Channel Coding for Wireless Image Transmission,
E. Bourtsoulatze, D. Burth Kurka, and D. G ¨und¨uz, “Deep Joint Source- Channel Coding for Wireless Image Transmission,”IEEE Transactions on Cognitive Communications and Networking, vol. 5, no. 3, pp. 567– 579, 2019
2019
-
[7]
DeepJSCC-f: Deep Joint Source-Channel Coding of Images With Feedback,
D. B. Kurka and D. G ¨und¨uz, “DeepJSCC-f: Deep Joint Source-Channel Coding of Images With Feedback,”IEEE Journal on Selected Areas in Information Theory, vol. 1, no. 1, pp. 178–193, 2020
2020
-
[8]
A robust image semantic communication system with multi-scale vision transformer,
X. Peng, Z. Qin, X. Tao, J. Lu, and K. B. Letaief, “A robust image semantic communication system with multi-scale vision transformer,” IEEE Journal on Selected Areas in Communications, vol. 43, no. 4, pp. 1278–1291, 2025
2025
-
[9]
BERT: Pre-training of deep bidirectional transformers for language understanding,
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers for language understanding,” inPro- ceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologies, volume 1 (long and short papers), 2019, pp. 4171–4186
2019
-
[10]
Language mod- els are few-shot learners,
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askellet al., “Language mod- els are few-shot learners,”Advances in neural information processing systems, vol. 33, pp. 1877–1901, 2020
1901
-
[12]
Learning transferable visual models from natural language supervision,
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clarket al., “Learning transferable visual models from natural language supervision,” inInternational conference on machine learning. PmLR, 2021, pp. 8748–8763
2021
-
[13]
Dynam- icvit: Efficient vision transformers with dynamic token sparsification,
Y . Rao, W. Zhao, B. Liu, J. Lu, J. Zhou, and C.-J. Hsieh, “Dynam- icvit: Efficient vision transformers with dynamic token sparsification,” Advances in neural information processing systems, vol. 34, pp. 13 937– 13 949, 2021
2021
-
[14]
Tokenlearner: What can 8 learned tokens do for images and videos?
M. S. Ryoo, A. Piergiovanni, A. Arnab, M. Dehghani, and A. Angelova, “Tokenlearner: What can 8 learned tokens do for images and videos?” arXiv preprint arXiv:2106.11297, 2021
Pith/arXiv arXiv 2021
-
[15]
Evo-ViT: Slow-fast token evolution for dynamic vision transformer,
Y . Xu, Z. Zhang, M. Zhang, K. Sheng, K. Li, W. Dong, L. Zhang, C. Xu, and X. Sun, “Evo-ViT: Slow-fast token evolution for dynamic vision transformer,” inProceedings of the AAAI conference on artificial intelligence, vol. 36, no. 3, 2022, pp. 2964–2972
2022
-
[16]
Cosine similarity to determine similarity measure: Study case in online essay assessment,
A. R. Lahitani, A. E. Permanasari, and N. A. Setiawan, “Cosine similarity to determine similarity measure: Study case in online essay assessment,” in2016 4th International conference on cyber and IT service management. IEEE, 2016, pp. 1–6
2016
-
[17]
A survey on non-orthogonal multiple access for 5G networks: Research challenges and future trends,
Z. Ding, X. Lei, G. K. Karagiannidis, R. Schober, J. Yuan, and V . K. Bhargava, “A survey on non-orthogonal multiple access for 5G networks: Research challenges and future trends,”IEEE Journal on Selected Areas in Communications, vol. 35, no. 10, pp. 2181–2195, 2017
2017
-
[18]
Attention is all you need,
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,”Advances in neural information processing systems, vol. 30, 2017
2017
-
[19]
IoT-Enabled Traffic Management System Using Vehicle Count Prediction in a Semantic Communication Frame- work,
S. Kadam and D. I. Kim, “IoT-Enabled Traffic Management System Using Vehicle Count Prediction in a Semantic Communication Frame- work,”IEEE Internet of Things Journal, vol. 12, no. 17, pp. 36 258– 36 273, 2025
2025
-
[20]
Knowledge-aware semantic communication system design,
——, “Knowledge-aware semantic communication system design,” in ICC 2023-IEEE International Conference on Communications. IEEE, 2023, pp. 6102–6107
2023
-
[21]
Knowledge-aware semantic communication system design and data allocation,
——, “Knowledge-aware semantic communication system design and data allocation,”IEEE Transactions on Vehicular Technology, vol. 73, no. 4, pp. 5755–5769, 2023
2023
-
[22]
Semantic communication-empowered vehicle count prediction for traffic management,
——, “Semantic communication-empowered vehicle count prediction for traffic management,” in2024 IEEE Wireless Communications and Networking Conference (WCNC). IEEE, 2024, pp. 1–6
2024
-
[23]
An image is worth 16x16 words: Transformers for image recognition at scale,
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gellyet al., “An image is worth 16x16 words: Transformers for image recognition at scale,”arXiv preprint arXiv:2010.11929, 2020
Pith/arXiv arXiv 2010
-
[24]
Long short-term memory,
S. Hochreiter and J. Schmidhuber, “Long short-term memory,”Neural computation, vol. 9, no. 8, pp. 1735–1780, 1997
1997
-
[25]
Adaptive semantic token selection for ai-native goal-oriented commu- nications,
A. Devoto, J. Pomponi, S. Petruzzi, P. Di Lorenzo, and S. Scardapane, “Adaptive semantic token selection for ai-native goal-oriented commu- nications,” in2024 IEEE Globecom Workshops (GC Wkshps). IEEE, 2024, pp. 1–6
2024
-
[26]
Efficient trans- formers with dynamic token pooling,
P. Nawrot, J. Chorowski, A. Lancucki, and E. M. Ponti, “Efficient trans- formers with dynamic token pooling,” inProceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2023, pp. 6403–6417
2023
-
[27]
Semantic communications for future internet: Fundamentals, applications, and challenges,
W. Yang, H. Du, Z. Q. Liew, W. Y . B. Lim, Z. Xiong, D. Niyato, X. Chi, X. Shen, and C. Miao, “Semantic communications for future internet: Fundamentals, applications, and challenges,”IEEE Communications Surveys & Tutorials, vol. 25, no. 1, pp. 213–250, 2022
2022
-
[28]
A lite distributed semantic communication system for internet of things,
H. Xie and Z. Qin, “A lite distributed semantic communication system for internet of things,”IEEE Journal on Selected Areas in Communica- tions, vol. 39, no. 1, pp. 142–153, 2020
2020
-
[29]
Semantic communication meets edge intelligence,
W. Yang, Z. Q. Liew, W. Y . B. Lim, Z. Xiong, D. Niyato, X. Chi, X. Cao, and K. B. Letaief, “Semantic communication meets edge intelligence,” IEEE wireless communications, vol. 29, no. 5, pp. 28–35, 2022
2022
-
[30]
Goal- oriented semantic communications for 6g networks,
H. Zhou, Y . Deng, X. Liu, N. Pappas, and A. Nallanathan, “Goal- oriented semantic communications for 6g networks,”IEEE Internet of Things Magazine, vol. 7, no. 5, pp. 104–110, 2024
2024
-
[31]
Atp-llava: Adaptive token pruning for large vision language models,
X. Ye, Y . Gan, Y . Ge, X.-P. Zhang, and Y . Tang, “Atp-llava: Adaptive token pruning for large vision language models,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2025, pp. 24 972–24 982
2025
-
[32]
Non-orthogonal multiple access enhanced multi-user semantic communication,
W. Li, H. Liang, C. Dong, X. Xu, P. Zhang, and K. Liu, “Non-orthogonal multiple access enhanced multi-user semantic communication,”IEEE Transactions on Cognitive Communications and Networking, vol. 9, no. 6, pp. 1438–1453, 2023
2023
-
[33]
A joint communication and computation design for probabilistic semantic com- munications,
Z. Zhao, Z. Yang, M. Chen, Z. Zhang, and H. V . Poor, “A joint communication and computation design for probabilistic semantic com- munications,”Entropy, vol. 26, no. 5, p. 394, 2024
2024
-
[34]
A survey on large language model (LLM) security and privacy: The good, the bad, and the ugly,
Y . Yao, J. Duan, K. Xu, Y . Cai, Z. Sun, and Y . Zhang, “A survey on large language model (LLM) security and privacy: The good, the bad, and the ugly,”High-Confidence Computing, vol. 4, no. 2, p. 100211, 2024
2024
-
[35]
Tokens-to-token ViT: Training vision transformers from scratch on imagenet,
L. Yuan, Y . Chen, T. Wang, W. Yu, Y . Shi, Z.-H. Jiang, F. E. Tay, J. Feng, and S. Yan, “Tokens-to-token ViT: Training vision transformers from scratch on imagenet,” inProceedings of the IEEE/CVF international conference on computer vision, 2021, pp. 558–567
2021
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.