{"id":"e6e3ed12-b966-4247-999e-3d2f4a92bbe8","arxiv_id":"2505.10946","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"ToDMA lets many devices share one wireless channel by transmitting token indices from a common codebook, recovering collisions with compressed sensing and masked-token prediction from pretrained models.","lead":"A new wireless access method lets many devices transmit compact 'tokens' on the same channel at the same time. The receiver separates the overlapping signals and uses pretrained AI to repair tokens lost to collisions.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The advertised 4x latency gain over Orth-Com does not follow from the paper's own rate formula at SNR=25 dB; at reasonable BERs the gain is closer to 1.5-2x.","rationale":"The reader's verdict is CONDITIONAL, and my analysis supports keeping that verdict: the paper's novel ToDMA architecture and its simulation evidence for collision mitigation are not invalidated, but the headline quantitative claim of a 4x latency reduction appears inconsistent with the paper's own rate formula at the stated SNR. I therefore agree with the reader's overall conditional assessment, though I focus on a different load-bearing point than the reader's stated weakest assumption. The reader did flag the 4x latency inconsistency in the rationale, so there is partial agreement. The suggested concrete test is purely analytic and does not require new experiments; it isolates whether the latency claim survives a correct use of Eq. (35). If the 4x claim is corrected to roughly 1.5-2x, the central qualitative claim of lower latency may still hold, but the magnitude and the specific '4 times' statement in the contributions need revision, so the paper should be accepted only after this correction or clarification.","tokens_in":19888,"tokens_out":13103,"duration_ms":144188,"concrete_test":"Recompute the Fig. 10 latency comparison using Eq. (35) with SNR=10^(25/10), the stated K values (20, 40, 60, 80, and K=KT), the same BER range, L=K+1, N=256, Q=1024, and B=10 MHz. Report the Orth-Com-to-ToDMA latency ratio as a function of K and BER. If the maximum ratio is below 2.5 across all plotted BERs, the '4 times lower' claim in Section III should be revised. Also check whether the original curve was generated with SNR=25 (linear) rather than 25 dB; if so, that is a dB/linear unit error and explains the discrepancy.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section VIII-D compares Orth-Com latency KN log2(Q)/ROrth with ToDMA latency LN/B, using ROrth from Eq. (35) and setting L=K+1, N=256, Q=1024, SNR=25 dB, B=10 MHz. With SNR=25 dB (linear SNR about 316) and BER=1e-3, Eq. (35) gives spectral efficiency log2(1+1.5*316/5.30) about 6.5 b/s/Hz. Then for K=20, Orth-Com latency is 20*256*10/(10e6*6.5) about 0.79 ms, whereas ToDMA latency is 21*256/10e6 about 0.54 ms; the ratio is about 1.5, not 4. Even at BER=1e-9 the spectral efficiency is about 4.7 b/s/Hz and the ratio is about 2.0. The only way Eq. (35) yields a ratio near 4 at the stated parameters is to insert SNR=25 as a linear value (14 dB) rather than 10^(25/10), or to assume an effective spectral efficiency near 2.5 b/s/Hz, which is inconsistent with 25 dB AWGN for uncoded MQAM. Since the 4x latency reduction is stated as a headline contribution in Section III and is the quantitative basis of Fig. 10, the central latency claim is not supported by the paper's own equations.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces ToDMA, a token-domain multiple access scheme for massive uncoordinated uplink transmissions. Each active device tokenizes its source signal (text or image) and modulates every token with a codeword from a common Gaussian modulation codebook. At the receiver, an AMP-based algorithm detects active tokens and estimates their associated CSI per time slot; a clustering step assigns tokens to devices; residual token collisions leave masked positions, which are filled by candidate-restricted masked-token prediction using pretrained bidirectional transformers (BERT for text, MaskGIT for images). Simulations on ImageNet-100 and QUOTES500K show that ToDMA reduces token error rate and improves PSNR/LPIPS/BERTScore relative to a context-unaware non-orthogonal baseline, and the paper claims a 4-times latency reduction relative to an orthogonal adaptive-QAM baseline. The receiver complexity is reported to be linear in K and M except for the transformer-based prediction stage.","tokens_in":20205,"tokens_out":6758,"duration_ms":63602,"significance":"The core idea—using contextual redundancy in token sequences as a collision-resolution mechanism for unsourced multiple access—is novel and timely, and the empirical validation with standard pretrained models is a strength. If the latency and quality claims survive scrutiny, ToDMA would be a practical interface between token-based source coding and unsourced random access. The paper also provides a useful AMP-based detection formulation and a clear complexity breakdown. However, the headline latency gain does not follow from the paper's own rate formula, and two structural assumptions (known number of active devices and perfect token detection) are not fully supported; these issues need to be resolved before the paper can be recommended for publication.","major_comments":[{"comment":"The claimed 4-times latency reduction over Orth-Com is not supported by the paper's own equations. Inserting SNR = 25 dB (linear SNR about 316), BER = 1e-3, K = 20, N = 256, Q = 1024, and B = 10 MHz into Eq. (35) gives ROrth ≈ 10 MHz × log2(1 + 1.5·316/5.30) ≈ 65 Mb/s, so the Orth-Com latency is 20·256·10/(65e6) ≈ 0.79 ms, whereas the ToDMA latency is 21·256/(10e6) ≈ 0.54 ms; the ratio is about 1.5, not 4. Even at BER = 1e-9 the ratio is about 2.0. Since the 4-times statement appears as a contribution in Section III and underlies Fig. 10, the latency comparison must be recomputed and the claimed gain revised to a value consistent with Eq. (35), or the comparison setup must be re-specified (for example, by including Orth-Com's signaling overhead) and defended.","section":"Section VIII-D, Eq. (35), Fig. 10"},{"comment":"The token-assignment stage requires the number of active devices K as an input to K-means++ clustering. In the grant-free unsourced scenario of Section IV-B, K is not known at the receiver; the paper neither provides an activity-count estimator nor analyzes the sensitivity of clustering and subsequent token assignment to K mismatch. Since all later masked-token recovery depends on this clustering, the paper should either introduce an estimation procedure for K or clearly state a separate assumption and evaluate the impact of K mismatch on the reported performance.","section":"Section VI-A, problem (25)"},{"comment":"The fine-grained assignment is developed under the assumption that token detection is perfect (bPn = Pn), justified by the claim that token detection error approaches zero as the number of antennas M grows. However, the main simulations use M = 256, where Fig. 5(a) shows nonzero TDER; the effect of detection errors on the candidate token sets and on the final TER/PSNR is not quantified. Please either analyze the imperfect-detection case or restrict the corresponding claims to the perfect-detection regime.","section":"Section VI-B.2, after Eq. (28)"},{"comment":"The notion of 'token-domain semantic orthogonality' is introduced only by example and is never defined formally. The entire advantage of ToDMA over the context-unaware baseline rests on the ability of BERT/MaskGIT to resolve masked positions from context, so the paper should state a quantitative condition (for example, in terms of the conditional entropy of tokens given their context, or a minimum prediction accuracy threshold) under which the proposed recovery is effective, and it should discuss regimes where the contextual redundancy is insufficient.","section":"Section I-C and Section VI-C"}],"minor_comments":[{"comment":"In the update for R^t_{q,m}, the codebook coefficient is written as u^*_{l,m}; it should be u^*_{l,q}.","section":"Section V-B, Eq. (13)"},{"comment":"The initialization line contains a stray semicolon and comma ('bBk = 0Q×N , ;'); please clean up the pseudocode.","section":"Algorithm 1, line 4"},{"comment":"The formula for ROrth is typeset ambiguously: the placement of SNR in the denominator makes it appear that spectral efficiency decreases with SNR at fixed BER; please rewrite it as ROrth = B log2(1 + 1.5·SNR / (−ln(5·BER))).","section":"Eq. (35)"},{"comment":"The statement 'we assume K = KT' contradicts the massive-access premise that only a small fraction of devices are active (K ≪ KT); if all devices are active, the sparsity argument in Eq. (2) and the 'massive' characterization need to be revisited.","section":"Section VIII-D"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is closely related to the authors' own INFOCOM 2025 Workshop paper (ref. [1]) and to their TokenCom preprint (ref. [15]); the editor may wish to confirm that the journal submission contains sufficient additional material beyond those works. In addition, the 4-times latency gain is repeated in the contribution list and in Fig. 10; if it cannot be supported after recomputation, the headline claim should be revised before acceptance."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know two things before reading. First, the core token-domain multiple access mechanism is already in their INFOCOM workshop paper [1]; this manuscript extends it with text experiments, complexity analysis, and a latency comparison. Second, the advertised 4x latency reduction over Orth-Com does not survive contact with their own equation. Using their stated SNR=25 dB as 25 dB, Eq. (35) gives about 6.5 b/s/Hz at BER=1e-3, so Orth-Com latency is roughly 0.79 ms versus ToDMA's 0.54 ms: a ratio around 1.5, not 4. Only if SNR=25 is treated as a linear value (14 dB) does the ratio approach 4. Since the 4x figure appears in Section III and is the basis of Fig. 10, the central quantitative claim is currently unsupported.\n\nWhat is genuinely good: the pipeline is clean and the simulations are plausible. Shared token codebook plus shared modulation codebook, AMP-based active token detection, CSI clustering, and candidate-restricted masked token prediction is a sensible combination. The image and text results show that ToDMA recovers most collided tokens, and the complexity analysis is useful. The paper uses standard external components (AMP, BERT, MaskGIT) with no apparent tuning-to-target, and it explicitly states the perfect-detection assumption used in Section VI.\n\nSoft spots beyond latency: the clustering step takes K as given, which is a real issue for massive unsourced random access; 'semantic orthogonality' is a metaphor rather than a defined quantity, and the method is only tested on highly redundant natural data, so the generalization claim is weaker than the text suggests. Also, the 'first paper to propose multiple access in the token domain' claim is hard to square with their own earlier workshop paper, even though the overlap is transparently acknowledged.\n\nBottom line: the core simulation results probably hold, and the paper deserves a serious referee. But the latency claim must be corrected or substantially qualified, and the known-K assumption needs discussion. Send it to review, with a referee asked to redo the latency calculation and to probe how the scheme behaves when K is unknown.","headline":"Worth reading, but the headline 4x latency gain does not follow from the paper's own formula, and the core scheme is already in the authors' workshop paper.","tokens_in":20721,"tokens_out":3166,"would_cite":true,"duration_ms":34668,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"ToDMA, a token-domain multiple access scheme, lets many uncoordinated devices share uplink resources by transmitting token indices, while pretrained transformers at the receiver repair collision-lost tokens, cutting latency fourfold.","keywords":["semantic communications","token communications","unsourced random access","multiple access","approximate message passing","masked token prediction","compressed sensing","massive MIMO"],"falsifier":"Feed ToDMA a source whose tokens are independent and uniformly distributed over the codebook, for example random strings or randomly shuffled image tokens, and run the default settings ($K=40$, $M=256$, $L=K+1$) at SNR $=25$ dB; if TER and PSNR/LPIPS match the context-unaware Non-Orth Com baseline, the masked-token prediction is doing no work and the semantic-orthogonality premise is falsified.","tokens_in":19719,"feed_emoji":"📡","tokens_out":8182,"duration_ms":66760,"temperature":0.7,"pith_summary":"ToDMA asks whether token-based processing can turn the uplink of a massive access network into a single shared token channel. The paper claims yes: when many uncoordinated devices transmit tokenized image or text sources over the same time-frequency resources, a receiver can first detect the active tokens by compressed sensing, then use the consistency of each device's channel across time slots to assign tokens to devices, and finally let a pretrained bidirectional transformer fill in tokens lost to collisions. The payoff, if true, is a grant-free non-orthogonal multiple access scheme whose latency is four times lower than an orthogonal token-communication baseline at comparable reconstruction quality. This matters because massive machine-type traffic in future networks needs low-latency access with minimal coordination, and token-domain redundancy offers a new resource to exploit.","feed_headline":"Token-domain multiple access cuts uplink latency fourfold","feed_subtitle":"Uncoordinated devices share a token codebook and a pretrained transformer repairs collision damage on images and text.","key_machinery":"The load-bearing mechanism is the shared token-modulation codebook $\\mathbf{U}\\in\\mathbb{C}^{L\\times Q}$, which turns each of the $Q$ token indices into a fixed length-$L$ codeword so that the received signal is a sparse superposition; the receiver's three-stage pipeline then converts that sparsity into reconstructed sequences: approximate message passing for active-token and CSI estimation, K-means++ clustering of per-token CSI across $N$ time slots to assign tokens to the $K$ devices, and candidate-restricted masked-token prediction by a pretrained bidirectional transformer to fill collision-induced gaps. The paper also names an enabling premise, 'semantic orthogonality': the contextual redundancy of natural text and images must be strong enough that a pretrained model can distinguish and recover a device's token sequence from the mixed and partially missing token stream.","core_discovery":"The central discovery is that token-domain semantic orthogonality can be turned into a multiple access dimension. Each token index from a shared codebook is mapped to a shared modulation codeword, so the superposition over the wireless channel is a sparse linear mixture of codewords. The receiver runs an AMP-based estimator to recover the active token set and per-token CSI in each time slot, clusters the CSI to assign tokens to devices, and uses pretrained BERT or MaskGIT to predict masked positions caused by token collisions. Simulations on ImageNet-100 and QUOTES500K show that ToDMA keeps token error rates and perceptual quality close to an error-free orthogonal baseline while cutting latency by a factor of four, and outperforms a context-unaware non-orthogonal baseline that randomly guesses collided tokens.","pith_inferences":["Inference: the decisive test is the strength of contextual redundancy. On low-entropy or adversarial sources such as random identifiers or shuffled sensor logs, the masked-token predictor will have nothing to condition on, so ToDMA should collapse to the context-unaware baseline; the paper does not quantify how much redundancy is needed.","Inference: ToDMA is effectively an unsourced random access code over token sequences, so its throughput and latency could be compared against information-theoretic bounds for unsourced multiple access; the paper stops short of that comparison.","Inference: the slow-fading assumption, channel vectors constant across all $N$ token slots, is likely the practical bottleneck; under mobility the CSI clustering step would degrade, so a Doppler-robust variant would be a natural follow-up."],"forward_implications":["If ToDMA's claims hold, uplink access for massive IoT can be grant-free and non-orthogonal: devices transmit when they have data, and the receiver separates them using token structure rather than per-device preambles.","The fourfold latency reduction over Orth-Com comes with comparable or better distortion and perceptual quality in the tested image and text tasks, so orthogonal token transmission may be unnecessary when sources are contextually redundant.","Receiver complexity scales linearly with the number of active devices and receive antennas, decoupling the large tokenizer dimension $Q$ from the number of devices, which is desirable for massive MIMO.","As the number of receive antennas $M$ grows, token detection error approaches zero without increasing the codeword length $L$, meaning detection accuracy can be bought with more antennas rather than more communication overhead.","ToDMA functions as a joint source-channel code in which the source's contextual redundancy is used to repair channel collisions, so stronger contextual models should further reduce token error."],"supporting_citations":[{"why":"Introduces token communications, the framework ToDMA builds on.","marker":"[15]"},{"why":"Provides the pretrained BERT bidirectional transformer used for masked token prediction in text.","marker":"[10]"},{"why":"Provides MaskGIT, the masked image transformer used for image token recovery.","marker":"[13]"},{"why":"Supplies the VQ-GAN tokenizer that maps images to the 1024-entry token codebook used in simulations.","marker":"[14]"},{"why":"Defines the unsourced random access paradigm that ToDMA extends into the token domain.","marker":"[51]"},{"why":"Establishes the AMP decoupling result underlying the active token detection algorithm.","marker":"[60]"},{"why":"Provides generalized approximate message passing used for posterior mean estimation in detection.","marker":"[61]"},{"why":"K-means++ clustering is used for coarse token assignment from estimated CSI.","marker":"[63]"}],"fun_headline_variants":["ToDMA: 4x faster uplink with token-sharing and AI collision repair","Token-domain access: shared codebooks cut latency 4x","Massive token communications: AI fixes collisions, cuts latency 4x","ToDMA: Pretrained models fix token collisions, 4x speedup","Token-domain multiple access: 4x lower latency with AI repair"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The recovery pipeline assumes the transmitted token sequences are predictable from their context, so that a pretrained transformer can reliably fill in tokens lost to collisions; if the source tokens carry little contextual redundancy, the whole advantage over a context-unaware baseline disappears.","fun_headline_variants_meta":{"raw":{"variants":["ToDMA: 4x faster uplink with token-sharing and AI collision repair","Token-domain access: shared codebooks cut latency 4x","Massive token communications: AI fixes collisions, cuts latency 4x","ToDMA: Pretrained models fix token collisions, 4x speedup","Token-domain multiple access: 4x lower latency with AI repair"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001285,"raw_usage":{"total_tokens":5248,"prompt_tokens":943,"completion_tokens":4305,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":559,"completion_tokens_details":{"reasoning_tokens":4207}},"tokens_in":559,"tokens_out":4305,"duration_ms":32007,"temperature":1.0,"reasoning_tokens":4207,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T21:02:30.338877+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Feed ToDMA a source whose tokens are independent and uniformly distributed over the codebook, for example random strings or randomly shuffled image tokens, and run the default settings ($K=40$, $M=256$, $L=K+1$) at SNR $=25$ dB; if TER and PSNR/LPIPS match the context-unaware Non-Orth Com baseline, the masked-token prediction is doing no work and the semantic-orthogonality premise is falsified.","supporting_citations":[{"cited_title":"Unsourced multiple access: A coding paradigm for massive random access,","cited_arxiv_id":null,"evidence_quote":"Defines the unsourced random access paradigm that ToDMA extends into the token domain."},{"cited_title":"MaskGIT: Masked generative image transformer,","cited_arxiv_id":null,"evidence_quote":"Provides MaskGIT, the masked image transformer used for image token recovery."},{"cited_title":"Taming transformers for high- resolution image synthesis,","cited_arxiv_id":null,"evidence_quote":"Supplies the VQ-GAN tokenizer that maps images to the 1024-entry token codebook used in simulations."},{"cited_title":"Message passing algo- rithms for compressed sensing: I. motivation and construction,","cited_arxiv_id":null,"evidence_quote":"Establishes the AMP decoupling result underlying the active token detection algorithm."},{"cited_title":"Generalized approximate message passing for estimation with random linear mixing,","cited_arxiv_id":null,"evidence_quote":"Provides generalized approximate message passing used for posterior mean estimation in detection."},{"cited_title":"k-means++: the advantages of careful seeding,","cited_arxiv_id":null,"evidence_quote":"K-means++ clustering is used for coarse token assignment from estimated CSI."}],"review_version":1}