{"id":"5c8129a6-979f-461e-a070-b5edb7253e71","arxiv_id":"2502.03949","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":2,"one_line_summary":"The paper proposes SFDMA, a learned multiple access scheme for digital semantic broadcast channels, and evaluates it on inference and image reconstruction tasks.","lead":"A multi-user broadcast scheme encodes each user's semantic features into nearly orthogonal digital signals so all users can transmit simultaneously over the same time-frequency resource, then fits a performance-versus-SINR curve to allocate power. The simulations show gains over a deep JSCC baseline, but the theoretical support is mostly empirical and several claims are overstated.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Approximate orthogonality is never enforced or analyzed; SFDMA's interference and privacy claims rest on an unverified spontaneous-emergence assumption tested only at N=2,3 and at operating points different from the main performance curves.","rationale":"The reader's weakest_assumption identifies exactly the same load-bearing point: approximate orthogonality is expected to emerge without an explicit constraint, and the central interference and privacy claims depend on it. I agree with that assessment. The experiments give some support, but they are narrow: two or three users, one dataset per task, one training SNR for the main curves, and orthogonality measured at qbit values not matched to the performance curves. The added privacy flaw strengthens the concern: the cross-decoding test uses a clean interfering signal rather than the superposition actually received, so the 'each receiver can only decode its own semantic information' claim is not established even where orthogonality is observed. The other issues raised by the reader, such as the ABG function being an empirical fit and the RIB derivation containing notation errors, are real but secondary; they affect the power-allocation and theoretical contributions, not the core SFDMA mechanism. Since the most load-bearing condition is unproven and the available evidence is too narrow to establish it, I do not see a basis for changing the reader's rejection. A focused N-scaling and operating-point experiment would be the check that could overturn or confirm this concern.","tokens_in":17789,"tokens_out":9538,"duration_ms":110512,"concrete_test":"Train the exact Algorithm 1 on MNIST and Algorithm 2 on CelebA with N=5 users (and N=3 as a control), using at least 3 random seeds, with training SNR 0 dB and with qbit = 32/64 for inference and qbit = 4096 for reconstruction. On a held-out set, report per-user accuracy/MS-SSIM, the maximum normalized pairwise inner product max_{i!=j} |x_i^H x_j|/(||x_i|| ||x_j||), and cross-decoder accuracy on the actual superposition y1 = sqrt(p1)x1 + sqrt(p2)x2 + noise rather than on x2 alone. If max inner product exceeds 1e-2 or per-user performance drops by more than the seed spread relative to the Upper Bound lines in Figs. 6/10, the spontaneous-orthogonality mechanism is the failure point.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's headline mechanism is Section II-B, Eq. (8): encoded BPSK signals x_i are approximately orthogonal so that simultaneous transmission in the same time-frequency resource is interference-free and each receiver can decode only its own semantics. Nothing in the RIB loss (Eq. (16)) or the MSE loss (Eq. (29)) includes a pairwise orthogonality term. Section II-C, after Algorithm 1, explicitly relies on the hope that joint training will 'drive them toward orthogonality.' That is an empirical regularity claim, not a design invariant. The supporting measurements in Tables IV and VI are reported for qbit = 128 and 4096 bits, whereas the accuracy/PSNR curves in Figs. 6 and 10 are produced at qbit = 32/64 and under a different training SNR; no orthogonality metric is given at the actual operating points of the headline curves. Only two- and three-user cases are tested, with no seeds or error bars. If the mechanism fails for larger N, heterogeneous user tasks, or mismatched channel statistics, Eq. (8) collapses and with it the interference-mitigation and 'only decode own information' claims. The privacy claim has an additional logical gap: Tables V and VII feed a clean x2 (or x1) to the other user's decoder, but the real received signal y1 is a superposition. Showing that f_theta1(x2) is near chance does not show that y1 leaks no information about user 2's semantics; a decoder trained or adapted on y1 could in principle extract it. These gaps are the load-bearing weak point of the central claim.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes SFDMA (semantic feature division multiple access) for multi-user digital semantic broadcast networks. The transmitter encodes each user's source into a discrete approximately orthogonal BPSK feature vector, and all users' vectors are broadcast simultaneously in the same time-frequency resource; each receiver is expected to decode only its own semantic information. For inference tasks the authors design a robust information bottleneck (RIB) objective, and for image reconstruction they use a Swin Transformer with an MSE loss. The paper further fits an Alpha-Beta-Gamma (ABG) curve relating task performance to SINR and uses it to formulate a linear power allocation problem. Experiments on MNIST and CelebA compare the proposed SFDMA with Deep JSCC and an upper bound, with t-SNE plots and inner-product/angle tables for orthogonality, along with CDFs for the power allocation scheme.","tokens_in":18105,"tokens_out":6347,"duration_ms":59905,"significance":"If validated, the SFDMA mechanism would be a useful contribution to semantic broadcast: it addresses multi-user interference without SIC and offers a form of semantic privacy, and the ABG-based power allocation would give a practical QoS control tool. The paper's strengths are its clear system model, the use of quantitative orthogonality metrics (inner products, angles), and the comparison against JSCC baselines and an upper bound. However, the central design assumption of spontaneous orthogonality is not enforced or analyzed, the privacy test does not match the actual received signal, and the ABG 'optimal' power allocation is fitted to simulation data from the same system on which it is evaluated. These gaps currently limit the claims to a specific simulation setup and prevent the paper from supporting its stated generalizations.","major_comments":[{"comment":"The privacy claim 'each receiver can only decode its own semantic information' is not supported by the reported experiments. In Tables V and VII the cross-decoding inputs are the clean codewords x2 or x1, but the actual received signal y_i in Eq. (5) is a superposition of x_i, the interfering x_j, and noise. A decoder that performs poorly on an isolated interfering codeword can nevertheless extract information about user j from y_i, where x_j is present as interference mixed with the intended signal. To support the privacy claim, the authors should evaluate cross-decoding from the actual equalized received signal y_i (or a noise-free version of it), or quantify leakage via mutual information between y_i and s_j/u_j.","section":"Section V-B/V-C, Tables V and VII"}],"minor_comments":[{"comment":"'Sematic broadcast network' should be 'Semantic broadcast network'.","section":"Index Terms"},{"comment":"The notation gi ⊙ xi is unclear if gi is a scalar channel gain; please define whether gi is a scalar or vector and specify how the equalizer in Eq. (5) depends on gi, including the treatment of noise when dividing by gi.","section":"Eq. (4)"},{"comment":"Eq. (17) is missing a closing parenthesis and the summation over c is written inconsistently with the indexing of x_{c,j}; the sentence describing f_epsilon as a combination of Bernoulli and Cauchy distributions is not precise enough to reproduce the computation.","section":"Eq. (17)"},{"comment":"Table II lists a parameter ζ that does not appear in Eq. (30), and only one set of ABG parameters is given although Eq. (30) is written per user; please clarify whether the parameters are shared by all users and what ζ represents.","section":"Table II"},{"comment":"The Fig. 10 caption says training SNR = 5 dB for both panels, but the text says Fig. 10(a) uses training SNR = 0 dB; please reconcile this inconsistency.","section":"Fig. 10"},{"comment":"In the paragraph after Table V, the sentence 'the classification accuracy of User 1 decoding User 2's semantic information is only 9.71%' should read 'User 2 decoding User 1's semantic information'; similarly, the sentence after Table VII that refers to 'Table V' should refer to 'Table VII'.","section":"Section V-B/V-C, text around Tables V and VII"},{"comment":"Fig. 6(b) contains a stray '4' on the vertical axis and the axis label is incomplete; please correct the figure.","section":"Fig. 6(b)"}],"recommendation":"major_revision","confidential_remarks":"The paper's central idea is interesting and the simulation setup is extensive, but the revision will require more than cosmetic changes: the orthogonality mechanism needs either an explicit constraint or much stronger empirical validation at the actual operating points, the privacy test must be run on the received superposition, and the ABG-based power allocation must be reframed as an empirical model with out-of-sample validation rather than an analytical relationship. The overlap with prior multi-user semantic multiple access works such as DeepMA [35] and MDMA [33] should also be clarified in the introduction, since those works already address multi-user semantic transmission."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a real extension of multi-user semantic communication and the measured inner products and angles are a nice piece of evidence. But the load-bearing claim that orthogonality emerges reliably from training is supported only by a few runs at N=2 and N=3, with no explicit constraint or analysis, and the privacy argument has a logical gap. The paper is not ready for acceptance, but the idea deserves a serious referee.\n\nWhat is new: the authors combine discrete BPSK semantic features with RIB-based training for inference and a Swin Transformer for reconstruction, then fit an ABG curve for power allocation. That combination is not in DeepMA or MDMA, and the direct orthogonality measurements (inner products around 1e-3, angles near 90 degrees) give concrete evidence that the learned features are close to orthogonal. The CDF plots for the power allocation are also a useful practical check.\n\nThe soft spots are where the reader said. Eq. (8) is the core assumption, but nothing in the RIB loss (Eq. (16)) or the MSE loss (Eq. (29)) enforces pairwise orthogonality. The text argues that training drives the features toward orthogonality, and the tables show that for qbit = 128/4096, but the accuracy and PSNR curves are at qbit = 32/64. The operating points do not match. There are no error bars, and N is only 2 or 3. Second, the privacy claim is weaker than stated: feeding a clean x2 to f_theta1 shows near-chance accuracy, but the actual received y1 is a superposition. That experiment does not rule out information leakage through y1. Third, the ABG relationship is an empirical fit, not an analytical expression; the 'optimal' power allocation is optimal relative to that fitted model. That is not fatal, but the framing overstates it. Fourth, the RIB derivation has undefined quantities (q_theta_i(s_i|y_i) appears without definition) and a summation index typo; these are fixable.\n\nThe most relevant baselines, DeepMA and MDMA, are cited but not compared in the simulations. The comparison is only to deep JSCC and an upper bound.\n\nWho this is for: readers in semantic communications or multiple access who want to see whether learned feature-domain separation can work. It is a useful existence proof at small scale, but not a conclusive one.\n\nI would not desk-reject this. A serious referee could push for an orthogonality regularizer or an analysis of why it emerges, tests at the operating points of the headline curves, error bars, and a cleaner separation of the privacy and power-allocation claims. With those changes it could become a solid paper. As it stands, it needs heavy revision rather than acceptance.","headline":"SFDMA is a plausible extension of multi-user semantic communication with genuine empirical evidence of near-orthogonality, but the central claims rest on an unenforced and under-analyzed assumption.","tokens_in":838,"tokens_out":2101,"would_cite":false,"duration_ms":42670,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Orthogonal semantic codes let many users share one channel while keeping each user's data private.","keywords":["semantic communication","multi-user broadcast channel","multiple access","orthogonal feature representations","information bottleneck","power allocation","image reconstruction","semantic privacy"],"falsifier":"Take a trained two-user SFDMA encoder and feed it images from a held-out class or a channel fading profile not used in training, then compute the normalized inner product of the two encoded feature vectors. If it rises above roughly 0.1, or if cross-decoding accuracy moves well above chance (for MNIST, above about 15% on 10 classes), the claimed emergent orthogonality has not generalized and the interference and privacy guarantees fail for those conditions.","tokens_in":17534,"feed_emoji":"📡","tokens_out":5982,"duration_ms":54426,"temperature":0.7,"pith_summary":"This paper tries to show that a base station can broadcast different semantic messages to several users at once, on the same time and frequency, without the usual multi-user interference. The trick is to encode each user's information as a discrete, approximately orthogonal feature vector, so the signals coexist in one channel and each receiver can pull out only its own message. If this works, it gives a digital, bandwidth-efficient multiple-access scheme for semantic communication, with a degree of privacy built in: cross-decoding another user's signal fails. The paper also derives an empirical performance-versus-SINR formula and uses it to allocate power so each user meets a quality target under fading.","feed_headline":"Orthogonal semantic codes let many users share one channel","feed_subtitle":"Joint training alone makes user features nearly orthogonal, cutting interference and hiding data from other users.","key_machinery":"The load-bearing mechanism is the learned approximate orthogonality of encoded semantic features, stated as $x_i^H x_j/(\\|x_i\\|\\|x_j\\|) \\to 0$ for $i\\neq j$ (Eq. (8)). In the proposed SFDMA networks, a semantic encoder maps each source to a continuous feature, a sign binarizer quantizes it to $\\pm 1$ using the straight-through estimator, and BPSK modulation normalizes power; the whole encoder is trained jointly across users so that interference between users is minimized. For inference tasks, the robust information bottleneck (RIB) objective, approximated by a variational upper bound on mutual information terms, trades off inference accuracy, compression, and interference. For image reconstruction, a Swin Transformer encoder-decoder with MSE loss plays the same role, and the ABG function $\\phi_i = \\alpha_i - \\gamma_i/(1+(\\beta_i \\mathrm{SINR}_i)^{\\tau_i})$ links task performance to SINR for power allocation.","core_discovery":"The central discovery is that multi-user interference in a semantic broadcast channel can be handled in the feature domain rather than the signal-processing domain. With user-specific encoders trained jointly, the quantized BPSK-modulated semantic features of different users become approximately orthogonal, with normalized inner products around $10^{-3}$ and angles near $90^\\circ$, even when the inputs are identical. This makes it possible to superpose all users' signals in the same time-frequency resource; each decoder recovers its own semantic information while decoding another user's signal yields near-chance accuracy or, for images, very low PSNR and MS-SSIM. The orthogonality is not imposed by a loss term: it is described as emerging spontaneously because minimizing reconstruction error rewards well-separated signals.","pith_inferences":["If the emergent orthogonality is the only thing separating users, then out-of-distribution inputs or channel conditions not seen during training could break the separation; a targeted stress test on unseen classes or fading statistics would show how much margin exists.","The near-chance cross-decoding results suggest statistical separation, not cryptographic secrecy; the scheme should be described as providing confidentiality against casual receivers, not as a secure physical-layer privacy mechanism.","The ABG fit is empirical and dataset-specific, so the power-allocation guarantee likely carries only as far as the fitted parameters; refitting on another dataset or channel model would be needed before deployment.","The same feature-domain division idea could be applied to other multi-user settings, such as uplink semantic access or over-the-air federated learning, though the paper does not explore those cases."],"forward_implications":["Users' signals can share the same time-frequency resource without successive interference cancellation, so receiver complexity no longer limits the number of superposed users to two.","Semantic privacy becomes a side effect of the coding: a receiver that tries to decode another user's feature gets near-chance results, so user data is not exposed to other users in the broadcast.","The ABG performance-SINR curve turns semantic quality-of-service constraints into a linear power-allocation problem, enabling adaptive power control in fading channels.","Because the transmitted features are binary and BPSK-modulated, the scheme is compatible with digital modulation chains rather than requiring analog transmission of continuous features.","With more users or new data domains, the encoders would need retraining to re-establish the near-orthogonality that the current experiments show for two and three users."],"supporting_citations":[{"why":"Supplies the robust information bottleneck formulation for digital task-oriented communication that the inference SFDMA network builds on.","marker":"[19]"},{"why":"Provides the Swin Transformer architecture used as the semantic encoder and decoder for image reconstruction tasks.","marker":"[36]"},{"why":"Provides the straight-through estimator that makes the sign quantizer trainable, enabling discrete binary features.","marker":"[37]"},{"why":"Provides the deep variational information bottleneck technique used to derive the tractable upper bound on the RIB objective.","marker":"[38]"},{"why":"Serves as the deep JSCC baseline that SFDMA is compared against in the multi-user experiments.","marker":"[40]"},{"why":"Supplies the MNIST dataset used to evaluate the inference-task SFDMA network.","marker":"[41]"},{"why":"Supplies the CelebA dataset used to evaluate the image-reconstruction SFDMA network.","marker":"[42]"},{"why":"Provides the t-SNE visualization used to show that encoded features of different users separate into distinct clusters.","marker":"[43]"}],"fun_headline_variants":["Orthogonal semantic features let users share one channel","Semantic features go orthogonal for multi-user broadcast","Feature-domain orthogonality ends multi-user interference","One time-frequency slot for many users via semantic codes","Semantic broadcast channel now serves multiple users at once"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that joint training alone will push different users' encoded signals to be nearly orthogonal in practice; if that emergent separation fails for new data, channels, or more users, the interference mitigation and privacy claims collapse.","fun_headline_variants_meta":{"raw":{"variants":["Orthogonal semantic features let users share one channel","Semantic features go orthogonal for multi-user broadcast","Feature-domain orthogonality ends multi-user interference","One time-frequency slot for many users via semantic codes","Semantic broadcast channel now serves multiple users at once"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000731,"raw_usage":{"total_tokens":3248,"prompt_tokens":901,"completion_tokens":2347,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":517,"completion_tokens_details":{"reasoning_tokens":2274}},"tokens_in":517,"tokens_out":2347,"duration_ms":16191,"temperature":1.0,"reasoning_tokens":2274,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-09T00:07:13.074175+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a trained two-user SFDMA encoder and feed it images from a held-out class or a channel fading profile not used in training, then compute the normalized inner product of the two encoded feature vectors. If it rises above roughly 0.1, or if cross-decoding accuracy moves well above chance (for MNIST, above about 15% on 10 classes), the claimed emergent orthogonality has not generalized and the interference and privacy guarantees fail for those conditions.","supporting_citations":[{"cited_title":"Swin transformer: Hierarchical vision transformer using shifted windows,","cited_arxiv_id":null,"evidence_quote":"Provides the Swin Transformer architecture used as the semantic encoder and decoder for image reconstruction tasks."},{"cited_title":"Visualizing data using t-SNE","cited_arxiv_id":null,"evidence_quote":"Provides the t-SNE visualization used to show that encoded features of different users separate into distinct clusters."}],"review_version":1}