{"id":"81b40e19-a3be-478f-8a4d-d44376637920","arxiv_id":"2607.07455","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":14,"one_line_summary":"AWGN-trained DeepSC generalizes to THz fading channels with a pilot-free neural detector that outperforms MMSE equalization with perfect CSI.","lead":"This paper shows that AI-based semantic communication systems, trained on simple noise, can work over complex terahertz channels without needing channel estimation. A smart business person might read this because it could simplify future 6G wireless hardware by removing costly channel measurement components.","discovery_kind":"unclear","skeptic_critique":{"model":"glm-5.2","headline":"The ND's 'CSI-free' advantage is contingent on having learned the TeraMIMO channel distribution during Stage 2 training; cross-environment generalization is the untested load-bearing assumption.","rationale":"The reader correctly identified the most load-bearing concern: the ND is trained and tested on the same TeraMIMO channel model, making the 'CSI-free' framing misleading. The ND has implicitly learned the channel distribution during Stage 2 training, and true cross-environment generalization is untested. This is the right concern and it is precisely stated. The reader's verdict of CONDITIONAL with MODERATE confidence is appropriate. The paper makes a legitimate technical contribution — the two-stage training strategy, the post-equalization noise argument for AWGN training, and the ND architecture are all reasonable. The simulation methodology includes proper statistical reporting (30 test passes, confidence intervals, ablations). The ND outperforming linear MMSE is plausible because the ND can exploit the non-Gaussian structure of DeepSC symbols and sentence-level dependencies through self-attention, which linear MMSE cannot. However, the 'CSI-free' claim is overstated: it should be clarified as 'no explicit CSI at inference' rather than 'channel-agnostic.' The 4096×4096 antenna configuration is extreme but is a modeling choice consistent with ultra-massive MIMO THz literature. The >50 percentage-point BLEU-1 gap over the traditional baseline is partly inflated by the use of ASCII source encoding, but this is a standard comparison in the semantic communication literature and the throughput is matched. No code or data is shipped, which limits reproducibility. The verdict should remain CONDITIONAL: the contribution is real but the generalization claims need to be either tested on different channel models or scoped more carefully.","tokens_in":9106,"tokens_out":5238,"duration_ms":274734,"concrete_test":"Train the ND on TeraMIMO indoor parameters as described, then test it on a held-out channel configuration with materially different parameters — e.g., outdoor THz scenario with different cluster/ray arrival rates, a different antenna array size (e.g., 256×256), or non-ideal beamforming with steering errors. Compare ND BLEU-1 against perfect-CSI MMSE on the new configuration. If the ND's BLEU-1 drops by more than 10 percentage points relative to its performance on the training-matched configuration while MMSE with perfect CSI remains stable, the 'CSI-free' advantage is distribution-specific and the generalization claim does not hold.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that the pilot-free ND outperforms perfect-CSI MMSE (Fig. 3a–b) rests on a specific experimental loop: the ND is trained on THz channel realizations from TeraMIMO with fixed indoor Saleh-Valenzuela parameters (T=298.15 K, p=1 atm, cluster/ray arrival rates 0.13/0.37 ns⁻¹, 4096×4096 antennas, ideal geometry-aligned beamforming) and tested on new independent realizations from the same simulator with the same parameters. The ND has no explicit CSI at inference, but it has learned the channel statistics during Stage 2 training. The 'CSI-free' label means 'no explicit CSI at inference,' not 'channel-agnostic.' The post-equalization noise argument (§3.2.1) motivates AWGN training of the DeepSC encoder-decoder, but it does not justify the ND's channel-agnostic operation — the ND is explicitly trained on THz-corrupted observations. If the real THz channel differs from TeraMIMO's model (outdoor scenarios, different antenna configurations, non-ideal beamforming, different absorption profiles), the ND may fail silently because it has no error signal to detect distribution shift. The ±50 MHz frequency offset test (Fig. 3b) is a very small perturbation within the same model family and does not address this. The claim 'outperforms MMSE with perfect CSI' is thus specific to the TeraMIMO channel distribution and may not hold for other THz environments.","agreement_with_reader":"agree"},"referee_report":{"model":"glm-5.2","summary":"This manuscript investigates deep learning-based semantic communication (DeepSC) over terahertz (THz) channels. The authors make three main claims: (1) DeepSC models trained solely under AWGN generalize well to tested THz block- and fast-fading channels when receiver-side compensation is applied, motivated by a post-equalization noise argument; (2) a proposed lightweight, pilot-free neural detector (ND) outperforms MMSE equalization with both perfect and imperfect CSI; and (3) DeepSC is more robust to CSI errors than a throughput-matched traditional coded baseline. The system is evaluated using the TeraMIMO channel simulator at 0.3 THz, and the ND is trained over THz channel realizations from this simulator while the DeepSC encoder-decoder is frozen after AWGN pre-training.","tokens_in":9397,"tokens_out":1173,"duration_ms":345525,"significance":"The paper presents a falsifiable and well-structured experimental methodology. The throughput-matched baseline comparison (4.89 vs 4.8 bits/symbol) is a notable strength, ensuring a fair comparison between semantic and traditional coded systems. The post-equalization noise argument in §3.2.1 provides a parameter-free motivation for AWGN training. The inclusion of 30 independent test passes with reported confidence intervals (max std 1.3e-2 BLEU-1) and receiver ablations (Conv-only, MLP-128) adds rigor to the empirical claims. The proposed pilot-free ND for continuous semantic channel symbols addresses a practical gap in THz communications where CSI acquisition is costly.","major_comments":[{"comment":"§3.2.1, Eq. (2) context: The post-equalization noise argument assumes that under perfect-CSI MMSE, the fading channel y=hx+n becomes x_hat = alpha*x + n_tilde where n_tilde remains white Gaussian. This motivates AWGN training of the DeepSC encoder-decoder. However, the argument is used to motivate the entire CSI-free system, including the ND. The ND (Stage 2, Algorithm 1) is explicitly trained on THz-corrupted observations from the TeraMIMO simulator with specific indoor Saleh-Valenzuela parameters (T=298.15 K, p=1 atm, 4096x4096 antennas). The 'CSI-free' label means 'no explicit CSI at inference,' not 'channel-agnostic.' The manuscript should clarify this distinction explicitly, as the ND's advantage is structurally dependent on having learned the TeraMIMO channel distribution during Stage 2 training. Without this clarification, the 'CSI-free' claim risks being overstated.","section":null},{"comment":"§4.3, Fig. 3(b): The claim that the ND 'outperforms MMSE with perfect CSI' is specific to the TeraMIMO channel distribution used in training and testing. The ND is trained and evaluated on independent realizations from the same simulator with the same parameters. Cross-environment generalization (e.g., outdoor scenarios, different antenna configurations, non-ideal beamforming) is not tested. The ±50 MHz frequency offset test is a very small perturbation within the same model family and does not address distribution shift across different THz environments. The authors should explicitly state that the perfect-CSI MMSE outperformance is specific to the tested TeraMIMO indoor scenario and acknowledge that the ND may fail silently under distribution shift since it has no error signal to detect it.","section":null}],"minor_comments":[{"comment":"§3.1: The ND architecture description mentions d_model=64, 8 attention heads, and FFN dimension 256, but the number of trainable parameters (100,802) seems surprisingly small for two Transformer encoder layers with these dimensions. A brief breakdown of parameter count would help readers verify this.","section":null},{"comment":"Fig. 2(a,b): The legend distinguishes 'AWGN-trained' (dashed) from 'same-channel trained' (solid) models, but the notation 'AWGN (AWGN)', 'Rayleigh (AWGN)', etc., is initially confusing. Consider relabeling for clarity, e.g., 'Trained: AWGN, Tested: Rayleigh'.","section":null},{"comment":"§4.2: The latency analysis reports processing times on a V100 GPU, but the batch size and sequence length assumptions for the timing comparison could be stated more explicitly to ensure the comparison between MMSE and ND is fair.","section":null},{"comment":"§4.3, Eq. (3): The high-SNR approximation is used to analyze CSI mismatch, but the tested SNR range includes 0-12 dB where the approximation may not hold well. A brief comment on the validity of this approximation at low SNR would strengthen the analysis.","section":null},{"comment":"References [14-16] on diffusion-based denoisers are cited but not deeply compared to the proposed ND. A sentence clarifying why diffusion-based approaches are unsuitable for pilot-free text semantic equalization (beyond targeting image transmission) would improve the related work.","section":null}],"recommendation":"minor_revision","confidential_remarks":"The reader's concern about the 'CSI-free' label being structurally dependent on the ND having learned the TeraMIMO channel distribution is valid and load-bearing. However, this is a framing/clarification issue rather than a fundamental flaw—the experimental setup is sound and the results are honestly reported within the tested scope. The paper's claims are defensible if the authors explicitly acknowledge that 'CSI-free' means 'no explicit CSI at inference' and that the ND is trained on THz channel data. I recommend minor revision to address this framing and the scope of the generalization claim. The paper is a reasonable fit for the journal's scope on THz and semantic communications."},"author_rebuttal":{"model":"glm-5.2","summary":"We thank the referee for the careful reading and constructive feedback. Both major comments concern the scope and generalizability of the 'CSI-free' label and the ND's outperformance of perfect-CSI MMSE. We agree that the manuscript should explicitly clarify that the ND is trained on the TeraMIMO channel distribution and that its advantages are specific to the tested indoor scenario. We will revise the manuscript accordingly.","responses":[{"response":"The referee is correct. The post-equalization noise argument in §3.2.1 motivates AWGN pre-training of the DeepSC encoder-decoder: under perfect-CSI MMSE, the effective channel seen by the decoder is approximately AWGN-like, so an AWGN-trained decoder should generalize. This argument does not, however, motivate the ND itself. The ND is trained on THz-corrupted observations drawn from the TeraMIMO simulator (Stage 2, Algorithm 1) and learns to invert the specific channel distribution it was trained on. The 'CSI-free' label accurately describes the inference-time property—no pilots, no explicit channel estimates, no matrix inversions—but it does not mean the ND is channel-agnostic. Its advantage over MMSE is structurally dependent on having learned the TeraMIMO indoor channel statistics during training. We will revise the manuscript to state this distinction explicitly in §3.2.1 and §3.2.2, and will add a clarifying note to the abstract and contributions list so that 'CSI-free' is understood as 'no explicit CSI at inference' rather than 'channel-distribution-independent.'","revision_made":"yes","referee_comment":"§3.2.1, Eq. (2) context: The post-equalization noise argument assumes that under perfect-CSI MMSE, the fading channel y=hx+n becomes x_hat = alpha*x + n_tilde where n_tilde remains white Gaussian. This motivates AWGN training of the DeepSC encoder-decoder. However, the argument is used to motivate the entire CSI-free system, including the ND. The ND (Stage 2, Algorithm 1) is explicitly trained on THz-corrupted observations from the TeraMIMO simulator with specific indoor Saleh-Valenzuela parameters. The 'CSI-free' label means 'no explicit CSI at inference,' not 'channel-agnostic.' The manuscript should clarify this distinction explicitly, as the ND's advantage is structurally dependent on having learned the TeraMIMO channel distribution during Stage 2 training. Without this clarification, the 'CSI-free' claim risks being overstated."},{"response":"The referee's assessment is accurate. The ND's outperformance of perfect-CSI MMSE in Fig. 3(b) is observed on independent realizations from the same TeraMIMO indoor scenario with the same Saleh-Valenzuela parameters used in training. We did not test cross-environment generalization (outdoor, different antenna configurations, non-ideal beamforming), and the ±50 MHz frequency offset is indeed a small perturbation within the same model family rather than a cross-distribution test. We agree that the manuscript should explicitly scope the claim to the tested TeraMIMO indoor scenario. We also acknowledge the important point about silent failure under distribution shift: since the ND has no explicit error or confidence signal, it may degrade without detection if deployed in an environment whose statistics differ from training. We will add a statement to §4.3 and the Conclusion explicitly noting that the perfect-CSI MMSE outperformance is specific to the tested TeraMIMO indoor scenario, that cross-environment generalization is untested, and that silent failure under distribution shift is a known limitation of the current design that warrants future work (e.g., distribution-shift detection or online adaptation).","revision_made":"yes","referee_comment":"§4.3, Fig. 3(b): The claim that the ND 'outperforms MMSE with perfect CSI' is specific to the TeraMIMO channel distribution used in training and testing. The ND is trained and evaluated on independent realizations from the same simulator with the same parameters. Cross-environment generalization (e.g., outdoor scenarios, different antenna configurations, non-ideal beamforming) is not tested. The ±50 MHz frequency offset test is a very small perturbation within the same model family and does not address distribution shift across different THz environments. The authors should explicitly state that the perfect-CSI MMSE outperformance is specific to the tested TeraMIMO indoor scenario and acknowledge that the ND may fail silently under distribution shift since it has no error signal to detect it."}],"tokens_in":8881,"tokens_out":967,"duration_ms":156254,"standing_objections":[]},"desk_editor":{"model":"glm-5.2","letter":"The main thing to know: this paper builds a pilot-free neural detector (ND) for continuous DeepSC channel symbols over THz channels, which is a genuine gap. Prior neural detectors (MMNet, OAMP-Net, VBINet) target discrete constellations, not continuous semantic embeddings. The two-stage training — AWGN pretraining of DeepSC, then ND training on TeraMIMO THz realizations — is clean, and the post-equalization noise argument (Section 3.2.1) is a simple, parameter-free motivation for why AWGN training should generalize. The experimental methodology is solid: throughput-matched baseline (4.8 vs 4.89 bits/symbol), 30 independent test passes, tight confidence intervals (max std 1.3e-2 BLEU-1), and ablations (Conv-only, MLP-128). The 50+ percentage-point BLEU-1 gain over LDPC-coded 64-QAM at 0-12 dB is a real result. Fig. 2(a,b) showing AWGN-trained models matching or beating channel-specifically trained models is the most interesting finding here — it suggests semantic encoders are surprisingly insensitive to fading profile when receiver compensation is applied. The noise-distribution ablation (Fig. 2c) confirming sensitivity to noise type but not fading rate is a nice supporting experiment. The soft spot is real but bounded. The stress-test concern lands: the ND is trained on TeraMIMO THz channels and tested on new realizations from the same simulator with the same parameters. So 'CSI-free' means no explicit CSI at inference, not channel-agnostic. The ND has implicitly learned the TeraMIMO channel distribution. The 50 MHz frequency offset test is a small perturbation within the same model family and doesn't address cross-environment generalization. The 4096x4096 antenna configuration is extreme, though the authors cite prior work justifying ultra-massive MIMO at THz. That said, the paper doesn't overclaim on this point in the body — the framing issue is mostly in the abstract's 'CSI-free' language. A reviewer should ask for clarification that this means inference-time CSI-free, not channel-agnostic, and ideally a test on a second simulator or measured data. No code or data is shipped, which limits reproducibility. The free-parameter list is long but most are standard Transformer hyperparameters; the architecture is compact at 100K parameters. This paper is for researchers working at the intersection of semantic communication and THz/beyond-6G systems. It deserves a serious referee who can push on the generalization claim and the CSI-free framing. The core result — that a lightweight ND can replace MMSE equalization for semantic symbols over simulated THz channels — holds up.","headline":"Pilot-free neural detector for semantic communication over THz channels — a real gap, but the CSI-free claim is narrower than it sounds.","tokens_in":10197,"tokens_out":646,"would_cite":true,"duration_ms":118687,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"glm-5.2","headline":"Pilot-free neural detector beats CSI-based equalization for THz semantic links","keywords":["semantic communication","terahertz communications","neural detector","channel state information","deep learning","equalization","DeepSC","THz MIMO"],"falsifier":"If the post-equalization residual noise in real THz channels is significantly non-Gaussian or colored due to effects not captured by the TeraMIMO simulator, the AWGN-trained DeepSC decoder would degrade, and the ND trained on simulated data would fail to compensate, causing the pilot-free system to underperform CSI-based equalization.","tokens_in":9299,"feed_emoji":"📡","tokens_out":1141,"duration_ms":202637,"temperature":0.7,"pith_summary":"This paper argues that a deep learning-based semantic communication system (DeepSC), trained only on simple additive white Gaussian noise (AWGN), can operate over realistic terahertz (THz) fading channels without channel-specific retraining. The key insight is a post-equalization argument: under perfect minimum mean square error (MMSE) equalization, a fading channel reduces to an equivalent AWGN channel with rescaled noise, so a decoder trained on AWGN should generalize. To make the system fully free of channel state information (CSI), the authors propose a lightweight Transformer-based neural detector (ND) that replaces conventional equalization. The ND is trained on simulated THz channel data to recover clean transmitted channel-symbol sequences directly from corrupted observations, without pilots. At 0.3 THz, the pilot-free ND matches perfect-CSI MMSE and outperforms pilot-based and noisy-CSI MMSE, while the overall DeepSC system achieves more than 50 percentage-point higher BLEU-1 than a throughput-matched LDPC-coded 64-QAM baseline over 0-12 dB SNR. The ND also remains robust to carrier-frequency offsets of up to 50 MHz without retraining.","feed_headline":"Pilot-free neural detector beats CSI-based equalization for THz links","feed_subtitle":"AWGN-trained semantic communication matches full-CSI MMSE at 0.3 THz without pilots, cutting retraining and pilot overhead for 6G.","key_machinery":"The mechanism is a two-stage training procedure: Stage 1 trains the DeepSC encoder-decoder end-to-end over an AWGN channel and freezes it; Stage 2 trains a Transformer-based neural detector (100,802 parameters, two encoder layers, eight attention heads) to map THz-corrupted received signals back to the clean channel-symbol distribution expected by the frozen decoder. The ND uses self-attention to exploit sentence-level dependencies among continuous channel symbols, distinguishing it from constellation-level detectors that assume independent discrete symbols. The THz channel is generated via the TeraMIMO simulator at 0.3 THz with indoor Saleh-Valenzuela parameters and 4096x4096 antenna arrays","core_discovery":"The central discovery is that AWGN-only training of a semantic communication system suffices for THz fading channels when paired with a receiver that compensates for the channel, and that a compact pilot-free neural detector can serve as that compensator while outperforming conventional MMSE equalization that has access to CSI. The paper also shows that semantic performance is more sensitive to the post-equalization noise distribution than to the fading profile itself, which is why AWGN training transfers across block fading, fast fading, and molecular absorption scenarios.","pith_inferences":["If the ND implicitly learns the THz channel statistics during Stage 2 training, then its CSI-free claim is contingent on the real channel matching the TeraMIMO simulator's model; deployment in environments not represented in training data could cause silent failures without an explicit error signal.","The two-stage approach separates semantic coding from channel compensation, which suggests the ND could be swapped or retrained independently for different THz scenarios while reusing the same AWGN-trained DeepSC backbone, though this modularity is not tested in the paper.","The post-equalization noise-whiteness argument may break down under spatially correlated MIMO channels or near-field propagation effects specific to ultra-massive antenna arrays, where residual interference after equalization may not be white or Gaussian."],"forward_implications":["If the post-equalization noise argument holds broadly, semantic communication systems could be deployed across diverse fading environments without costly channel-specific retraining, reducing the engineering burden for new frequency bands.","Pilot-free neural detection could eliminate the pilot overhead and CSI estimation pipelines that are especially costly at THz frequencies, simplifying transceiver design for future 6G systems.","The robustness to frequency offsets suggests that the ND learns channel-symbol-level structure rather than memorizing exact channel parameters, which could make it adaptable to hardware impairments like oscillator drift.","The finding that semantic performance depends more on noise distribution than fading profile implies that future semantic coding efforts should prioritize robustness to non-Gaussian post-equalization residual noise."],"fun_headline_variants":["AWGN-trained semantic models transfer to THz fading without retraining","Pilot-free neural detector outperforms full-CSI MMSE at 0.3 THz","Semantic communication survives THz block and fast fading with receiver compensation","DeepSC beats coded baseline by 50+ BLEU points over THz channels","Compact pilot-free detector stays robust to 50 MHz frequency offsets at THz"],"cache_read_input_tokens":0,"weakest_assumption_plain":"The argument for AWGN training rests on the claim that after perfect-CSI MMSE equalization, the residual noise remains white and Gaussian, so the decoder sees an AWGN-like channel. The ND trained on TeraMIMO-simulated indoor THz channels is then assumed to generalize to real THz environments, but this is only tested for one simulator configuration, one indoor scenario, and one antenna array size.","fun_headline_variants_meta":{"raw":{"variants":["AWGN-trained semantic models transfer to THz fading without retraining","Pilot-free neural detector outperforms full-CSI MMSE at 0.3 THz","Semantic communication survives THz block and fast fading with receiver compensation","DeepSC beats coded baseline by 50+ BLEU points over THz channels","Compact pilot-free detector stays robust to 50 MHz frequency offsets at THz"]},"model":"glm-5.2","effort":"low","cost_usd":0.0,"raw_usage":{"total_tokens":612,"prompt_tokens":511,"completion_tokens":101,"prompt_tokens_details":null},"tokens_in":511,"tokens_out":101,"duration_ms":91230,"temperature":1.0,"reasoning_tokens":null,"cache_read_input_tokens":0,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-09T10:19:27.504191+00:00","model_set":{"reader":"glm-5.2"},"falsifier":"If the post-equalization residual noise in real THz channels is significantly non-Gaussian or colored due to effects not captured by the TeraMIMO simulator, the AWGN-trained DeepSC decoder would degrade, and the ND trained on simulated data would fail to compensate, causing the pilot-free system to underperform CSI-based equalization.","supporting_citations":[],"review_version":1}