{"id":"1c449d40-c3e6-4406-b37b-813c854a276c","arxiv_id":"2501.17888","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"RadioLLM shows that a GPT-2 backbone with token reprogramming, hybrid prompts, and CNN fusion outperforms task-specific networks on most radio classification and denoising benchmarks.","lead":"RadioLLM rewrites raw radio I/Q samples into token-like embeddings and feeds them to GPT-2 with compact prompts drawn from word embeddings, plus a CNN branch for high-frequency detail. It reports top accuracy on six of seven radio classification benchmarks and better denoising scores, but the comparison is weakened because RadioLLM is pre-trained on the same datasets most baselines never see.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Asymmetric pretraining on RML18A and Wi-Fi may explain reported classification wins; equalized baselines needed.","rationale":"The reader's weakest assumption is precisely the asymmetry: RadioLLM pretrains on the same datasets on which it is evaluated, while most baselines are restricted to 100 labeled examples. Section IV-G partially addresses this for two SSL baselines on three RML datasets, and the results there favor RadioLLM, so the concern is not automatically fatal. Yet the same equalization is missing for RML18A and Wi-Fi, the two datasets where pretraining exposure is strongest and where RadioLLM claims wins. RML22 offers a fair held-out test and RadioLLM wins, which supports the method, but the central claim is a majority-of-benchmarks statement that must hold across the full table. The ablation inconsistency on SSIM is a further internal contradiction that weakens confidence in the component attribution, though it is secondary to the comparison fairness issue. The proposed test directly isolates whether pretraining exposure explains the RML18A and Wi-Fi margins; if RadioLLM still wins after baseline pretraining, the central claim is materially strengthened. I therefore see no reason to change the reader's CONDITIONAL verdict, only to emphasize that the missing equalized comparison on RML18A and Wi-Fi is the key unresolved threat.","tokens_in":20193,"tokens_out":5320,"duration_ms":48908,"concrete_test":"Re-run the 100-shot classification comparison on RML18A and Wi-Fi with TcssAMR and SemiAMC pretrained on the same unlabeled training splits (using their own self-supervised objectives and the same 8:1:1 partition), then fine-tune with the same 100 labeled examples and report OA/Kappa over at least 5 seeds. If either baseline matches or surpasses RadioLLM on these datasets, the reported superiority on those benchmarks is not established.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim rests on Table I, where RadioLLM reports the highest OA/Kappa on six of seven datasets. However, Section IV-A pretrains RadioLLM on RML16A/B/C, RML18A, ADS-B, and Wi-Fi, and Section IV-F then evaluates 100-shot classification on those same datasets. Section IV-G equalizes only TcssAMR and SemiAMC, and only on RML16A/B/C; for those three datasets the equalized baselines still trail RadioLLM, which is evidence in the paper's favor. But no equalization is reported for RML18A or Wi-Fi, where RadioLLM claims gains of 52.03 vs 47.66 and 35.41 vs 34.59. Pretraining on the full RML18A training set and 5% of the Wi-Fi training set gives RadioLLM substantial unsupervised feature learning that baselines do not receive, so those two wins may reflect data exposure rather than HPTR and FAF. RML22, which was not pretrained, provides a fairer test and RadioLLM wins there, but the majority-of-benchmarks claim still includes RML18A and Wi-Fi wins that are not yet fairly established. The ablation table also undercuts the component story: full HTRP+FAF SSIM is 0.838, below FAF-only 0.857, contradicting the text's claim of consistent improvements across all metrics. Without error bars or an equalized comparison on the un-equalized datasets, the central comparative claim is not fully supported.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes RadioLLM, a framework that reprograms GPT-2 to process raw I/Q radio signals via Hybrid Prompt and Token Reprogramming (HPTR) and a Frequency-Attuned Fusion (FAF) module. The model is pretrained on six public radio datasets and then evaluated on 100-shot classification over seven datasets and on denoising over three datasets. The paper claims that RadioLLM achieves superior performance over task-specific RSC and denoising baselines in the majority of testing scenarios, with Table I reporting the highest OA/Kappa on six of seven classification benchmarks and Table II reporting the highest SSIM on all three denoising benchmarks.","tokens_in":20409,"tokens_out":4591,"duration_ms":39857,"significance":"If the result holds, RadioLLM would be a meaningful demonstration that a relatively small LLM backbone can be reprogrammed into a unified radio-signal processing front-end, which is a plausible and timely direction for cognitive radio. The claim is made more credible by the deliberate holdout of RML22 from pretraining, by the equalization experiment in Section IV-G showing that two pretrained baselines still trail on three RML datasets, and by the inclusion of an ablation table. However, the fairness of the benchmark comparison is the central load-bearing issue: most evaluation datasets are also pretraining datasets for RadioLLM but not for the baselines. The paper does not yet provide the evidence needed to attribute the RML18A and Wi-Fi wins to HPTR and FAF rather than to asymmetric data exposure.","major_comments":[{"comment":"Pretraining in Section IV-A uses RML16A, RML16B, RML16C, RML18A, ADS-B, and Wi-Fi, and classification results in Section IV-F are then reported on those same datasets, while the baselines receive only 100 labeled samples. The reported gains on RML18A (52.03 vs. 47.66) and Wi-Fi (35.41 vs. 34.59) may therefore reflect the pretraining exposure rather than the efficacy of HPTR and FAF. Section IV-G equalizes only TcssAMR and SemiAMC, and only on RML16A/B/C; the equalized baselines on those three datasets still trail RadioLLM, which is evidence in the paper's favor, but no equalized comparison is provided for RML18A or Wi-Fi. An equalized comparison on those datasets, or removal of those wins from the central claim, is required. The RML22 result, which is not pretrained, is a fair test and supports the method.","section":"IV-A, IV-F"},{"comment":"The ablation contradicts the text's claim of consistent improvements when HTRP and FAF are combined: the SSIM of the full HTRP+FAF configuration (0.838) is lower than that of FAF-only (0.857). The paragraph in Section IV-I states that the joint configuration 'delivers consistent improvements across all evaluation metrics,' which is directly contradicted by the same table. This needs an explanation, for example if SSIM is evaluated only on the denoising task while the primary target is classification, and the wording should be corrected to match the numeric results.","section":"IV-I, Table III"},{"comment":"The paper reports a single run for each configuration without variance, confidence intervals, or significance testing. Several reported margins are small (for example, 58.10 vs. 57.60 for the equalized TcssAMR on RML16A, and 35.41 vs. 34.59 on Wi-Fi), so the 'superior performance' claim is not yet quantitatively grounded. At least a small number of random seeds and a statement of variance are needed to establish that the observed differences are not noise.","section":"IV-F, IV-H, Tables I and II"},{"comment":"The denoising evaluation on RML16A, RML16B, and RML16C is subject to the same asymmetry as the classification comparison: RadioLLM was pretrained on these datasets, while the denoising baselines SGFilter and DNCNet were not. The qualitative and quantitative superiority claimed in Section IV-H would be more convincing if the pretrained-equivalent baselines from Section IV-G were also evaluated for denoising, or if denoising were reported on a dataset not used in pretraining.","section":"IV-H, Table II"}],"minor_comments":[{"comment":"The abbreviation 'HTRP' appears in several places (the contributions bullet in Section I and the Table III caption and surrounding text) but the paper consistently defines the module as HPTR; please unify the abbreviation.","section":"I, IV-I"},{"comment":"The Index Terms contain the typo 'Technolog'; it should read 'Technology'.","section":"Index Terms"},{"comment":"In Section III-E, 'fine-tune GPT-2 using the LoRA technique [24]' is correct in the reference list, but the later sentence 'with only a subset updated via LoRA [7]' cites the network optimization reference [7] instead of the LoRA reference; please fix the citation.","section":"III-E"},{"comment":"For RML18A, the text states 2,555,904 total samples, 24 modulation types, and 26 SNR levels with 4096 samples per class per SNR; these numbers imply 24*26*4096 = 2,555,904, which is consistent, but the phrase '26 SNR levels from -20 dB to 30 dB in 2 dB increments' actually lists 26 values, so the count is consistent; please confirm the SNR range wording is intended.","section":"IV-B"},{"comment":"Figure 4 has panels labeled (a)-(e), (f)-(j), and (k)-(o), but the text refers to 'Fig. 4 (o)' both for the RML18A confusion matrix and for the misclassification discussion; please renumber or refer to the specific panel of the confusion matrix more clearly.","section":"IV-F, Fig. 4"},{"comment":"The text in Section IV-G discusses the results shown in Fig. 5 but does not explicitly cite the figure number in the paragraph; please add a citation to Fig. 5.","section":"IV-G, Fig. 5"}],"recommendation":"major_revision","confidential_remarks":"The paper is within scope for a cognitive-radio or signal-processing journal and does not appear to involve any misconduct; the RML22 holdout and the Section IV-G equalization experiment show good-faith effort. The main obstacle to acceptance is experimental fairness: the central claim rests on wins over baselines that were not given the same pretraining exposure on several evaluation datasets. If the authors add equalized comparisons for RML18A and Wi-Fi (and ideally for the denoising benchmarks) and correct the ablation narrative, the contribution could become supportable. The absence of variance reporting is also a standard but important fix."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth a look if you care about reprogramming LLMs for non-text signals. The paper takes the Time-LLM/TEMPO token-and-prompt reprogramming recipe, applies it to raw I/Q radio signals with GPT-2 as the frozen backbone, and reports SOTA-ish numbers across seven classification datasets and three denoising benchmarks. That domain transfer is the genuinely new bit, and the paper is honest about its lineage—it cites Time-LLM, TEMPO, and S2IP-LLM. The engineering is thorough: seven datasets, ten baselines, ablations, parameter sensitivity, LLM choice, and a control experiment where two baselines are given multi-domain pretraining.\n\nThe soft spot is real but narrower than a fatal flaw. RadioLLM is pretrained on RML16A/B/C, RML18A, and Wi-Fi, the same datasets it is later evaluated on, while most baselines only get the 100-shot labeled split. Section IV-G equalizes only TcssAMR and SemiAMC on three RML datasets; on those, RadioLLM still wins, which is evidence in its favor. But there is no equalization on RML18A or Wi-Fi, where the claimed wins are 52.03 vs 47.66 and 35.41 vs 34.59. RML22, which was not pretrained, gives the fairest test, and RadioLLM wins there too—so the core idea is not hollow, but the majority-of-benchmarks claim is not yet fairly established.\n\nSecond soft spot: the ablation table contradicts its own narrative. Full model SSIM is 0.838, below FAF-only 0.857, yet the text says the combined model delivers consistent improvements across all metrics. That is an internal inconsistency the authors should fix. It does not kill the classification claim, but it shows the two modules trade off rather than compose cleanly.\n\nMinor: no error bars, no code, single seed. The gains over equalized baselines are under a percentage point on RML16A/B, so the quantitative claim is fragile even where the comparison is fair.\n\nBottom line: this is a serious attempt at an interesting idea, not a sloppy one. The math is standard, the self-citations are not load-bearing, and the central claim is plausible. It deserves a referee, but the referee should ask for equalized baselines on all pretrained datasets, multiple seeds, released code, and a corrected ablation discussion.","headline":"Plausible but not yet proven: the first real application of LLM token reprogramming to raw I/Q radio signals, undermined by asymmetric pretraining and an ablation inconsistency.","tokens_in":21053,"tokens_out":3840,"would_cite":false,"duration_ms":27694,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A single pretrained language model, adapted with hybrid prompts and frequency fusion, beats task-specific radio classifiers and denoisers on most benchmarks.","keywords":["cognitive radio","large language models","radio signal classification","token reprogramming","hybrid prompting","frequency-attuned fusion","signal denoising","few-shot learning"],"falsifier":"Retrain every baseline under the same multi-dataset pretraining protocol that RadioLLM receives and compare on the held-out test splits; if the baselines match or exceed RadioLLM's reported OA, Kappa, and SSIM numbers, the claimed advantage of HPTR and FAF over existing methods is not established.","tokens_in":19900,"feed_emoji":"📡","tokens_out":12175,"duration_ms":93216,"temperature":0.7,"pith_summary":"This paper tries to establish that a large language model can act as a single universal engine for cognitive radio tasks. It proposes RadioLLM, which feeds raw I/Q radio samples to a GPT-2 backbone after reprogramming them into token embeddings and prefixing them with hybrid prompts that compress expert knowledge into a few retrieved semantic anchors. A frequency-attuned fusion module injects CNN-extracted high-frequency features so the transformer does not lose transient and phase-detail information. Across seven classification benchmarks and three denoising benchmarks, the paper reports the highest overall accuracy and Kappa on six of the seven classification sets and the highest SSIM on all three denoising sets. If correct, radio-signal applications would no longer need a separate task-specific network for every modulation family and noise regime.","feed_headline":"One pretrained language model tops most radio benchmarks","feed_subtitle":"Hybrid prompts and frequency fusion let it beat task-specific classifiers and denoisers.","key_machinery":"The load-bearing machinery is the pair of modules that make an LLM accept radio signals. Hybrid Prompt and Token Reprogramming (HPTR) does two things: it selects the top-K embeddings from a pretrained word-token space that are most similar to a concise hardware prompt, forming a short hybrid prefix, and it uses a multi-head cross-attention layer, with raw I/Q patches as queries and the anchor embeddings as keys and values, to turn signal patches into LLM-compatible tokens without a natural-language intermediate. Frequency-Attuned Fusion (FAF) runs the raw signal through three convolutional high-frequency extraction layers and fuses those local features with the reprogrammed tokens before the LLM, compensating for the attention mechanism's bias toward low-frequency global structure. The output side is a lightweight decoder that can either reconstruct the denoised I/Q signal or feed pooled features to a linear classification head.","core_discovery":"The central claim is that a GPT-2 model, mostly frozen and adapted with LoRA, can master radio signal classification and denoising when its input is reprogrammed and its attention is supplemented with high-frequency information. Specifically, RadioLLM obtains the best overall accuracy on RML16A (58.10), RML16B (58.35), RML16C (68.19), RML22 (59.39), RML18A (52.03), and Wi-Fi (35.41) under a 100-shot labeling regime, and the best SSIM on all three RML denoising tasks (0.838, 0.893, and 0.846). The authors attribute the gains to pretraining across multiple radio datasets, the hybrid prompt's replacement of verbose templates with top-K semantic anchors, and the FAF module's correction of the transformer's low-frequency bias. They also report that on ADS-B the model is not the top performer, and that the remaining confusions are between similar modulation families such as 16QAM/64QAM and AM-DSB/WBFM.","pith_inferences":["The reported margins may largely reflect in-distribution pretraining rather than HPTR and FAF, because RadioLLM is pretrained on the same benchmark datasets it is later tested on; an out-of-distribution holdout would settle this.","The same hybrid-prompt and token-reprogramming design could be applied to other physical-layer signals such as radar, sonar, or biomedical I/Q streams by swapping the anchor embedding space, though the paper does not test these.","If the semantic anchors are the main source of the improvement, a smaller non-LLM encoder conditioned on the same anchors might reproduce most of the gain, which would weaken the case that a full language model is necessary; this is an untested hypothesis.","The FAF design suggests a direct modification of transformer attention to be frequency-aware could yield similar or better low-SNR performance without a separate CNN branch, which the paper does not explore."],"forward_implications":["If correct, a single LLM backbone with lightweight task heads can replace separate radio classification and denoising networks, reducing deployment complexity in cognitive radio systems.","The hybrid prompt's top-K retrieval shortens the prompt prefix, and the paper reports 31.85% faster inference with a 0.85% accuracy gain, implying prompt compression matters for latency-sensitive radio settings.","FAF's success implies transformer-based radio models should not rely on attention alone; injecting convolutional high-frequency features is a transferable recipe for low-SNR robustness.","The gains grow as labeled samples increase, so the model is most useful in semi-supervised regimes where a large unlabeled signal pool can be pretrained on, rather than in extreme 1-shot settings."],"supporting_citations":[{"why":"Supplies the cross-attention token-reprogramming method that maps non-text sequences into LLM-compatible embeddings, the basis of RadioLLM's signal tokenization.","marker":"[11]"},{"why":"Motivates retrieving semantic anchors from the LLM's embedding space to build prompts, the idea behind the hybrid prompt's top-K retrieval.","marker":"[23]"},{"why":"Provides the low-rank adaptation technique used to fine-tune the GPT-2 backbone while keeping most parameters frozen.","marker":"[24]"},{"why":"Defines the RadioML2016.10a and 10b datasets used for both classification benchmarks and pretraining.","marker":"[25]"},{"why":"Source of the RadioML2016.04c dataset used as a classification and denoising benchmark.","marker":"[26]"},{"why":"Source of the RadioML2022 dataset used as a held-out transfer benchmark on which RadioLLM is not pretrained.","marker":"[27]"},{"why":"Source of the RadioML2018.01a large-scale classification benchmark.","marker":"[28]"},{"why":"Source of the real-world ADS-B dataset used to test classification outside the RML signal families.","marker":"[29]"},{"why":"Source of the real-world Wi-Fi dataset used to test over-the-air classification performance.","marker":"[30]"},{"why":"The strongest semi-supervised classification baseline that RadioLLM must beat, and one of the two baselines given a multi-domain pretraining comparison.","marker":"[38]"}],"fun_headline_variants":["GPT-2 with hybrid prompts tops most radio benchmarks","Frozen LLM beats task-specific radio models via reprogramming","RadioLLM: language model excels at signal classification and denoising","Token reprogramming turns LLM into radio signal specialist","Hybrid prompt and frequency fusion give LLM radio superiority"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The evaluation assumes that pretraining RadioLLM on the same datasets on which it is later benchmarked (RML16A/B/C, RML18A, Wi-Fi, and ADS-B) does not give it an unfair advantage, since most baselines receive only the 100 labeled examples and only TcssAMR and SemiAMC receive a multi-domain pretraining comparison on three of the datasets.","fun_headline_variants_meta":{"raw":{"variants":["GPT-2 with hybrid prompts tops most radio benchmarks","Frozen LLM beats task-specific radio models via reprogramming","RadioLLM: language model excels at signal classification and denoising","Token reprogramming turns LLM into radio signal specialist","Hybrid prompt and frequency fusion give LLM radio superiority"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000228,"raw_usage":{"total_tokens":1463,"prompt_tokens":921,"completion_tokens":542,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":537,"completion_tokens_details":{"reasoning_tokens":473}},"tokens_in":537,"tokens_out":542,"duration_ms":5281,"temperature":1.0,"reasoning_tokens":473,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T10:53:06.650297+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Retrain every baseline under the same multi-dataset pretraining protocol that RadioLLM receives and compare on the held-out test splits; if the baselines match or exceed RadioLLM's reported OA, Kappa, and SSIM numbers, the claimed advantage of HPTR and FAF over existing methods is not established.","supporting_citations":[{"cited_title":"Time-LLM: Time series forecasting by reprogramming large language models,","cited_arxiv_id":null,"evidence_quote":"Supplies the cross-attention token-reprogramming method that maps non-text sequences into LLM-compatible embeddings, the basis of RadioLLM's signal tokenization."},{"cited_title":"S2ip-llm: Semantic space informed prompt learning with llm for time series forecasting,","cited_arxiv_id":null,"evidence_quote":"Motivates retrieving semantic anchors from the LLM's embedding space to build prompts, the idea behind the hybrid prompt's top-K retrieval."},{"cited_title":"Rml22: Realistic dataset generation for wireless modulation classification,","cited_arxiv_id":null,"evidence_quote":"Source of the RadioML2022 dataset used as a held-out transfer benchmark on which RadioLLM is not pretrained."},{"cited_title":"Large-scale real-world radio signal recognition with deep learning,","cited_arxiv_id":null,"evidence_quote":"Source of the real-world ADS-B dataset used to test classification outside the RML signal families."},{"cited_title":"Oracle: Optimized radio classification through convo- lutional neural networks,","cited_arxiv_id":null,"evidence_quote":"Source of the real-world Wi-Fi dataset used to test over-the-air classification performance."},{"cited_title":"A transformer- based contrastive semi-supervised learning framework for automatic modulation recognition,","cited_arxiv_id":null,"evidence_quote":"The strongest semi-supervised classification baseline that RadioLLM must beat, and one of the two baselines given a multi-domain pretraining comparison."}],"review_version":1}