Pith. sign in

REVIEW 4 cited by

Building 6G Radio Foundation Models with Transformer Architectures

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.09996 v1 pith:BFPZ2EVD submitted 2024-11-15 eess.SP cs.AIcs.NI

classification eess.SPcs.AIcs.NI
keywords foundationmodelmodelsspectrogramlearningacrossactivitycompetitive
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Foundation deep learning (DL) models are general models, designed to learn general, robust and adaptable representations of their target modality, enabling finetuning across a range of downstream tasks. These models are pretrained on large, unlabeled datasets using self-supervised learning (SSL). Foundation models have demonstrated better generalization than traditional supervised approaches, a critical requirement for wireless communications where the dynamic environment demands model adaptability. In this work, we propose and demonstrate the effectiveness of a Vision Transformer (ViT) as a radio foundation model for spectrogram learning. We introduce a Masked Spectrogram Modeling (MSM) approach to pretrain the ViT in a self-supervised fashion. We evaluate the ViT-based foundation model on two downstream tasks: Channel State Information (CSI)-based Human Activity sensing and Spectrogram Segmentation. Experimental results demonstrate competitive performance to supervised training while generalizing across diverse domains. Notably, the pretrained ViT model outperforms a four-times larger model that is trained from scratch on the spectrogram segmentation task, while requiring significantly less training time, and achieves competitive performance on the CSI-based human activity sensing task. This work demonstrates the effectiveness of ViT with MSM for pretraining as a promising technique for scalable foundation model development in future 6G networks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. IQFM A Wireless Foundational Model for I/Q Streams in AI-Native 6G

    eess.SP 2025-06 conditional novelty 6.0 of 10

    A self-supervised encoder trained on raw multi-antenna I/Q data reaches strong few-shot accuracy on modulation, angle-of-arrival, beam prediction, and RF fingerprinting tasks.

  2. Radio-FM: A Foundation Model for Radio Signal Representation Learning and Its Applications

    eess.SP 2026-08 reject novelty 5.0 of 10

    Radio-FM pretrains dual-channel transformers on 15 radio datasets and claims state-of-the-art transfer on 13 of 15 benchmarks, though several evaluation datasets overlap with the pretraining data.

  3. Towards channel foundation models (CFMs): Motivations, methodologies and opportunities

    eess.SP 2025-07 conditional novelty 4.0 of 10

    A survey and position paper proposing channel foundation models, with experiments on two pretrained CSI models showing gains over a vanilla ViT baseline.

  4. Foundation Model Empowered Synesthesia of Machines (SoM): AI-native Intelligent Multi-Modal Sensing-Communication Integration

    eess.SP 2025-06 conditional novelty 4.0 of 10

    The paper proposes a systematic classification and two roadmaps for using foundation models (LLMs and wireless foundation models) to design Synesthesia of Machines systems for 6G, with preliminary case-study evidence ...

Pith tools