Pith. sign in

REVIEW 4 cited by

Multimodal Transformers for Wireless Communications: A Case Study in Beam Prediction

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.11811 v1 pith:2GXTN4K4 submitted 2023-09-21 eess.SP cs.AI

classification eess.SPcs.AI
keywords databeamfeaturemodalitiesmultimodalpredictionradarwireless
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Wireless communications at high-frequency bands with large antenna arrays face challenges in beam management, which can potentially be improved by multimodality sensing information from cameras, LiDAR, radar, and GPS. In this paper, we present a multimodal transformer deep learning framework for sensing-assisted beam prediction. We employ a convolutional neural network to extract the features from a sequence of images, point clouds, and radar raw data sampled over time. At each convolutional layer, we use transformer encoders to learn the hidden relations between feature tokens from different modalities and time instances over abstraction space and produce encoded vectors for the next-level feature extraction. We train the model on a combination of different modalities with supervised learning. We try to enhance the model over imbalanced data by utilizing focal loss and exponential moving average. We also evaluate data processing and augmentation techniques such as image enhancement, segmentation, background filtering, multimodal data flipping, radar signal transformation, and GPS angle calibration. Experimental results show that our solution trained on image and GPS data produces the best distance-based accuracy of predicted beams at 78.44%, with effective generalization to unseen day scenarios near 73% and night scenarios over 84%. This outperforms using other modalities and arbitrary data processing techniques, which demonstrates the effectiveness of transformers with feature fusion in performing radio beam prediction from images and GPS. Furthermore, our solution could be pretrained from large sequences of multimodality wireless data, on fine-tuning for multiple downstream radio network tasks.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Vehicle-to-Infrastructure Collaborative Spatial Perception via Multimodal Large Language Models

    cs.LG 2025-09 conditional novelty 5.0 of 10

    Multi-car bird's-eye view tokens injected into a frozen vision-language model improve simulated V2I link prediction accuracy by up to 13.9 points on average.

  2. Unreal is all you need: Multimodal ISAC Data Simulation with Only One Engine

    cs.CV 2025-07 conditional novelty 5.0 of 10

    Great-X reproduces Sionna's ray-tracing channel model inside Unreal Engine and provides Great-MSD, a multimodal UAV dataset for CSI localization.

  3. Multi-Modal Beamforming with Model Compression and Modality Generation for V2X Networks

    eess.SP 2025-06 conditional novelty 4.0 of 10

    A multi-modal transformer with module-aware pruning and a conditional VAE for missing-sensor reconstruction improves beam prediction accuracy on the DeepSense 6G dataset.

  4. From Large AI Models to Agentic AI: A Tutorial on Future Intelligent Communications

    cs.AI 2025-05 conditional novelty 2.0 of 10

    This paper is a broad tutorial on applying LAMs and agentic AI to 6G, largely restating existing research rather than introducing new results.

Pith tools