REVIEW 4 cited by
BeamLLM: Vision-Empowered mmWave Beam Prediction with Large Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
In this paper, we propose BeamLLM, a vision-aided millimeter-wave (mmWave) beam prediction framework leveraging large language models (LLMs) to address the challenges of high training overhead and latency in mmWave communication systems. By combining computer vision (CV) with LLMs' cross-modal reasoning capabilities, the framework extracts user equipment (UE) positional features from RGB images and aligns visual-temporal features with LLMs' semantic space through reprogramming techniques. Evaluated on a realistic vehicle-to-infrastructure (V2I) scenario, the proposed method achieves 61.01% top-1 accuracy and 97.39% top-3 accuracy in standard prediction tasks, significantly outperforming traditional deep learning models. In few-shot prediction scenarios, the performance degradation is limited to 12.56% (top-1) and 5.55% (top-3) from time sample 1 to 10, demonstrating superior prediction capability.
Forward citations
Cited by 4 Pith papers
-
LVM4CSI: Enabling Direct Application of Pre-Trained Large Vision Models for Wireless Channel Tasks
A frozen pre-trained vision model can extract wireless channel paths and features, beating conventional estimators in channel estimation and matching specialized networks in sensing with far fewer trainable parameters.
-
BERT4beam: Large AI Model Enabled Generalized Beamforming Optimization
A BERT-based transformer, BERT4beam, learns to output beamforming vectors from CSI and achieves near-SCA performance across multiple MU-MISO tasks and system scales.
-
Vehicle-to-Infrastructure Collaborative Spatial Perception via Multimodal Large Language Models
Multi-car bird's-eye view tokens injected into a frozen vision-language model improve simulated V2I link prediction accuracy by up to 13.9 points on average.
-
M2BeamLLM: Multimodal Sensing-empowered mmWave Beam Prediction with Large Language Models
M2BeamLLM combines four sensing modalities with a lightly fine-tuned GPT-2 backbone and reports 68.9% top-1 beam prediction accuracy on DeepSense 6G Scenario 32, beating the compared baselines by up to 13.9 percentage points.
Discussion (0). Continue with ORCID to comment.