Pith. sign in

REVIEW 4 cited by

BeamLLM: Vision-Empowered mmWave Beam Prediction with Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.10432 v2 pith:JGQL5GVT submitted 2025-03-13 cs.LG cs.CL

classification cs.LGcs.CL
keywords predictionllmsmmwavemodelsaccuracybeambeamllmfeatures
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this paper, we propose BeamLLM, a vision-aided millimeter-wave (mmWave) beam prediction framework leveraging large language models (LLMs) to address the challenges of high training overhead and latency in mmWave communication systems. By combining computer vision (CV) with LLMs' cross-modal reasoning capabilities, the framework extracts user equipment (UE) positional features from RGB images and aligns visual-temporal features with LLMs' semantic space through reprogramming techniques. Evaluated on a realistic vehicle-to-infrastructure (V2I) scenario, the proposed method achieves 61.01% top-1 accuracy and 97.39% top-3 accuracy in standard prediction tasks, significantly outperforming traditional deep learning models. In few-shot prediction scenarios, the performance degradation is limited to 12.56% (top-1) and 5.55% (top-3) from time sample 1 to 10, demonstrating superior prediction capability.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LVM4CSI: Enabling Direct Application of Pre-Trained Large Vision Models for Wireless Channel Tasks

    cs.IT 2025-07 conditional novelty 6.0 of 10

    A frozen pre-trained vision model can extract wireless channel paths and features, beating conventional estimators in channel estimation and matching specialized networks in sensing with far fewer trainable parameters.

  2. BERT4beam: Large AI Model Enabled Generalized Beamforming Optimization

    eess.SY 2025-09 conditional novelty 5.0 of 10

    A BERT-based transformer, BERT4beam, learns to output beamforming vectors from CSI and achieves near-SCA performance across multiple MU-MISO tasks and system scales.

  3. Vehicle-to-Infrastructure Collaborative Spatial Perception via Multimodal Large Language Models

    cs.LG 2025-09 conditional novelty 5.0 of 10

    Multi-car bird's-eye view tokens injected into a frozen vision-language model improve simulated V2I link prediction accuracy by up to 13.9 points on average.

  4. M2BeamLLM: Multimodal Sensing-empowered mmWave Beam Prediction with Large Language Models

    cs.CL 2025-06 conditional novelty 5.0 of 10

    M2BeamLLM combines four sensing modalities with a lightly fine-tuned GPT-2 backbone and reports 68.9% top-1 beam prediction accuracy on DeepSense 6G Scenario 32, beating the compared baselines by up to 13.9 percentage points.

Pith tools