Pith. sign in

REVIEW 1 cited by

Beam Selection in ISAC using Contextual Bandit with Multi-modal Transformer and Transfer Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.08937 v1 pith:EQI6NTOT submitted 2025-03-11 eess.SP cs.LG

classification eess.SPcs.LG
keywords isacsensingbeamcommunicationlearningmodelmulti-modaltransformer
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Sixth generation (6G) wireless technology is anticipated to introduce Integrated Sensing and Communication (ISAC) as a transformative paradigm. ISAC unifies wireless communication and RADAR or other forms of sensing to optimize spectral and hardware resources. This paper presents a pioneering framework that leverages ISAC sensing data to enhance beam selection processes in complex indoor environments. By integrating multi-modal transformer models with a multi-agent contextual bandit algorithm, our approach utilizes ISAC sensing data to improve communication performance and achieves high spectral efficiency (SE). Specifically, the multi-modal transformer can capture inter-modal relationships, enhancing model generalization across diverse scenarios. Experimental evaluations on the DeepSense 6G dataset demonstrate that our model outperforms traditional deep reinforcement learning (DRL) methods, achieving superior beam prediction accuracy and adaptability. In the single-user scenario, we achieve an average SE regret improvement of 49.6% as compared to DRL. Furthermore, we employ transfer reinforcement learning to reduce training time and improve model performance in multi-user environments. In the multi-user scenario, this approach enhances the average SE regret, which is a measure to demonstrate how far the learned policy is from the optimal SE policy, by 19.7% compared to training from scratch, even when the latter is trained 100 times longer.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Foundation Model-Aided Deep Reinforcement Learning for RIS-Assisted Wireless Communication

    eess.SP 2025-06 reject novelty 4.0 of 10

    A fine-tuned wireless foundation model provides channel embeddings that feed a DDPG agent, which reportedly improves spectral efficiency over DRL with raw CSI and over beam sweeping in DeepMIMO simulation.

Pith tools