REVIEW 5 cited by
ORAN-Bench-13K: An Open Source Benchmark for Assessing LLMs in Open Radio Access Networks
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Large Language Models (LLMs) can revolutionize how we deploy and operate Open Radio Access Networks (O-RAN) by enhancing network analytics, anomaly detection, and code generation and significantly increasing the efficiency and reliability of a plethora of O-RAN tasks. In this paper, we present ORAN-Bench-13K, the first comprehensive benchmark designed to evaluate the performance of Large Language Models (LLMs) within the context of O-RAN. Our benchmark consists of 13,952 meticulously curated multiple-choice questions generated from 116 O-RAN specification documents. We leverage a novel three-stage LLM framework, and the questions are categorized into three distinct difficulties to cover a wide spectrum of ORAN-related knowledge. We thoroughly evaluate the performance of several state-of-the-art LLMs, including Gemini, Chat-GPT, and Mistral. Additionally, we propose ORANSight, a Retrieval-Augmented Generation (RAG)-based pipeline that demonstrates superior performance on ORAN-Bench-13K compared to other tested closed-source models. Our findings indicate that current popular LLM models are not proficient in O-RAN, highlighting the need for specialized models. We observed a noticeable performance improvement when incorporating the RAG-based ORANSight pipeline, with a Macro Accuracy of 0.784 and a Weighted Accuracy of 0.776, which was on average 21.55% and 22.59% better than the other tested LLMs.
Forward citations
Cited by 5 Pith papers
-
AI5GTest: AI-Driven Specification-Aware Automated Testing and Validation of 5G O-RAN Components
An LLM-based framework that generates expected O-RAN and 3GPP procedural flows from standards and validates captured signaling logs against them, reporting 100% accuracy on 15 testbed instances and under an hour per t...
-
NextG-GPT: Leveraging GenAI for Advancing Wireless Networks and Communication Research
A RAG-enhanced LLM assistant for wireless research testbeds is built and evaluated, with LLaMa3.1-70B scoring best, though the abstract mislabels a faithfulness score as correctness.
-
Benchmarking Vector, Graph and Hybrid Retrieval Augmented Generation (RAG) Pipelines for Open Radio Access Networks (ORAN)
On a 600-question subset of ORAN-Bench-13K, GraphRAG and Hybrid GraphRAG beat plain vector RAG on factual accuracy, but Hybrid GraphRAG scored below vector RAG on context relevance.
-
ORAN-GUIDE: RAG-Driven Prompt Learning for LLM-Augmented Reinforcement Learning in O-RAN Network Slicing
ORAN-GUIDE couples a domain-specific LLM prompt generator with a frozen GPT-2 encoder and learnable prompt tokens to improve multi-agent SAC sample efficiency in O-RAN slicing.
-
Prompt-Tuned LLM-Augmented DRL for Dynamic O-RAN Network Slicing
Prompt-tuned ORANSight state representations improve convergence and slice-level QoS for multi-agent SAC in a simulated O-RAN slicing environment, according to the reported ablation.
Discussion (0). Sign in to comment.