Pith. sign in

REVIEW 2 cited by

TeleQnA: A Benchmark Dataset to Assess Large Language Models Telecommunications Knowledge

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.15051 v1 pith:4T6727KS submitted 2023-10-23 cs.IT cs.AIcs.LGmath.IT

classification cs.ITcs.AIcs.LGmath.IT
keywords datasetllmsknowledgetelecommodelsperformancequestionsactive
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce TeleQnA, the first benchmark dataset designed to evaluate the knowledge of Large Language Models (LLMs) in telecommunications. Comprising 10,000 questions and answers, this dataset draws from diverse sources, including standards and research articles. This paper outlines the automated question generation framework responsible for creating this dataset, along with how human input was integrated at various stages to ensure the quality of the questions. Afterwards, using the provided dataset, an evaluation is conducted to assess the capabilities of LLMs, including GPT-3.5 and GPT-4. The results highlight that these models struggle with complex standards related questions but exhibit proficiency in addressing general telecom-related inquiries. Additionally, our results showcase how incorporating telecom knowledge context significantly enhances their performance, thus shedding light on the need for a specialized telecom foundation model. Finally, the dataset is shared with active telecom professionals, whose performance is subsequently benchmarked against that of the LLMs. The findings illustrate that LLMs can rival the performance of active professionals in telecom knowledge, thanks to their capacity to process vast amounts of information, underscoring the potential of LLMs within this domain. The dataset has been made publicly accessible on GitHub.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. NextG-GPT: Leveraging GenAI for Advancing Wireless Networks and Communication Research

    cs.ET 2025-05 conditional novelty 5.0 of 10

    A RAG-enhanced LLM assistant for wireless research testbeds is built and evaluated, with LLaMa3.1-70B scoring best, though the abstract mislabels a faithfulness score as correctness.

  2. Edge Agentic AI Framework for Autonomous Network Optimisation in O-RAN

    eess.SP 2025-07 conditional novelty 4.0 of 10

    A simulated edge agentic AI framework with LSTM traffic prediction and tiered Tx power control reports zero network outages in high-stress 5G scenarios.

Pith tools