Pith. sign in

REVIEW 1 cited by

An Empirical Study of NetOps Capability of Pre-Trained Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.05557 v3 pith:RLQ4NY3A submitted 2023-09-11 cs.CL cs.AIcs.NI

classification cs.CLcs.AIcs.NI
keywords netopsllmsnetevalcapabilitiesmodelscapabilityhoweverlanguage
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Nowadays, the versatile capabilities of Pre-trained Large Language Models (LLMs) have attracted much attention from the industry. However, some vertical domains are more interested in the in-domain capabilities of LLMs. For the Networks domain, we present NetEval, an evaluation set for measuring the comprehensive capabilities of LLMs in Network Operations (NetOps). NetEval is designed for evaluating the commonsense knowledge and inference ability in NetOps in a multi-lingual context. NetEval consists of 5,732 questions about NetOps, covering five different sub-domains of NetOps. With NetEval, we systematically evaluate the NetOps capability of 26 publicly available LLMs. The results show that only GPT-4 can achieve a performance competitive to humans. However, some open models like LLaMA 2 demonstrate significant potential.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Hierarchical Debate-Based Large Language Model (LLM) for Complex Task Planning of 6G Network Management

    eess.SY 2025-06 conditional novelty 4.0 of 10

    A hierarchical debate framework, in which LLMs first decompose a 6G task and then refine each sub-task, improves keyword coverage over one-shot and regular single-level debate on the 6GPlan benchmark.

Pith tools