Pith. sign in

REVIEW 2 cited by

Evaluating ChatGPT on Nuclear Domain-Specific Data

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.00090 v1 pith:4V6IWH2N submitted 2024-08-26 cs.CL cs.AI

classification cs.CLcs.AI
keywords chatgptdomain-specificnuclearaccuracydataevaluatinggeneratedgenerating
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper examines the application of ChatGPT, a large language model (LLM), for question-and-answer (Q&A) tasks in the highly specialized field of nuclear data. The primary focus is on evaluating ChatGPT's performance on a curated test dataset, comparing the outcomes of a standalone LLM with those generated through a Retrieval Augmented Generation (RAG) approach. LLMs, despite their recent advancements, are prone to generating incorrect or 'hallucinated' information, which is a significant limitation in applications requiring high accuracy and reliability. This study explores the potential of utilizing RAG in LLMs, a method that integrates external knowledge bases and sophisticated retrieval techniques to enhance the accuracy and relevance of generated outputs. In this context, the paper evaluates ChatGPT's ability to answer domain-specific questions, employing two methodologies: A) direct response from the LLM, and B) response from the LLM within a RAG framework. The effectiveness of these methods is assessed through a dual mechanism of human and LLM evaluation, scoring the responses for correctness and other metrics. The findings underscore the improvement in performance when incorporating a RAG pipeline in an LLM, particularly in generating more accurate and contextually appropriate responses for nuclear domain-specific queries. Additionally, the paper highlights alternative approaches to further refine and improve the quality of answers in such specialized domains.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. AI-Assisted Knowledge Access for Legacy Enterprise Asset Management in Energy Operations: A Practical Retrieval System

    cs.IR 2026-06 conditional novelty 4.0 of 10

    A semantic-enrichment-based retrieval assistant improved retrieval quality and cut median task time from 14.2 to 8.3 minutes for legacy asset-management knowledge tasks in an energy utility pilot.

  2. From Naive RAG to Deep Agentic Retrieval: An Evolving Context Engineering Pipeline for Regulatory Compliance

    cs.IR 2026-06 conditional novelty 3.0 of 10

    OPG's production RAG system evolved into a cost-aware multi-agent retrieval pipeline (PEA-CAE), which the authors argue is a better investment than fine-tuning for evolving regulatory corpora.

Pith tools