NuclearQAv2 is a hybrid-constructed benchmark dataset for evaluating LLM competence in nuclear engineering knowledge using three question types.
NuclearQA: A human-made benchmark for language models for the nuclear domain.arXiv preprint arXiv:2310.10920, 2023
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
years
2026 2verdicts
UNVERDICTED 2representative citing papers
LLM planning agent with dynamic KG state achieves 81.5% accuracy on 200 multi-hop questions from NuScale FSAR documents, outperforming non-planning RAG baselines by up to 38pp.
citing papers explorer
-
NuclearQAv2: A Structured Benchmark for Evaluating Domain-Science Competence in Large Language Models
NuclearQAv2 is a hybrid-constructed benchmark dataset for evaluating LLM competence in nuclear engineering knowledge using three question types.
-
LLM-Guided Planning for Multi-hop Reasoning over Multimodal Nuclear Regulatory Documents
LLM planning agent with dynamic KG state achieves 81.5% accuracy on 200 multi-hop questions from NuScale FSAR documents, outperforming non-planning RAG baselines by up to 38pp.