StabilizerBench is a new benchmark for evaluating AI agents on generating, optimizing, and making fault-tolerant stabilizer circuits for quantum error correction, with efficient verification and multi-tier scoring.
Qiskit HumanEval: An evaluation benchmark for quantum code generative models
6 Pith papers cite this work. Polarity classification is still indexing.
years
2026 6representative citing papers
QBugLM introduces an agentic benchmarking framework for LLM debugging of quantum software and shows iterative prompting greatly improves repair rates on injected bugs.
A taxonomy-guided RAG system with LLMs reduces hallucinations and improves migration suggestions for Qiskit code compared to unconstrained retrieval.
A layered framework with physical gatekeepers, fidelity analysis against reference VQE circuits, and a consistency metric identifies five LLM failure modes in quantum circuit generation and reveals that some apparent model errors originated in the evaluation harness itself.
Adapts QuantumKatas to Qiskit yielding a 350-task benchmark across 26 categories and evaluates 16 LLMs in 39,200 runs, reporting performance gaps and prompting effects.
Retrieval over a 13,389-example verified PennyLane corpus raises QHack pass@5 from 36/43/24% to 64/68/52% across 2022–2024 with Claude Sonnet 4.6.
citing papers explorer
-
StabilizerBench: A Benchmark for AI-Assisted Quantum Error Correction Circuit Synthesis
StabilizerBench is a new benchmark for evaluating AI agents on generating, optimizing, and making fault-tolerant stabilizer circuits for quantum error correction, with efficient verification and multi-tier scoring.
-
QBugLM: An Agentic Benchmarking Framework for LLM-based Quantum Software Debugging
QBugLM introduces an agentic benchmarking framework for LLM debugging of quantum software and shows iterative prompting greatly improves repair rates on injected bugs.
-
Qiskit Code Migration with LLMs
A taxonomy-guided RAG system with LLMs reduces hallucinations and improves migration suggestions for Qiskit code compared to unconstrained retrieval.
-
Gatekeepers and Hallucinations: A Layered Evaluation Framework for LLM-Driven Quantum Circuit Generation
A layered framework with physical gatekeepers, fidelity analysis against reference VQE circuits, and a consistency metric identifies five LLM failure modes in quantum circuit generation and reveals that some apparent model errors originated in the evaluation harness itself.
-
Qiskit QuantumKatas: Adapting Microsoft's Quantum Computing exercises for LLM evaluation
Adapts QuantumKatas to Qiskit yielding a 350-task benchmark across 26 categories and evaluates 16 LLMs in 39,200 runs, reporting performance gaps and prompting effects.
-
PennySynth: RAG-Driven Data Synthesis for Automated Quantum Code Generation
Retrieval over a 13,389-example verified PennyLane corpus raises QHack pass@5 from 36/43/24% to 64/68/52% across 2022–2024 with Claude Sonnet 4.6.