Pith. sign in

REVIEW 4 cited by

Qiskit HumanEval: An Evaluation Benchmark For Quantum Code Generative Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.14712 v1 pith:CVECA3U5 submitted 2024-06-20 quant-ph cs.AI

classification quant-phcs.AI
keywords quantumcodeqiskitbenchmarkdatasetdevelopmenthumanevalllms
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Quantum programs are typically developed using quantum Software Development Kits (SDKs). The rapid advancement of quantum computing necessitates new tools to streamline this development process, and one such tool could be Generative Artificial intelligence (GenAI). In this study, we introduce and use the Qiskit HumanEval dataset, a hand-curated collection of tasks designed to benchmark the ability of Large Language Models (LLMs) to produce quantum code using Qiskit - a quantum SDK. This dataset consists of more than 100 quantum computing tasks, each accompanied by a prompt, a canonical solution, a comprehensive test case, and a difficulty scale to evaluate the correctness of the generated solutions. We systematically assess the performance of a set of LLMs against the Qiskit HumanEval dataset's tasks and focus on the models ability in producing executable quantum code. Our findings not only demonstrate the feasibility of using LLMs for generating quantum code but also establish a new benchmark for ongoing advancements in the field and encourage further exploration and development of GenAI-driven tools for quantum code generation.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Benchmarking API Drift in LLM-Generated Quantum Code Across Successive SDK Versions

    cs.SE 2026-07 conditional novelty 6.5 of 10

    Across 17 models and three Qiskit versions, version-aligned Pass@1 ranges from 0.02 to 0.85, with v1.3 hardest and documentation repair only partly effective.

  2. QuTuner: Feature- and Learning-Guided Optimization Pass Tuning for Quantum Compilers

    quant-ph 2026-07 conditional novelty 6.0 of 10

    QuTuner retrieves and ranks full-space quantum optimization pass sequences from a large BO-built dataset using static features plus pass-response embeddings, then refines them with short Bayesian search.

  3. QHackBench: Benchmarking Large Language Models for Quantum Code Generation Using PennyLane Hackathon Challenges

    cs.AI 2025-06 conditional novelty 6.0 of 10

    A benchmark of 49 QHack PennyLane challenges shows LLMs solve at most 49 percent of tasks, retrieval augmentation usually does not help, and a multi-agent retry loop improves results.

  4. PennySynth: RAG-Driven Data Synthesis for Automated Quantum Code Generation

    cs.CL 2026-05 unverdicted novelty 5.0 of 10

    Retrieval over a 13,389-example verified PennyLane corpus raises QHack pass@5 from 36/43/24% to 64/68/52% across 2022–2024 with Claude Sonnet 4.6.

Pith tools