Pith. sign in

REVIEW 1 cited by

Can ChatGPT advance software testing intelligence? An experience report on metamorphic testing

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.19204 v2 pith:M4L4DH6S submitted 2023-10-30 cs.SE cs.AIcs.HC

classification cs.SEcs.AIcs.HC
keywords testingchatgptintelligencesoftwarecandidateshumanmetamorphicused
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

While ChatGPT is a well-known artificial intelligence chatbot being used to answer human's questions, one may want to discover its potential in advancing software testing. We examine the capability of ChatGPT in advancing the intelligence of software testing through a case study on metamorphic testing (MT), a state-of-the-art software testing technique. We ask ChatGPT to generate candidates of metamorphic relations (MRs), which are basically necessary properties of the object program and which traditionally require human intelligence to identify. These MR candidates are then evaluated in terms of correctness by domain experts. We show that ChatGPT can be used to generate new correct MRs to test several software systems. Having said that, the majority of MR candidates are either defined vaguely or incorrect, especially for systems that have never been tested with MT. ChatGPT can be used to advance software testing intelligence by proposing MR candidates that can be later adopted for implementing tests; but human intelligence should still inevitably be involved to justify and rectify their correctness.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. System Test Case Design from Requirements Specifications: Insights and Challenges of Using ChatGPT

    cs.SE 2024-12 conditional novelty 4.0 of 10

    An empirical study reports that ChatGPT-4o Turbo generates mostly valid system test cases from SRS documents, with about 15% of valid tests being novel, but its redundancy detection has a 30% false positive rate.

Pith tools