Pith. sign in

REVIEW 6 cited by

Towards Evaluation Guidelines for Empirical Studies involving LLMs

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.07668 v3 pith:HRARFLBG submitted 2024-11-12 cs.SE

classification cs.SE
keywords llmsstudiesresearchempiricalengineeringsoftwareguidelinesinvolving
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In the short period since the release of ChatGPT, large language models (LLMs) have changed the software engineering research landscape. While there are numerous opportunities to use LLMs for supporting research or software engineering tasks, solid science needs rigorous empirical evaluations. However, so far, there are no specific guidelines for conducting and assessing studies involving LLMs in software engineering research. Our focus is on empirical studies that either use LLMs as part of the research process or studies that evaluate existing or new tools that are based on LLMs. This paper contributes the first set of holistic guidelines for such studies. Our goal is to start a discussion in the software engineering research community to reach a common understanding of our standards for high-quality empirical studies involving LLMs.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Methodological Framework for LLM-Based Mining of Software Repositories

    cs.SE 2025-08 conditional novelty 5.0 of 10

    A rapid review and survey of LLM-based repository mining yield a threat-mitigation map and the six-stage PRIMES 2.0 framework for conducting such studies.

  2. Investigating the Use of LLMs for Evidence Briefings Generation in Software Engineering

    cs.SE 2025-07 unverdicted novelty 5.0 of 10

    A registered report protocol for comparing LLM-generated and human-made software engineering evidence briefings is laid out, but no experimental results are reported yet.

  3. ReqBrain: Task-Specific Instruction Tuning of LLMs for AI-Assisted Requirements Generation

    cs.SE 2025-05 conditional novelty 5.0 of 10

    ReqBrain, a LoRA-fine-tuned Zephyr-7b-beta model, produces software requirements that human evaluators could not reliably tell apart from human-authored ones, with automatic metrics favoring it over untuned ChatGPT-4o.

  4. Preliminary Guidelines for Using and Evaluating GenAI Tools to Support Systematic Literature Reviews

    cs.SE 2026-07 conditional novelty 4.5 of 10

    GUEST gives SE researchers process recommendations for planning, conducting, and reporting GenAI-supported SLRs and independent GenAI tool evaluations under mandatory human oversight.

  5. Empowering Computing Education Researchers Through LLM-Assisted Content Analysis

    cs.CL 2025-08 conditional novelty 4.0 of 10

    The paper proposes LACA, a reproducible protocol in which humans build a codebook, an LLM performs deductive coding, and interrater reliability checks gate whether the LLM can code the full dataset.

  6. Get on the Train or be Left on the Station: Using LLMs for Software Engineering Research

    cs.SE 2025-06 conditional novelty 4.0 of 10

    The paper maps LLM impacts onto the SE research pipeline with McLuhan's Tetrad, predicting enhancements, obsolescence, retrievals, and reversals, and calls for community action.

Pith tools