Pith. sign in

REVIEW 1 cited by

Evaluating the Process Modeling Abilities of Large Language Models -- Preliminary Foundations and Results

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.13520 v1 pith:SES2ZQC3 submitted 2025-03-14 cs.CL cs.LGcs.SE

classification cs.CLcs.LGcs.SE
keywords processabilitieschallengeslanguagemodelingmodelsresultstaken
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large language models (LLM) have revolutionized the processing of natural language. Although first benchmarks of the process modeling abilities of LLM are promising, it is currently under debate to what extent an LLM can generate good process models. In this contribution, we argue that the evaluation of the process modeling abilities of LLM is far from being trivial. Hence, available evaluation results must be taken carefully. For example, even in a simple scenario, not only the quality of a model should be taken into account, but also the costs and time needed for generation. Thus, an LLM does not generate one optimal solution, but a set of Pareto-optimal variants. Moreover, there are several further challenges which have to be taken into account, e.g. conceptualization of quality, validation of results, generalizability, and data leakage. We discuss these challenges in detail and discuss future experiments to tackle these challenges scientifically.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. What is the Best Process Model Representation? A Comparative Analysis for Process Modeling with Large Language Models

    cs.CL 2025-07 conditional novelty 6.0 of 10

    A new dataset and head-to-head comparison of nine process model representations with LLMs finds Mermaid best for general use and BPMN text best for generation.

Pith tools