Pith. sign in

REVIEW 5 major objections 5 minor 124 references

Large Language Models for Software Engineering Diagrams: A Systematic Review of UML and ER modelling

T0 review · 5 major / 5 minor · reviewed 2026-08-01 · deepseek-v4-flash

Pith's one-line read This review of 64 studies claims that LLM-based software diagramming is concentrated on UML class-diagram generation from natural language, while behavioural and ER data modelling, transformation, quality assurance, and shared benchmarks ar

desk verdict A useful first systematic map of LLM-based diagram modelling, with a defensible taxonomy and plausible qualitative findings — but the corpus needs auditing before the headline numbers can be trusted. read the letter →

arxiv 2607.26100 v1 pith:UOCP5XN5 submitted 2026-07-28 cs.SE

classification cs.SE
keywords systematicliteraturereviewlargelanguagemodelsUMLdiagramsentity-relationshipsoftwaremodellingdiagramgenerationevaluationpracticesclass
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This systematic review synthesises 64 studies from 2023–2025 that apply large language models to UML and entity-relationship diagram tasks. Its central claim is a map of the field: research strongly favours UML software modelling—class diagrams above all—while behavioural diagrams and ER data modelling lag; natural-language diagram construction is the dominant task; GPT-style models dominate the technical choices; and evaluation is heterogeneous, with diverse metrics, custom datasets, and little benchmark reuse. The paper also documents recurring reported limitations: hallucinated diagram elements, semantic inaccuracies, sensitivity to prompt formulation, and reproducibility constraints. If this map is right, it gives the community a concrete agenda: developing shared benchmarks, strengthening evaluation protocols, broadening diagram coverage, and linking generation to formal modelling semantics.

What carries the argument

The working object is the multi-dimensional classification scheme (Fig. 4 taxonomy) into which every one of the 64 studies is placed: diagram type (behavioural vs structural UML vs ER), task category (construction, transformation, evaluation, understanding/assistance), technical configuration (model family, prompting strategy, architecture), and evaluation approach (metrics and datasets). The survey's distributions—which diagram types dominate, which tasks are rare, which models are used—are simply the counts over this taxonomy, so the taxonomy is the load-bearing instrument: the findings are true only if the categories are applied consistently and the corpus is complete. A secondary mechani

What would settle it

Re-run the stated protocol—same search string, databases, inclusion and exclusion criteria, and iterative snowballing—and check that the corpus reproduces exactly these 64 studies; then have a second team independently apply the Fig. 4 taxonomy to the same papers and compare category assignments. As a spot check, test the peer-review-only exclusion against the reference list, which includes a preprint-only study; if adding or removing that study changes the diagram-type or task distributions, the map is not stable under a defensible variant of the criteria.

Watch

Extended reading notes

Core claim

The authors claim that, across 64 studies from 2023–2025, LLM-based diagramming is concentrated on UML class-diagram generation from natural language using GPT-family models, while behavioural diagrams, entity-relationship and other data modelling, transformation, quality assurance, and multi-view consistency checking are underrepresented. They further claim that evaluation is heterogeneous—diverse metrics, custom datasets, little benchmark reuse, inconsistent statistical reporting—and that this review is the first systematic, diagram-centric synthesis of the literature, with its taxonomy of diagram types, tasks, techniques, and evaluation practices as the organising structure.

Load-bearing premise

The load-bearing premise is that the 64 studies assembled through the stated search string, screening criteria, and snowballing are a complete and representative sample of LLM diagram-modelling research; if retrieval or screening systematically selects GPT-focused, class-diagram papers, the reported concentrations describe the search rather than the field.

Editorial extensions

If this is right

  • Research effort and funding in LLM diagramming will need to shift if the field is to cover behavioural diagrams and ER data modelling, the corners the review shows are most neglected.
  • Because no shared benchmarks or common matching criteria exist, no current result is directly comparable to another; building a multi-diagram, multi-solution benchmark is the most direct structural improvement the review implies.
  • Most systems are standalone, prompt-only invocations of GPT-family models, so integrating validation infrastructure—formal diagram rules, constraint checkers, statistical reporting—is the most plausible route from demonstrations to reliable tools.
  • The prevalence of educational applications (tutoring, grading) suggests LLM diagram tools are most mature in teaching, so near-term deployment may happen in the classroom before industrial design.
  • The persistent limitation reports across diagram types indicate that hallucinated elements, prompt sensitivity, and non-determinism are current-generation LLM properties rather than artifacts of a single approach.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The GPT-prevalence result may be partly an artifact of the search string, which explicitly includes 'GPT' and 'ChatGPT' as query terms but no open-weights model names; a replication with open-model terms would test whether vendor concentration is a property of the field or of the retrieval.
  • The ER-underrepresentation result rests on just two included studies, both dealing with ER evaluation; a broader search of database and data-modelling venues would show whether the ER gap is a property of the literature or of this review's selection.
  • The paper's call for a benchmark with multiple valid reference solutions per specification implies a testable prediction: LLMs' measured quality relative to expert humans should improve under such a benchmark, because single-gold-standard evaluation penalises legitimate design alternatives.
  • The absence of an auditable corpus—no screening counts, no excluded list, no extraction data—means the map can be falsified only by replicating the protocol; publishing the extraction data would turn this review into a reusable resource.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 5 minor

Summary. This manuscript presents a systematic literature review of 64 studies (claimed) on the use of large language models for UML and entity-relationship diagram modelling. It reports a protocol using six digital libraries, a defined search string, inclusion/exclusion criteria, two-reviewer screening with Cohen's kappa 0.738, and backward/forward snowballing. The results are organized around five research questions covering diagram types, modelling tasks, technical configurations, evaluation practices, and reported limitations. The paper's central claims are that class diagrams dominate, that diagram construction from natural language is the primary task, that GPT-based models are heavily prevalent, that behavioural diagrams and ER modelling are underrepresented, and that evaluation is heterogeneous with little benchmark reuse. The paper positions itself as the first systematic synthesis of LLM-based diagram modelling research.

Significance. If the corpus is complete and the classification is reliable, this would be a useful first systematic map of a young but fast-growing area. The paper contributes a taxonomy of diagram types and tasks, a structured overview of technical integration patterns, a summary of evaluation metrics and public datasets, and a set of future-research directions. The authors report a transparent search protocol, snowballing, and a threats-to-validity section, and Table 6 lists concrete public dataset URLs, which strengthens reproducibility. However, the value of the contribution rises or falls on the integrity of the 64-study corpus and the consistency of the manual classification that feeds every distributional headline.

major comments (5)
  1. [§4.1 and reference list] The corpus is not auditable as presented. References [61] and [79] are the same study (Al-Ahmad et al., Information 16(7), 565, 2025), so the reported total of 64 unique studies is at least one too high. This duplicate also inflates counts for class diagrams, deployment diagrams, and educational support in Fig. 4 and Table 2. Additionally, the manuscript ships no screening counts, no excluded-paper list, and no extraction dataset, so completeness and representativeness cannot be checked. Please provide a PRISMA-style flow diagram, a full list of excluded studies with reasons, a deduplicated reference list, and re-computed distributional counts.
  2. [§3.1 search string] The search string includes 'GPT' and 'ChatGPT' as explicit OR terms, while the abstract's headline finding states that GPT-based models are 'heavily prevalent'. This operationalization can preferentially retrieve papers that merely mention these vendor names, and it may bias the model-prevalence result. Please report how many included studies were retrieved only because of the vendor-specific terms, or rerun the search with a generic LLM-only string, and discuss whether the GPT-dominance claim survives that sensitivity analysis.
  3. [Table 1 and references [48], [97]] There are two inclusion-consistency defects that affect the corpus definition. Exclusion criterion (5) forbids preprints, yet [48] is explicitly an arXiv preprint (arXiv:2404.17739). Also, the abstract and §3.1 describe the window as 2023–2025, but [97] carries a 2026 publication date (Information and Software Technology 190, 107955, 2026). Please clarify whether [48] has since appeared in a peer-reviewed venue and whether [97] is an online-first/issue-date artifact; otherwise the inclusion decisions and the temporal distribution are inconsistent with the stated rules.
  4. [§3.2 and Fig. 4 taxonomy] The reported Cohen's kappa of 0.738 concerns title/abstract study selection only. No inter-rater reliability is reported for the manual classification of studies into the Fig. 4 taxonomy (diagram types, task categories, technical approaches, evaluation practices). Since every RQ1–RQ5 distributional claim is computed from that classification, please either report agreement for data extraction/classification, or describe a consensus-coding procedure with a resolved-disagreement log. The assignment of borderline scope cases also needs to be documented: for example, [42] addresses SysML diagrams, which are outside the UML/ER scope stated in the abstract, and the inclusion criteria in Table 1 should make clear whether such studies are covered under 'domain-specific modelling notations'.
  5. [§4.4.1 and Table 5] Some technical claims reference the wrong items. §4.4.1 states 'In [25], LoRA is used to fine-tune open-source LLMs...', but [25] is Zhang et al.'s survey 'A survey on large language models for software engineering', not a primary empirical study. This suggests a reference indexing error that should be corrected. In §4.5.1, 'Mean Absolute Percentage Error (MAPE)' is described as an element-level deviation measure for diagram generation; as defined, it is unusual in this context and its operationalization and suitability should be clarified.
minor comments (5)
  1. [§4.4.1] The text refers to 'Table X for full references', but no Table X appears in the manuscript. Add the missing table or remove the pointer.
  2. [§4.5.1 / Table 5] Table 5 labels 'Human Evaluation (Rubrics / Likert)' and 'Agreement Measures' as distinct metric categories, but the latter is a subtype of the former; consider a hierarchy or merged row to avoid double-counting.
  3. [Throughout] There are several minor typographical issues: 'T able 1' in §3.1, 'modelling Accuracy' in §4.6, and inconsistent capitalization such as 'F ew-shot' in §4.4.2. A careful copy-edit is needed.
  4. [§4.2 / Fig. 5] Fig. 5's distribution would benefit from explicit numeric counts per diagram type; the text and Fig. 4 contain the citations, but the visual is hard to interpret without the actual frequencies.
  5. [References] Several accessible-dates (e.g., 'Accessed 2026-03-22') are inconsistent with the stated 2023–2025 review window; align the access-date policy with the search and snowballing deadline described in §3.2.

Circularity Check

1 steps flagged · score 3.0 of 10

Partially circular: the 'GPT models are heavily prevalent' headline is partially baked into the search string; the remaining distributional findings are corpus-grounded and not query-forced.

  1. fitted input called prediction [Section 3.1 (Paper Collection Strategy); Abstract and Section 4.4.1 / RQ3.1 (LLM Model Families, Fig. 9)]
    "We used the following search string consistently across all selected digital libraries: ("large language model" OR LLM OR "GPT" OR "ChatGPT" OR "generative AI" OR "Gen AI") AND ("UML" OR "class diagram" OR ... "ER diagram" OR "entity relationship model"). ... Abstract: "GPT-based models are heavily prevalent, raising concerns about reproducibility and vendor dependence.""

    The headline claim of GPT prevalence is computed over a corpus retrieved with a query that explicitly lists 'GPT' and 'ChatGPT' as OR search terms. Papers whose titles/abstracts mention GPT/ChatGPT are therefore preferentially surfaced, while LLM-based diagram work that does not name GPT in its metadata is less likely to be captured by the database searches. The model-family distribution reported in Fig. 9 / RQ3.1 is thus partially a projection of the query's own term choices rather than an independent measurement of the field, i.e., the 'prediction' (vendor dominance) is partly baked into the selection input. The reduction is partial, not complete: generic terms ('LLM', 'large language model', 'generative AI') also retrieve non-GPT studies, so the diagram-type, task, and evaluation findin

full rationale

This is a systematic review, so its derivation chain is selection -> classification -> synthesis. The central distributional claims (class-diagram dominance, behavioural/ER underrepresentation, construction-from-NL as the primary task, heterogeneous evaluation, limited benchmark reuse) are computed from the 64-study corpus and do not reduce to the search instrument: the query weights all UML/ER diagram terms equally and includes generic LLM terms, so these findings are genuine corpus properties rather than artifacts of the query. The taxonomy follows the external UML Reference Manual [103], selection agreement is quantified (kappa 0.738), and no uniqueness theorem or ansatz is imported from the authors' prior work. The one partial circularity is the 'GPT-based models are heavily prevalent' headline: the Section 3.1 search string explicitly contains 'GPT' and 'ChatGPT' as OR retrieval terms, so GPT-mentioning papers are preferentially surfaced and the reported model-family dominance is in part constructed by the selection instrument. The remaining flagged issues are validity/correctness concerns, not circularity: references [61] and [79] are the same study (inflating N to 64), [48] is an arXiv preprint despite exclusion criterion (5), [97] is dated 2026 while the abstract states 2023-2025, [42] treats SysML rather than UML/ER, and [93] is the authors' own ER study (one of only two), making 'ER underrepresentation' mildly self-referential but robust even without that inclusion. The absence of screening counts, an excluded-paper list, and an extraction dataset limits auditability, but it does not make the central map equivalent to its inputs by construction.

Assumptions & free parameters 4 free parameters · 8 assumptions · 2 invented entities

The review's findings are entirely a function of methodological choices: the §3.1 search string with vendor-specific OR terms, the Table 1 exclusion rules (peer-review-only, violated by [48]), and the authors' manual classification into the Fig. 4 taxonomy. No numbers are fitted to data, but hand-chosen cutoffs and stop rules act as free parameters of the corpus. The main constructs (task taxonomy, integration typology) are invented analytic entities without independent validation. Background assumptions are standard SLR assumptions (Keele guidelines, UML Reference Manual classification, kappa-based reliability); one is ad hoc to this paper: that 'GPT'/'ChatGPT' search terms give a representative sample of LLM-based diagram research.

free parameters (4)
  • Publication window start (January 2023) = 2023-01-01
    Hand-selected cutoff justified by GPT-3.5/4 release; determines corpus composition and underlies the '2023-2025' claim, which conflicts with included 2026-dated study [97].
  • Search-string vendor terms ('GPT', 'ChatGPT', 'generative AI') = included as OR terms
    Hand-chosen terms guarantee GPT-mentioning papers are retrieved, loading the 'GPT dominance / vendor dependence' finding; a corpus built on 'LLM' alone could differ.
  • Kappa agreement threshold interpretation = 0.738 ('substantial agreement')
    Reported to support selection reliability; the threshold reading is conventional and the screening counts behind it are not disclosed.
  • Snowballing stop rule = iterations until Jan 1, 2026; stop when no new papers
    Hand-set termination rule; corpus completeness depends on this judgment call, and the number of rounds is not reported.
assumptions (8)
  • domain assumption UML and ER are the two dominant modelling notations; BPMN/SysML can be excluded without distorting the picture
    Intro §1; contradicted internally by the inclusion of SysML study [42], so the scope premise is applied inconsistently.
  • domain assumption LLM-based diagram research effectively began with the public release of GPT-3.5/GPT-4; earlier work is out of scope
    §3.1 time-frame justification; hand-selected cutoff that shapes the corpus.
  • domain assumption Peer-reviewed publication is a valid quality filter for included studies
    Table 1 exclusion criterion (5); contradicted by inclusion of arXiv preprint [48].
  • domain assumption Two-reviewer screening with Cohen's kappa 0.738 yields reliable inclusion decisions
    §3.1; kappa is reported, but full screening counts and disagreement resolutions are not.
  • domain assumption The UML Reference Manual [103] classification of diagrams into structural/behavioral is the correct organizing scheme
    §4.2; a standard external taxonomy, though ER diagrams and 'architecture models' do not fit cleanly under it.
  • domain assumption Diagram type, task, technique, and evaluation categories can be reliably extracted from paper descriptions; borderline cases resolved by author judgment
    §7 construct validity admits interpretive judgment; no reliability metric is given for classification itself.
  • domain assumption Snowballing until no new papers appear yields a complete corpus
    §3.2; completeness claim depends on this rule plus the search-string coverage.
  • ad hoc to paper 'GPT'/'ChatGPT' as explicit search terms operationalize 'research using large language models for diagrams'
    §3.1; this query choice guarantees vendor-specific papers enter the corpus and partially forces the 'GPT prevalence' conclusion.
invented entities (2)
  • Four-category task taxonomy (construction, transformation, evaluation, understanding & assistance)
    purpose: Organize the reviewed studies into comparable categories; underlies all RQ2 findings and the 'construction dominates' conclusion
    An analytic construct defined by the authors for this review; no inter-rater reliability is reported for the classification itself (kappa covers study selection only), so the categories are not independently validated.
  • Technical integration typology (standalone, multi-stage, agent-based, RAG, human-in-the-loop)
    purpose: Classify how LLMs are embedded in modelling workflows (RQ3.3)
    A second analytic construct; categories are derived from the corpus and not validated against an external scheme.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Large Language Models for Software Engineering Diagrams: A Systematic Review of UML and ER modelling." pith.science (2026). https://pith.science/paper/UOCP5XN5

@misc{pith2026260726100,
  author       = {Pith},
  title        = {Pith review of: Large Language Models for Software Engineering Diagrams: A Systematic Review of UML and ER modelling},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/UOCP5XN5}},
  note         = {Machine review of arXiv:2607.26100}
}
read the original abstract

Large language models (LLMs) are increasingly applied to diagram-based software and data modelling. Among various modelling notations, UML and entity-relationship (ER) diagrams are the most widely adopted for software modelling and data modelling, respectively. Recent literature has investigated various applications of LLMs in diagram modelling; however, their effectiveness and limitations have not been extensively discussed. This systematic literature review analyses 64 studies published between 2023 and 2025, examining diagram coverage, modelling tasks, technical approaches, evaluation practices, and limitations. Our findings reveal significant concentration patterns and gaps. UML-based software modelling strongly dominates, with class diagrams receiving the most attention whilst behavioural diagrams and data modelling remain underrepresented. Diagram construction from natural language is the primary focus, with limited work on transformation, quality assurance, and consistency checking. GPT-based models are heavily prevalent, raising concerns about reproducibility and vendor dependence. Evaluation practices are heterogeneous, employing diverse metrics and custom datasets with limited benchmark reuse and inconsistent reporting of robustness and statistical significance. Common limitations include semantic inaccuracies, hallucinated diagram elements, sensitivity to prompt formulation, and reproducibility constraints. This survey provides the first systematic synthesis of LLM-based diagram modelling research, highlighting needs for standardised benchmarks, stronger evaluation protocols, broader diagram coverage, and techniques for improving semantic reliability and multi-view consistency.

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

124 extracted references · 8 canonical work pages

  1. [61]

    Information16(7), 565 (2025)

    Al-Ahmad, B., Alsobeh, A., Meqdadi, O., Shaikh, N.: A student-centric evalua- tion survey to explore the impact of llms on uml modeling. Information16(7), 565 (2025)

  2. [79]

    Information 16(7), 565 (2025) https://doi.org/10.3390/info16070565

    Al-Ahmad, B., Alsobeh, A., Meqdadi, O., Shaikh, N.: A Student-Centric Eval- uation Survey to Explore the Impact of LLMs on UML Modeling. Information 16(7), 565 (2025) https://doi.org/10.3390/info16070565 . Accessed 2026-01-01

  3. [48]

    arxiv 2024

    Wang, B., Wang, C., Liang, P., Li, B., Zeng, C.: How llms aid in uml mod- eling: An exploratory study with novice analysts. arxiv 2024. arXiv preprint arXiv:2404.17739

  4. [42]

    In: Proceedings of the ACM/IEEE 27th International Conference on Model Driven Engineering Languages and Systems, pp

    Sultan, B., Apvrille, L.: Ai-driven consistency of sysml diagrams. In: Proceedings of the ACM/IEEE 27th International Conference on Model Driven Engineering Languages and Systems, pp. 149–159 (2024)

  5. [97]

    Information and Software Technology190, 107955 (2026) https://doi.org/10.1016/j.infsof

    Liu, X., Liu, Y., Zhuang, Y., Hou, W.: UCD-LLM: A use case diagram require- ment modeling multi-agent framework with large language model. Information and Software Technology190, 107955 (2026) https://doi.org/10.1016/j.infsof. 2025.107955 . Accessed 2026-02-09

  6. [25]

    arXiv preprint arXiv:2312.15223 (2023) 34

    Zhang, Q., Fang, C., Xie, Y., Zhang, Y., Yang, Y., Sun, W., Yu, S., Chen, Z.: A survey on large language models for software engineering. arXiv preprint arXiv:2312.15223 (2023) 34

  7. [1]

    Addison-Wesley, Upper Saddle River, NJ (2005)

    Booch, G., Rumbaugh, J., Jacobson, I.: The Unified Modeling Language User Guide, 2nd ed edn. Addison-Wesley, Upper Saddle River, NJ (2005)

  8. [2]

    Computer 39(02), 25–31 (2006)

    Schmidt, D.C.: Guest editor’s introduction: Model-driven engineering. Computer 39(02), 25–31 (2006). Publisher: IEEE Computer Society

Show all 124 references
  1. [3]

    OMG http://www

    SysML, O.: Systems modeling language. OMG http://www. sysmlomg. org (2006) 32

  2. [4]

    Addison-Wesley [u.a.], Harlow (2006)

    Jackson, M.A.: Problem Frames: Analysing and Structuring Software Develop- ment Problems, Transferred to digital print edn. Addison-Wesley [u.a.], Harlow (2006)

  3. [5]

    America: Pearson Education Inc (2011)

    Sommerville, I.: Software engineering (ed.). America: Pearson Education Inc (2011)

  4. [6]

    In: 25th International Con- ference on Software Engineering, 2003

    Clements, P., Garlan, D., Little, R., Nord, R., Stafford, J.: Document- ing software architectures: views and beyond. In: 25th International Con- ference on Software Engineering, 2003. Proceedings., pp. 740–741. IEEE, Portland, OR, USA (2003). https://doi.org/10.1109/ICSE.2003...

  5. [7]

    Addison-Wesley Object Technology Series

    Fowler, M.: UML Distilled: A Brief Guide to the Standard Object Modeling Lan- guage. Addison-Wesley Object Technology Series. Pearson Education, Boston, MA (2018).https://books.google.co.uk/books?id=VTdtDwAAQBAJ

  6. [8]

    print edn

    Rumbaugh, J., Jacobson, I., Booch, G.: The Unified Modeling Language Ref- erence Manual, 2. print edn. The Addison-Wesley object technology series. Addison-Wesley, Boston (2006)

  7. [9]

    In: 2013 35th International Conference on Software Engineering (ICSE), pp

    Petre, M.: UML in practice. In: 2013 35th International Conference on Software Engineering (ICSE), pp. 722–731. IEEE, San Fran- cisco, CA, USA (2013). https://doi.org/10.1109/ICSE.2013.6606618 . http://ieeexplore.ieee.org/document/6606618/Accessed 2026-03-09

  8. [10]

    In: Proceedings of the 33rd International Conference on Software Engineering, pp

    Hutchinson, J., Whittle, J., Rouncefield, M., Kristoffersen, S.: Empirical assess- ment of MDE in industry. In: Proceedings of the 33rd International Conference on Software Engineering, pp. 471–480 (2011)

  9. [11]

    arXiv preprint arXiv:2107.03374 (2021)

    Chen, M.: Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374 (2021)

  10. [12]

    Empirical Software Engineering30(3), 65 (2025)

    Tambon, F., Moradi-Dakhel, A., Nikanjam, A., Khomh, F., Desmarais, M.C., Antoniol, G.: Bugs in large language models generated code: An empirical study. Empirical Software Engineering30(3), 65 (2025). Publisher: Springer

  11. [13]

    In: 2023 IEEE/ACM 45th Inter- national Conference on Software Engineering (ICSE), pp

    Mastropaolo, A., Pascarella, L., Guglielmi, E., Ciniselli, M., Scalabrino, S., Oliveto, R., Bavota, G.: On the Robustness of Code Generation Techniques: An Empirical Study on GitHub Copilot. In: 2023 IEEE/ACM 45th Inter- national Conference on Software Engineering (ICSE), pp. ...

  12. [14]

    The Addison-Wesley object technology series

    Kleppe, A.G., Warmer, J.B., Bast, W.: MDA Explained: the Model Driven Architecture: Practice and Promise. The Addison-Wesley object technology series. Addison-Wesley, Boston (2003) 33

  13. [15]

    Software & Systems Modeling11(4), 513–526 (2012)

    Selic, B.: What will it take? A view on adoption of model-based methods in practice. Software & Systems Modeling11(4), 513–526 (2012). Publisher: Springer

  14. [16]

    329–380 (2001)

    Spanoudakis, G., Zisman, A.: INCONSISTENCY MANAGEMENT IN SOFT- W ARE ENGINEERING: SUR VEY AND OPEN RESEARCH ISSUES, pp. 329–380 (2001). https://doi.org/10.1142/9789812389718 0015 . http://www. worldscientific.com/doi/abs/10.1142/9789812389718 0015 Accessed 2026-03-22

  15. [17]

    Computer33(4), 24–29 (2002)

    Nuseibeh, B., Easterbrook, S., Russo, A.: Leveraging inconsistency in software development. Computer33(4), 24–29 (2002). Publisher: IEEE

  16. [18]

    Software Quality Journal22(1), 121–149 (2014)

    Landh¨ außer, M., K¨ orner, S.J., Tichy, W.F.: From requirements to UML models and back: how automatic processing of text can support requirements engineer- ing. Software Quality Journal22(1), 121–149 (2014). Publisher: Springer

  17. [19]

    Ilieva, M.G., Ormandjieva, O.: Automatic Transition of Natural Language Soft- ware Requirements Specification into Formal Presentation. In: Hutchison, D., Kanade, T., Kittler, J., Kleinberg, J.M., Mattern, F., Mitchell, J.C., Naor, M., Nierstrasz, O., Pandu Rangan, C., Steffen...

  18. [20]

    In: CS&P, vol

    Gr¨ opler, R., Sudhi, V., Garc ´ ıa, E.J.C., Bergmann, A.: NLP-Based Requirements Formalization for Automatic Test Case Generation. In: CS&P, vol. 21, pp. 18–30 (2021)

  19. [21]

    IEEE Transactions on Software Engineering50(9), 2269–2293 (2024)

    Marc´ en, A.C., Iglesias, A., Lape˜ na, R., P´ erez, F., Cetina, C.: A systematic literature review of model-driven engineering using machine learning. IEEE Transactions on Software Engineering50(9), 2269–2293 (2024)

  20. [22]

    ACM Transactions on Database Systems1(1), 9–36 (1976)

    Chen, P.P.-S.: The entity-relationship model—toward a unified view of data. ACM Transactions on Database Systems1(1), 9–36 (1976)

  21. [23]

    ISBN-10137035152, 18 (2011)

    Sommerville, I.: Software engineering 9th edition. ISBN-10137035152, 18 (2011)

  22. [24]

    ACM Transactions on Software Engineering and Methodology 33(8), 1–79 (2024)

    Hou, X., Zhao, Y., Liu, Y., Yang, Z., Wang, K., Li, L., Luo, X., Lo, D., Grundy, J., Wang, H.: Large language models for software engineering: A systematic literature review. ACM Transactions on Software Engineering and Methodology 33(8), 1–79 (2024)

  23. [26]

    In: 2023 IEEE/ACM International Conference on Software Engineering: Future of Software Engineering (ICSE-FoSE), pp

    Fan, A., Gokkaya, B., Harman, M., Lyubarskiy, M., Sengupta, S., Yoo, S., Zhang, J.M.: Large language models for software engineering: Survey and open prob- lems. In: 2023 IEEE/ACM International Conference on Software Engineering: Future of Software Engineering (ICSE-FoSE), pp....

  24. [27]

    arXiv preprint arXiv:2409.02977 (2024)

    Liu, J., Wang, K., Chen, Y., Peng, X., Chen, Z., Zhang, L., Lou, Y.: Large language model-based agents for software engineering: A survey. arXiv preprint arXiv:2409.02977 (2024)

  25. [28]

    Information and Software Technology169, 107423 (2024)

    Naveed, H., Arora, C., Khalajzadeh, H., Grundy, J., Haggag, O.: Model driven engineering for machine learning components: A systematic literature review. Information and Software Technology169, 107423 (2024)

  26. [29]

    Software and Systems Modeling 24(2), 445–469 (2025)

    R¨ adler, S., Berardinelli, L., Winter, K., Rahimi, A., Rinderle-Ma, S.: Bridging mde and ai: a systematic review of domain-specific languages and model-driven practices in ai software systems engineering. Software and Systems Modeling 24(2), 445–469 (2025)

  27. [30]

    di rocco et al

    Di Rocco, J., Di Ruscio, D., Di Sipio, C., Nguyen, P.T., Rubei, R.: On the use of large language models in model-driven engineering: J. di rocco et al. Software and Systems Modeling24(3), 923–948 (2025)

  28. [31]

    ACM Transactions on Software Engineering and Methodology34(5), 1–25 (2025)

    Burgue˜ no, L., Di Ruscio, D., Sahraoui, H., Wimmer, M.: Automation in model- driven engineering: A look back, and ahead. ACM Transactions on Software Engineering and Methodology34(5), 1–25 (2025)

  29. [32]

    Frontiers in Computer Science7, 1519437 (2025)

    Hemmat, A., Sharbaf, M., Kolahdouz-Rahimi, S., Lano, K., Tehrani, S.Y.: Research directions for using llm in software requirement engineering: A systematic review. Frontiers in Computer Science7, 1519437 (2025)

  30. [33]

    arXiv preprint arXiv:2509.11446 (2025)

    Zadenoori, M.A., Dabrowski, J., Alhoshan, W., Zhao, L., Ferrari, A.: Large language models (llms) for requirements engineering (re): A systematic literature review. arXiv preprint arXiv:2509.11446 (2025)

  31. [34]

    Software: Practice and Experience56(2), 141–170 (2026)

    Cheng, H., Husen, J.H., Lu, Y., Racharak, T., Yoshioka, N., Ubayashi, N., Washizaki, H.: Generative ai for requirements engineering: A systematic litera- ture review. Software: Practice and Experience56(2), 141–170 (2026)

  32. [35]

    applica- tions, challenges, and future directions

    Esposito, M., Li, X., Moreschini, S., Ahmad, N., Cerny, T., Vaidhyanathan, K., Lenarduzzi, V., Taibi, D.: Generative ai for software architecture. applica- tions, challenges, and future directions. Journal of Systems and Software, 112607 (2025)

  33. [36]

    arXiv preprint arXiv:2505.16697 (2025) 35

    Schmid, L., Hey, T., Armbruster, M., Corallo, S., Fuchß, D., Keim, J., Liu, H., Koziolek, A.: Software architecture meets llms: A systematic literature review. arXiv preprint arXiv:2505.16697 (2025) 35

  34. [37]

    Technical report, Technical report, ver

    Keele, S., et al.: Guidelines for performing systematic literature reviews in soft- ware engineering. Technical report, Technical report, ver. 2.3 ebse technical report. ebse (2007)

  35. [38]

    ACM Transactions on Software Engineering and Methodology33(5), 1–59 (2024)

    Chen, Z., Zhang, J.M., Hort, M., Harman, M., Sarro, F.: Fairness testing: A comprehensive survey and analysis of trends. ACM Transactions on Software Engineering and Methodology33(5), 1–59 (2024). Publisher: ACM New York, NY

  36. [39]

    In: International Conference on Advanced Information Systems Engineering, pp

    Reinhartz-Berger, I., Ali, S.J., Bork, D.: Leveraging llms for domain modeling: The impact of granularity and strategy on quality. In: International Conference on Advanced Information Systems Engineering, pp. 3–19 (2025). Springer

  37. [40]

    Procedia Computer Science246, 1346–1354 (2024)

    Naimi, L., Jakimi, A., Saadane, R., Chehri, A.,et al.: Automating software doc- umentation: Employing llms for precise use case description. Procedia Computer Science246, 1346–1354 (2024)

  38. [41]

    In: Proceedings of the ACM/IEEE 27th International Conference on Model Driven Engineering Languages and Systems, pp

    Tabassum, M.R., Ritchie, M.J., Mustafiz, S., Kienzle, J.: Using llms for use case modelling of iot systems: An experience report. In: Proceedings of the ACM/IEEE 27th International Conference on Model Driven Engineering Languages and Systems, pp. 611–619 (2024)

  39. [43]

    In: 37th International BCS Human-Computer Interaction Conference, pp

    O’Neill, I.: Getting gpt to answer like me. In: 37th International BCS Human-Computer Interaction Conference, pp. 7–13 (2024). BCS Learning & Development

  40. [44]

    In: Proceedings of the 33rd ACM International Conference on the Foundations of Software Engineering, pp

    Klimek, R.: Re-oriented model development with llm support and deduction- based verification. In: Proceedings of the 33rd ACM International Conference on the Foundations of Software Engineering, pp. 1297–1304 (2025)

  41. [45]

    In: SoutheastCon 2025, pp

    Ramachandran, R.: Transforming software architecture design with intelligent assistants-a comparative analysis. In: SoutheastCon 2025, pp. 1446–1454 (2025). IEEE

  42. [46]

    In: 2025 25th International Conference on Software Quality, Reliability and Security (QRS), pp

    Lu, J., Sun, P., Chen, Y., Yin, G., Ye, P.: Aug: an interactive tool for clarify- ing and generating uml models based on large language models. In: 2025 25th International Conference on Software Quality, Reliability and Security (QRS), pp. 78–85 (2025). IEEE

  43. [47]

    In: 2025 IEEE International Confer- ence on Software Analysis, Evolution and Reengineering (SANER), pp

    Hassine, J.: Evaluating multi-modal llms for automatically recognizing semantic elements in uml use case diagram images. In: 2025 IEEE International Confer- ence on Software Analysis, Evolution and Reengineering (SANER), pp. 861–866 (2025). IEEE 36

  44. [49]

    Machine Learning with Applications, 100660 (2025)

    Bates, A., Vavricka, R., Carleton, S., Shao, R., Pan, C.: Unified modeling lan- guage code generation from diagram images using multimodal large language models. Machine Learning with Applications, 100660 (2025)

  45. [50]

    In: Proceedings of the 2024 IEEE/ACM 46th International Conference on Software Engineering: Companion Proceedings, pp

    Siala, H.A.: Enhancing model-driven reverse engineering using machine learn- ing. In: Proceedings of the 2024 IEEE/ACM 46th International Conference on Software Engineering: Companion Proceedings, pp. 173–175 (2024)

  46. [51]

    In: Proceedings of the 30th ACM Conference on Innovation and Technology in Computer Science Education V

    Ouh, E.L., Tan, K.W., Lo, S.L., Gan, B.K.S.: Evaluating chatgpt to answer multi-modal exercises in computer science education. In: Proceedings of the 30th ACM Conference on Innovation and Technology in Computer Science Education V. 1, pp. 58–64 (2025)

  47. [52]

    a rule-based approach

    Jahan, M., Hassan, M.M., Golpayegani, R., Ranjbaran, G., Roy, C., Roy, B., Schneider, K.: Automated derivation of uml sequence diagrams from user stories: Unleashing the power of generative ai vs. a rule-based approach. In: Proceedings of the ACM/IEEE 27th International Confer...

  48. [53]

    In: 2024 36th International Con- ference on Software Engineering Education and Training (CSEE&T), pp

    Speth, S., Meißner, N., Becker, S.: Chatgpt’s aptitude in utilizing uml diagrams for software engineering exercise generation. In: 2024 36th International Con- ference on Software Engineering Education and Training (CSEE&T), pp. 1–5 (2024). IEEE

  49. [54]

    In: 2025 IEEE/ACM 47th International Conference on Software Engineering: Software Engineering in Practice (ICSE-SEIP), pp

    Xiao, C., St ˚ ahl, D., Bosch, J.: Uml sequence diagram generation: A multi-model, multi-domain evaluation. In: 2025 IEEE/ACM 47th International Conference on Software Engineering: Software Engineering in Practice (ICSE-SEIP), pp. 272–283 (2025). IEEE

  50. [55]

    In: 2024 IEEE 32nd International Requirements Engineering Conference Workshops (REW), pp

    Ferrari, A., Abualhaija, S., Arora, C.: Model generation with llms: From requirements to uml sequence diagrams. In: 2024 IEEE 32nd International Requirements Engineering Conference Workshops (REW), pp. 291–300 (2024). IEEE

  51. [56]

    Theoretical Computer Science1021, 114879 (2024)

    Zhao, Z., Zhang, N., Yu, B., Duan, Z.: Generating java code pairing with chatgpt. Theoretical Computer Science1021, 114879 (2024)

  52. [57]

    In: 2023 IEEE/ACM 45th Interna- tional Conference on Software Engineering: New Ideas and Emerging Results (ICSE-NIER), pp

    Chaaben, M.B., Burgue˜ no, L., Sahraoui, H.: Towards using few-shot prompt learning for automating model completion. In: 2023 IEEE/ACM 45th Interna- tional Conference on Software Engineering: New Ideas and Emerging Results (ICSE-NIER), pp. 7–12 (2023). IEEE

  53. [58]

    In: 37 Proceedings of the ACM/IEEE 27th International Conference on Model Driven Engineering Languages and Systems, pp

    Ben Chaaben, M.: Software modeling assistance with large language models. In: 37 Proceedings of the ACM/IEEE 27th International Conference on Model Driven Engineering Languages and Systems, pp. 188–191 (2024)

  54. [59]

    In: 2025 IEEE 22nd International Conference on Software Architecture Companion (ICSA-C), pp

    Moezkarimi, Z., Eriksson, K., Johansson, A.A., Bucaioni, A., Sirjani, M.: Har- nessing chatgpt for model transformation in software architecture: From uml state diagrams to rebeca models for formal verification. In: 2025 IEEE 22nd International Conference on Software Architect...

  55. [60]

    Software and Systems Modeling22(3), 781–793 (2023)

    C´ amara, J., Troya, J., Burgue˜ no, L., Vallecillo, A.: On the assessment of gener- ative ai in modeling tasks: an experience report with chatgpt and uml. Software and Systems Modeling22(3), 781–793 (2023)

  56. [62]

    Computer Applications in Engineering Education33(5), 70080 (2025)

    Ib´ a˜ nez, M.B., Barr´ on-Estrada, M.L., Zatarain-Cabada, R.: Can multimodal large language models grade like an expert? a study on uml class diagram assess- ment accuracy. Computer Applications in Engineering Education33(5), 70080 (2025)

  57. [63]

    In: Proceedings of the ACM/IEEE 27th International Conference on Model Driven Engineering Languages and Systems, pp

    Yang, Y., Chen, B., Chen, K., Mussbacher, G., Varr´ o, D.: Multi-step iterative automated domain modeling with large language models. In: Proceedings of the ACM/IEEE 27th International Conference on Model Driven Engineering Languages and Systems, pp. 587–595 (2024)

  58. [64]

    In: Proceedings of the 33rd ACM International Conference on the Foundations of Software Engineering, pp

    Wang, T., Trimble, M., Brown, C.: Devcoach: Supporting students learning the software development life cycle with a generative ai powered multi-agent system. In: Proceedings of the 33rd ACM International Conference on the Foundations of Software Engineering, pp. 987–998 (2025)

  59. [65]

    In: Proceed- ings of the 18th ACM/IEEE International Symposium on Empirical Software Engineering and Measurement, pp

    De Bari, D., Garaccione, G., Coppola, R., Torchiano, M., Ardito, L.: Evaluating large language models in exercises of uml class diagram modeling. In: Proceed- ings of the 18th ACM/IEEE International Symposium on Empirical Software Engineering and Measurement, pp. 393–399 (2024)

  60. [66]

    In: Proceedings of the ACM/IEEE 27th International Conference on Model Driven Engineering Lan- guages and Systems, pp

    Ardimento, P., Bernardi, M.L., Cimitile, M., Scalera, M.: Enhancing soft- ware modeling learning with ai-powered scaffolding. In: Proceedings of the ACM/IEEE 27th International Conference on Model Driven Engineering Lan- guages and Systems, pp. 103–106 (2024)

  61. [67]

    In: Proceedings of the ACM/IEEE 27th International Conference on Model Driven Engineering Languages and Systems, pp

    Ardimento, P., Bernardi, M.L., Cimitile, M., Scalera, M.: A rag-based feedback tool to augment uml class diagram learning. In: Proceedings of the ACM/IEEE 27th International Conference on Model Driven Engineering Languages and Systems, pp. 26–30 (2024) 38

  62. [68]

    In: Proceedings of the 1st International Workshop on Large Language Models for Code, pp

    Antal, G., Voz´ ar, R., Ferenc, R.: Toward a new era of rapid development: Assess- ing gpt-4-vision’s capabilities in uml-based code generation. In: Proceedings of the 1st International Workshop on Large Language Models for Code, pp. 84–87 (2024)

  63. [69]

    In: Proceedings of the 27th International Conference on Evaluation and Assessment in Software Engineering, pp

    Ahmad, A., Waseem, M., Liang, P., Fahmideh, M., Aktar, M.S., Mikkonen, T.: Towards human-bot collaborative software architecting with chatgpt. In: Proceedings of the 27th International Conference on Evaluation and Assessment in Software Engineering, pp. 279–285 (2023)

  64. [70]

    In: Proceedings of the 27th ACM Inter- national Systems and Software Product Line Conference-Volume B, pp

    Acher, M., Martinez, J.: Generative ai for reengineering variants into software product lines: an experience report. In: Proceedings of the 27th ACM Inter- national Systems and Software Product Line Conference-Volume B, pp. 57–66 (2023)

  65. [71]

    In: Proceedings of the 2024 IEEE/ACM First International Conference on AI Foundation Models and Software Engineer- ing, pp

    Abukhalaf, S., Hamdaqa, M., Khomh, F.: Pathocl: path-based prompt augmen- tation for ocl generation with gpt-4. In: Proceedings of the 2024 IEEE/ACM First International Conference on AI Foundation Models and Software Engineer- ing, pp. 108–118 (2024)

  66. [72]

    IEEE Access (2025)

    Babaalla, Z., Jakimi, A., Oualla, M.: Llm-driven mda pipeline for generating uml class diagrams and code. IEEE Access (2025)

  67. [73]

    In: Proceedings of the 46th International Conference on Software Engineering: Software Engineering Education and Training, pp

    Xue, Y., Chen, H., Bai, G.R., Tairas, R., Huang, Y.: Does chatgpt help with introductory programming? an experiment of students using chatgpt in cs1. In: Proceedings of the 46th International Conference on Software Engineering: Software Engineering Education and Training, pp. ...

  68. [74]

    In: 2024 IEEE 48th Annual Computers, Software, and Applications Conference (COMPSAC), pp

    Li, Y., Keung, J., Ma, X., Chong, C.Y., Zhang, J., Liao, Y.: Llm-based class diagram derivation from user stories with chain-of-thought promptings. In: 2024 IEEE 48th Annual Computers, Software, and Applications Conference (COMPSAC), pp. 45–50 (2024). IEEE

  69. [75]

    In: 2023 IEEE/ACM 20th International Conference on Mining Software Repositories (MSR), pp

    Abukhalaf, S., Hamdaqa, M., Khomh, F.: On codex prompt engineering for ocl generation: an empirical study. In: 2023 IEEE/ACM 20th International Conference on Mining Software Repositories (MSR), pp. 148–157 (2023). IEEE

  70. [76]

    In: 2024 International Joint Conference on Neural Networks (IJCNN), pp

    Ardimento, P., Bernardi, M.L., Cimitile, M.: Teaching uml using a rag-based llm. In: 2024 International Joint Conference on Neural Networks (IJCNN), pp. 1–8 (2024). IEEE

  71. [77]

    In: 2024 IEEE Working Conference on Software Visualization (VISSOFT), pp

    Shehata, M., Lepore, B., Cummings, H., Parra, E.: Creating uml class dia- grams with general-purpose llms. In: 2024 IEEE Working Conference on Software Visualization (VISSOFT), pp. 157–158 (2024). IEEE

  72. [78]

    In: 2025 8th International Conference on Software and 39 System Engineering (ICoSSE), pp

    Siala, H.A., Lano, K.: Using large language models to extract uml class diagrams from java programs. In: 2025 8th International Conference on Software and 39 System Engineering (ICoSSE), pp. 70–74 (2025). IEEE

  73. [80]

    In: 2025 International Conference on Next Generation Information System Engineering (NGISE), pp

    Ojha, M., Gupta, S., Sharma, R.: Towards Class Diagram Generation from User Stories Using LLMs. In: 2025 International Conference on Next Generation Information System Engineering (NGISE), pp. 1–6. IEEE, Ghaziabad, Delhi (NCR), India (2025). https://doi.org/10.1109/NGISE64126....

  74. [81]

    IEEE Access13, 211605–211619 (2025) https://doi.org/10.1109/ACCESS.2025

    Nair, R.P., Thushara, M.G., Sugumaran, V.: Fine-Tuned LLMs Versus Rule- Based NLP for UML Diagram Generation: An Educational Evaluation. IEEE Access13, 211605–211619 (2025) https://doi.org/10.1109/ACCESS.2025. 3638372 . Accessed 2026-01-01

  75. [82]

    In: 2025 IEEE/ACM International Workshop on Nat- ural Language-Based Software Engineering (NLBSE), pp

    Garaccione, G., Vega Carrazan, P.F., Coppola, R., Ardito, L.: Eval- uating Large Language Models in Exercises of UML Use Case Dia- grams Modeling. In: 2025 IEEE/ACM International Workshop on Nat- ural Language-Based Software Engineering (NLBSE), pp. 41–44. IEEE, Ottawa, ON, Ca...

  76. [83]

    In: Proceedings of the 13th International Conference on Model-Based Software and Systems Engineering, pp

    Kop, C.: Evaluating the Quality of Class Diagrams Created by a Genera- tive AI: Findings, Guidelines and Automation Options:. In: Proceedings of the 13th International Conference on Model-Based Software and Systems Engineering, pp. 150–157. SCITEPRESS - Science and Technology ...

  77. [84]

    In: Maass, W., Han, H., Yasar, H., Multari, N

    Silva, J., Ma, Q., Cabot, J., Kelsen, P., Proper, H.A.: Application of the Tree-of- Thoughts Framework to LLM-Enabled Domain Modeling. In: Maass, W., Han, H., Yasar, H., Multari, N. (eds.) Conceptual Modeling vol. 15238, pp. 94–111. Springer, Cham (2025). https://doi.org/10.10...

  78. [85]

    In: Proceedings of the 17th International Conference on Computer Supported Education, pp

    Bouali, N., Gerhold, M., Rehman, T., Ahmed, F.: Toward Automated UML Dia- gram Assessment: Comparing LLM-Generated Scores with Teaching Assistants:. In: Proceedings of the 17th International Conference on Computer Supported Education, pp. 158–169. SCITEPRESS - Science and Tech...

  79. [86]

    Enterprise Mod- elling and Information Systems Architectures (EMISAJ), 3–115 (2023) https: //doi.org/10.18417/EMISA.18.3

    Fill, H.-G., Fettke, P., K¨ opke, J.: Conceptual Modeling and Large Language 40 Models: Impressions From First Experiments With ChatGPT. Enterprise Mod- elling and Information Systems Architectures (EMISAJ), 3–115 (2023) https: //doi.org/10.18417/EMISA.18.3 . Artwork Size: 3:1...

  80. [87]

    Journal of Systems and Software, 112709 (2025)

    Wang, C., Wang, B., Liang, P., Liang, J.: Assessing uml diagrams by gpt: Implications for education. Journal of Systems and Software, 112709 (2025)

  81. [88]

    In 2023 ACM/IEEE 26th International Conference on Model Driven Engineering Languages and Systems (MODELS)

    Chen, K., Yang, Y., Chen, B., L´ opez, J.A.H., Mussbacher, G., Varr´ o, D.: Auto- mated Domain Modeling with Large Language Models: A Comparative Study. In 2023 ACM/IEEE 26th International Conference on Model Driven Engineering Languages and Systems (MODELS). 162–172 (2023)

  82. [89]

    In: International Conference on Future Data and Security Engineering, pp

    Nguyen, V.-V., Nguyen, H.-K., Nguyen, K.-S., Luong Thi, M.-H., Nguyen, T.-V., Vu, D.-Q.: Automated uml generation: A framework for class diagram synthesis and multimodal validation. In: International Conference on Future Data and Security Engineering, pp. 212–224 (2025). Springer

  83. [90]

    In: Proceedings of the IEEE/ACM 47th International Conference on Software Engineering

    Fuchß, D., Hey, T., Keim, J., Liu, H., Ewald, N., Thirolf, T., Koziolek, A.: Lissa: toward generic traceability link recovery through retrieval-augmented gen- eration. In: Proceedings of the IEEE/ACM 47th International Conference on Software Engineering. ICSE, vol. 25 (2025)

  84. [91]

    In: 2025 IEEE 22nd Interna- tional Conference on Software Architecture Companion (ICSA-C), pp

    Tagliaferro, A., Corboe, S., Guindani, B.: Leveraging llms to automate software architecture design from informal specifications. In: 2025 IEEE 22nd Interna- tional Conference on Software Architecture Companion (ICSA-C), pp. 291–299 (2025). IEEE

  85. [92]

    In: 2023 IEEE 11th International Conference on Systems and Control (ICSC), pp

    Omar, M.A.: Measurement of chatgpt performance in mapping natural lan- guage speficaction into an entity relationship diagram. In: 2023 IEEE 11th International Conference on Systems and Control (ICSC), pp. 530–535 (2023). IEEE

  86. [93]

    Cogent Education12(1), 2590901 (2025) https://doi

    Rahmanian, M., Sami, A., Yu, Y.: Challenges and feasibility of multimodal LLMs in ER diagram evaluation. Cogent Education12(1), 2590901 (2025) https://doi. org/10.1080/2331186X.2025.2590901 . Accessed 2026-01-02

  87. [94]

    In: Proceedings of the 16th ACM International Conference on Future and Sustainable Energy Systems, pp

    Kuchenbuch, R., Lehnhoff, S., Sauer, J.: Smart grid assistive ai in requirement engineering: Improving the modeling of use cases and architecture models with llms. In: Proceedings of the 16th ACM International Conference on Future and Sustainable Energy Systems, pp. 495–504 (2025)

  88. [95]

    In: Proceedings of the 3rd International Conference on Futuristic Technology, pp

    G S, N.K., S, A., Thushara, M.G.: Comparative Analysis of Large Lan- guage Models for Automated Use Case Diagram Generation:. In: Proceedings of the 3rd International Conference on Futuristic Technology, pp. 465–

  89. [96]

    In: Guizzardi, R., Pufahl, L., Sturm, A., Van Der Aa, H

    Calamo, M., Mecella, M., Snoeck, M.: Assessing the Suitability of Large Lan- guage Models in Generating UML Class Diagrams as Conceptual Models. In: Guizzardi, R., Pufahl, L., Sturm, A., Van Der Aa, H. (eds.) Enterprise, Business-Process and Information Systems Modeling vol. 5...

  90. [98]

    In: 2025 40th IEEE/ACM International Conference on Automated Software Engineering Workshops (ASEW), pp

    Eisenreich, T., Friedlaender, N., Wagner, S.: Leveraging large language mod- els for use case model generation from software requirements. In: 2025 40th IEEE/ACM International Conference on Automated Software Engineering Workshops (ASEW), pp. 221–227 (2025). IEEE

  91. [99]

    In: International Conference on Conceptual Modeling, pp

    Gavric, A., Bork, D., Proper, H.A.: How does uml look and sound? using ai to interpret uml diagrams through multimodal evidence. In: International Conference on Conceptual Modeling, pp. 187–197 (2024). Springer

  92. [100]

    Big Data and Cognitive Computing 10(1), 2 (2025) https://doi.org/10.3390/bdcc10010002

    Ramachandran, R., Vijayan, P., Anilkumar, A., Gangadharan, V.: AI Assisted System for Automated Evaluation of Entity-Relationship Diagram and Schema Diagram Using Large Language Models. Big Data and Cognitive Computing 10(1), 2 (2025) https://doi.org/10.3390/bdcc10010002 . Acc...

  93. [101]

    Information16(5), 368 (2025) https://doi.org/10.3390/info16050368

    Avignone, A., Tierno, A., Fiori, A., Chiusano, S.: Exploring Large Language Models’ Ability to Describe Entity-Relationship Schema-Based Conceptual Data Models. Information16(5), 368 (2025) https://doi.org/10.3390/info16050368 . Accessed 2026-01-01

  94. [102]

    CAAI Transactions on Intelligence Technology9(1), 250–263 (2024)

    Yang, Y., Liu, Y., Bao, T., Wang, W., Niu, N., Yin, Y.: Deepocl: A deep neu- ral network for object constraint language generation from unrestricted nature language. CAAI Transactions on Intelligence Technology9(1), 250–263 (2024)

  95. [103]

    Pearson Higher Education, ??? (2004)

    Rumbaugh, J., Jacobson, I., Booch, G.: Unified Modeling Language Reference Manual, The (2nd Edition). Pearson Higher Education, ??? (2004)

  96. [104]

    Advances in Neural Information Processing Systems37, 50528–50652 (2024) 42

    Yang, J., Jimenez, C.E., Wettig, A., Lieret, K., Yao, S., Narasimhan, K., Press, O.: Swe-agent: Agent-computer interfaces enable automated software engineer- ing. Advances in Neural Information Processing Systems37, 50528–50652 (2024) 42

  97. [105]

    arXiv preprint arXiv:2207.10397 (2022)

    Chen, B., Zhang, F., Nguyen, A., Zan, D., Lin, Z., Lou, J.-G., Chen, W.: Codet: Code generation with generated tests. arXiv preprint arXiv:2207.10397 (2022)

  98. [106]

    experience: Evalu- ating the usability of code generation tools powered by large language models

    Vaithilingam, P., Zhang, T., Glassman, E.L.: Expectation vs. experience: Evalu- ating the usability of code generation tools powered by large language models. In: Chi Conference on Human Factors in Computing Systems Extended Abstracts, pp. 1–7 (2022)

  99. [107]

    arXiv preprint arXiv:2108.07732 (2021)

    Austin, J., Odena, A., Nye, M., Bosma, M., Michalewski, H., Dohan, D., Jiang, E., Cai, C., Terry, M., Le, Q., et al.: Program synthesis with large language models. arXiv preprint arXiv:2108.07732 (2021)

  100. [108]

    In: Future of Software Engineering (FOSE’07), pp

    France, R., Rumpe, B.: Model-driven development of complex software: A research roadmap. In: Future of Software Engineering (FOSE’07), pp. 37–54 (2007). IEEE

  101. [109]

    In: Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pp

    Rajpurkar, P., Zhang, J., Lopyrev, K., Liang, P.: Squad: 100,000+ questions for machine comprehension of text. In: Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pp. 2383–2392 (2016)

  102. [110]

    In: Proceedings of the 2018 EMNLP Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP, pp

    Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., Bowman, S.: Glue: A multi-task benchmark and analysis platform for natural language understand- ing. In: Proceedings of the 2018 EMNLP Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP, pp. 353–355 (2018)

  103. [111]

    Springer (2003)

    Handbook, A.: From contract drafting to software specification: Linguistic sources of ambiguity. Springer (2003)

  104. [112]

    ACM Transactions on Software Engineering and Methodology (TOSEM)27(3), 1–51 (2018)

    Stol, K.-J., Fitzgerald, B.: The abc of software engineering research. ACM Transactions on Software Engineering and Methodology (TOSEM)27(3), 1–51 (2018)

  105. [113]

    In: Findings of the Association for Computational Linguistics ACL 2024, pp

    Dhuliawala, S., Komeili, M., Xu, J., Raileanu, R., Li, X., Celikyilmaz, A., Weston, J.: Chain-of-Verification Reduces Hallucination in Large Language Models. In: Findings of the Association for Computational Linguistics ACL 2024, pp. 3563–3578. Association for Computational Li...

  106. [114]

    ACM Computing Surveys55(12), 1–38 (2023) https://doi.org/10.1145/3571730

    Ji, Z., Lee, N., Frieske, R., Yu, T., Su, D., Xu, Y., Ishii, E., Bang, Y.J., Madotto, A., Fung, P.: Survey of Hallucination in Natural Language Generation. ACM Computing Surveys55(12), 1–38 (2023) https://doi.org/10.1145/3571730 . Accessed 2026-03-05

  107. [115]

    Tonmoy, S.M.T.I., Zaman, S.M.M., Jain, V., Rani, A., Rawte, V., Chadha, A., Das, A.: A Comprehensive Survey of Hallucination Mitigation Techniques in Large Language Models. arXiv. Version Number: 3 (2024). https://doi. 43 org/10.48550/ARXIV.2401.01313 . https://arxiv.org/abs/2...

  108. [116]

    Iclr1(2), 3 (2022)

    Hu, E.J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., Chen, W.,et al.: Lora: Low-rank adaptation of large language models. Iclr1(2), 3 (2022)

  109. [117]

    Advances in neural information processing systems36, 10088–10115 (2023)

    Dettmers, T., Pagnoni, A., Holtzman, A., Zettlemoyer, L.: Qlora: Efficient fine- tuning of quantized llms. Advances in neural information processing systems36, 10088–10115 (2023)

  110. [118]

    Rozi` ere, B., Gehring, J., Gloeckle, F., Sootla, S., Gat, I., Tan, X.E., Adi, Y., Liu, J., Sauvestre, R., Remez, T., Rapin, J., Kozhevnikov, A., Evtimov, I., Bitton, J., Bhatt, M., Ferrer, C.C., Grattafiori, A., Xiong, W., D´ efossez, A., Copet, J., Azhar, F., Touvron, H., Ma...

  111. [119]

    arXiv (2023)

    Wang, X., Wei, J., Schuurmans, D., Le, Q., Chi, E., Narang, S., Chowdh- ery, A., Zhou, D.: Self-Consistency Improves Chain of Thought Reasoning in Language Models. arXiv (2023). https://doi.org/10.48550/arXiv.2203.11171 . http://arxiv.org/abs/2203.11171 Accessed 2026-03-05

  112. [120]

    In: Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pp

    Shin, R., Lin, C., Thomson, S., Chen Jr, C., Roy, S., Platanios, E.A., Pauls, A., Klein, D., Eisner, J., Van Durme, B.: Constrained language models yield few-shot semantic parsers. In: Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pp. ...

  113. [121]

    In: Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pp

    Scholak, T., Schucher, N., Bahdanau, D.: Picard: Parsing incrementally for con- strained auto-regressive decoding from language models. In: Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pp. 9895–9901 (2021)

  114. [122]

    In: Trec, vol

    Voorhees, E.M.,et al.: The trec-8 question answering track report. In: Trec, vol. 99, pp. 77–82 (1999) 44

  115. [397]

    https://doi.org/10.1007/11428817 45

    Springer, Berlin, Heidelberg (2005). https://doi.org/10.1007/11428817 45 . Series Title: Lecture Notes in Computer Science

  116. [471]

    https://doi.org/10.5220/0013594700004664

    SCITEPRESS - Science and Technology Publications, Hotel Crowne 41 Plaza pune, India (2025). https://doi.org/10.5220/0013594700004664 . https://www.scitepress.org/DigitalLibrary/Link.aspx?doi=10.5220/0013594700004664 Accessed 2026-01-01

Pith tools

Reviewed August 1, 2026 · model on record in the stance chip above.