REVIEW 5 major objections 5 minor 124 references
Large Language Models for Software Engineering Diagrams: A Systematic Review of UML and ER modelling
T0 review · 5 major / 5 minor · reviewed 2026-08-01 · deepseek-v4-flash
Pith's one-line read This review of 64 studies claims that LLM-based software diagramming is concentrated on UML class-diagram generation from natural language, while behavioural and ER data modelling, transformation, quality assurance, and shared benchmarks ar
desk verdict A useful first systematic map of LLM-based diagram modelling, with a defensible taxonomy and plausible qualitative findings — but the corpus needs auditing before the headline numbers can be trusted. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The working object is the multi-dimensional classification scheme (Fig. 4 taxonomy) into which every one of the 64 studies is placed: diagram type (behavioural vs structural UML vs ER), task category (construction, transformation, evaluation, understanding/assistance), technical configuration (model family, prompting strategy, architecture), and evaluation approach (metrics and datasets). The survey's distributions—which diagram types dominate, which tasks are rare, which models are used—are simply the counts over this taxonomy, so the taxonomy is the load-bearing instrument: the findings are true only if the categories are applied consistently and the corpus is complete. A secondary mechani
What would settle it
Re-run the stated protocol—same search string, databases, inclusion and exclusion criteria, and iterative snowballing—and check that the corpus reproduces exactly these 64 studies; then have a second team independently apply the Fig. 4 taxonomy to the same papers and compare category assignments. As a spot check, test the peer-review-only exclusion against the reference list, which includes a preprint-only study; if adding or removing that study changes the diagram-type or task distributions, the map is not stable under a defensible variant of the criteria.
Extended reading notes
Core claim
The authors claim that, across 64 studies from 2023–2025, LLM-based diagramming is concentrated on UML class-diagram generation from natural language using GPT-family models, while behavioural diagrams, entity-relationship and other data modelling, transformation, quality assurance, and multi-view consistency checking are underrepresented. They further claim that evaluation is heterogeneous—diverse metrics, custom datasets, little benchmark reuse, inconsistent statistical reporting—and that this review is the first systematic, diagram-centric synthesis of the literature, with its taxonomy of diagram types, tasks, techniques, and evaluation practices as the organising structure.
Load-bearing premise
The load-bearing premise is that the 64 studies assembled through the stated search string, screening criteria, and snowballing are a complete and representative sample of LLM diagram-modelling research; if retrieval or screening systematically selects GPT-focused, class-diagram papers, the reported concentrations describe the search rather than the field.
Editorial extensions
If this is right
- Research effort and funding in LLM diagramming will need to shift if the field is to cover behavioural diagrams and ER data modelling, the corners the review shows are most neglected.
- Because no shared benchmarks or common matching criteria exist, no current result is directly comparable to another; building a multi-diagram, multi-solution benchmark is the most direct structural improvement the review implies.
- Most systems are standalone, prompt-only invocations of GPT-family models, so integrating validation infrastructure—formal diagram rules, constraint checkers, statistical reporting—is the most plausible route from demonstrations to reliable tools.
- The prevalence of educational applications (tutoring, grading) suggests LLM diagram tools are most mature in teaching, so near-term deployment may happen in the classroom before industrial design.
- The persistent limitation reports across diagram types indicate that hallucinated elements, prompt sensitivity, and non-determinism are current-generation LLM properties rather than artifacts of a single approach.
Reading between the lines
- The GPT-prevalence result may be partly an artifact of the search string, which explicitly includes 'GPT' and 'ChatGPT' as query terms but no open-weights model names; a replication with open-model terms would test whether vendor concentration is a property of the field or of the retrieval.
- The ER-underrepresentation result rests on just two included studies, both dealing with ER evaluation; a broader search of database and data-modelling venues would show whether the ER gap is a property of the literature or of this review's selection.
- The paper's call for a benchmark with multiple valid reference solutions per specification implies a testable prediction: LLMs' measured quality relative to expert humans should improve under such a benchmark, because single-gold-standard evaluation penalises legitimate design alternatives.
- The absence of an auditable corpus—no screening counts, no excluded list, no extraction data—means the map can be falsified only by replicating the protocol; publishing the extraction data would turn this review into a reusable resource.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This manuscript presents a systematic literature review of 64 studies (claimed) on the use of large language models for UML and entity-relationship diagram modelling. It reports a protocol using six digital libraries, a defined search string, inclusion/exclusion criteria, two-reviewer screening with Cohen's kappa 0.738, and backward/forward snowballing. The results are organized around five research questions covering diagram types, modelling tasks, technical configurations, evaluation practices, and reported limitations. The paper's central claims are that class diagrams dominate, that diagram construction from natural language is the primary task, that GPT-based models are heavily prevalent, that behavioural diagrams and ER modelling are underrepresented, and that evaluation is heterogeneous with little benchmark reuse. The paper positions itself as the first systematic synthesis of LLM-based diagram modelling research.
Significance. If the corpus is complete and the classification is reliable, this would be a useful first systematic map of a young but fast-growing area. The paper contributes a taxonomy of diagram types and tasks, a structured overview of technical integration patterns, a summary of evaluation metrics and public datasets, and a set of future-research directions. The authors report a transparent search protocol, snowballing, and a threats-to-validity section, and Table 6 lists concrete public dataset URLs, which strengthens reproducibility. However, the value of the contribution rises or falls on the integrity of the 64-study corpus and the consistency of the manual classification that feeds every distributional headline.
major comments (5)
- [§4.1 and reference list] The corpus is not auditable as presented. References [61] and [79] are the same study (Al-Ahmad et al., Information 16(7), 565, 2025), so the reported total of 64 unique studies is at least one too high. This duplicate also inflates counts for class diagrams, deployment diagrams, and educational support in Fig. 4 and Table 2. Additionally, the manuscript ships no screening counts, no excluded-paper list, and no extraction dataset, so completeness and representativeness cannot be checked. Please provide a PRISMA-style flow diagram, a full list of excluded studies with reasons, a deduplicated reference list, and re-computed distributional counts.
- [§3.1 search string] The search string includes 'GPT' and 'ChatGPT' as explicit OR terms, while the abstract's headline finding states that GPT-based models are 'heavily prevalent'. This operationalization can preferentially retrieve papers that merely mention these vendor names, and it may bias the model-prevalence result. Please report how many included studies were retrieved only because of the vendor-specific terms, or rerun the search with a generic LLM-only string, and discuss whether the GPT-dominance claim survives that sensitivity analysis.
- [Table 1 and references [48], [97]] There are two inclusion-consistency defects that affect the corpus definition. Exclusion criterion (5) forbids preprints, yet [48] is explicitly an arXiv preprint (arXiv:2404.17739). Also, the abstract and §3.1 describe the window as 2023–2025, but [97] carries a 2026 publication date (Information and Software Technology 190, 107955, 2026). Please clarify whether [48] has since appeared in a peer-reviewed venue and whether [97] is an online-first/issue-date artifact; otherwise the inclusion decisions and the temporal distribution are inconsistent with the stated rules.
- [§3.2 and Fig. 4 taxonomy] The reported Cohen's kappa of 0.738 concerns title/abstract study selection only. No inter-rater reliability is reported for the manual classification of studies into the Fig. 4 taxonomy (diagram types, task categories, technical approaches, evaluation practices). Since every RQ1–RQ5 distributional claim is computed from that classification, please either report agreement for data extraction/classification, or describe a consensus-coding procedure with a resolved-disagreement log. The assignment of borderline scope cases also needs to be documented: for example, [42] addresses SysML diagrams, which are outside the UML/ER scope stated in the abstract, and the inclusion criteria in Table 1 should make clear whether such studies are covered under 'domain-specific modelling notations'.
- [§4.4.1 and Table 5] Some technical claims reference the wrong items. §4.4.1 states 'In [25], LoRA is used to fine-tune open-source LLMs...', but [25] is Zhang et al.'s survey 'A survey on large language models for software engineering', not a primary empirical study. This suggests a reference indexing error that should be corrected. In §4.5.1, 'Mean Absolute Percentage Error (MAPE)' is described as an element-level deviation measure for diagram generation; as defined, it is unusual in this context and its operationalization and suitability should be clarified.
minor comments (5)
- [§4.4.1] The text refers to 'Table X for full references', but no Table X appears in the manuscript. Add the missing table or remove the pointer.
- [§4.5.1 / Table 5] Table 5 labels 'Human Evaluation (Rubrics / Likert)' and 'Agreement Measures' as distinct metric categories, but the latter is a subtype of the former; consider a hierarchy or merged row to avoid double-counting.
- [Throughout] There are several minor typographical issues: 'T able 1' in §3.1, 'modelling Accuracy' in §4.6, and inconsistent capitalization such as 'F ew-shot' in §4.4.2. A careful copy-edit is needed.
- [§4.2 / Fig. 5] Fig. 5's distribution would benefit from explicit numeric counts per diagram type; the text and Fig. 4 contain the citations, but the visual is hard to interpret without the actual frequencies.
- [References] Several accessible-dates (e.g., 'Accessed 2026-03-22') are inconsistent with the stated 2023–2025 review window; align the access-date policy with the search and snowballing deadline described in §3.2.
Circularity Check
Partially circular: the 'GPT models are heavily prevalent' headline is partially baked into the search string; the remaining distributional findings are corpus-grounded and not query-forced.
-
fitted input called prediction
[Section 3.1 (Paper Collection Strategy); Abstract and Section 4.4.1 / RQ3.1 (LLM Model Families, Fig. 9)]
"We used the following search string consistently across all selected digital libraries: ("large language model" OR LLM OR "GPT" OR "ChatGPT" OR "generative AI" OR "Gen AI") AND ("UML" OR "class diagram" OR ... "ER diagram" OR "entity relationship model"). ... Abstract: "GPT-based models are heavily prevalent, raising concerns about reproducibility and vendor dependence.""
The headline claim of GPT prevalence is computed over a corpus retrieved with a query that explicitly lists 'GPT' and 'ChatGPT' as OR search terms. Papers whose titles/abstracts mention GPT/ChatGPT are therefore preferentially surfaced, while LLM-based diagram work that does not name GPT in its metadata is less likely to be captured by the database searches. The model-family distribution reported in Fig. 9 / RQ3.1 is thus partially a projection of the query's own term choices rather than an independent measurement of the field, i.e., the 'prediction' (vendor dominance) is partly baked into the selection input. The reduction is partial, not complete: generic terms ('LLM', 'large language model', 'generative AI') also retrieve non-GPT studies, so the diagram-type, task, and evaluation findin
full rationale
This is a systematic review, so its derivation chain is selection -> classification -> synthesis. The central distributional claims (class-diagram dominance, behavioural/ER underrepresentation, construction-from-NL as the primary task, heterogeneous evaluation, limited benchmark reuse) are computed from the 64-study corpus and do not reduce to the search instrument: the query weights all UML/ER diagram terms equally and includes generic LLM terms, so these findings are genuine corpus properties rather than artifacts of the query. The taxonomy follows the external UML Reference Manual [103], selection agreement is quantified (kappa 0.738), and no uniqueness theorem or ansatz is imported from the authors' prior work. The one partial circularity is the 'GPT-based models are heavily prevalent' headline: the Section 3.1 search string explicitly contains 'GPT' and 'ChatGPT' as OR retrieval terms, so GPT-mentioning papers are preferentially surfaced and the reported model-family dominance is in part constructed by the selection instrument. The remaining flagged issues are validity/correctness concerns, not circularity: references [61] and [79] are the same study (inflating N to 64), [48] is an arXiv preprint despite exclusion criterion (5), [97] is dated 2026 while the abstract states 2023-2025, [42] treats SysML rather than UML/ER, and [93] is the authors' own ER study (one of only two), making 'ER underrepresentation' mildly self-referential but robust even without that inclusion. The absence of screening counts, an excluded-paper list, and an extraction dataset limits auditability, but it does not make the central map equivalent to its inputs by construction.
Assumptions & free parameters
free parameters (4)
- Publication window start (January 2023) =
2023-01-01
- Search-string vendor terms ('GPT', 'ChatGPT', 'generative AI') =
included as OR terms
- Kappa agreement threshold interpretation =
0.738 ('substantial agreement')
- Snowballing stop rule =
iterations until Jan 1, 2026; stop when no new papers
assumptions (8)
- domain assumption UML and ER are the two dominant modelling notations; BPMN/SysML can be excluded without distorting the picture
- domain assumption LLM-based diagram research effectively began with the public release of GPT-3.5/GPT-4; earlier work is out of scope
- domain assumption Peer-reviewed publication is a valid quality filter for included studies
- domain assumption Two-reviewer screening with Cohen's kappa 0.738 yields reliable inclusion decisions
- domain assumption The UML Reference Manual [103] classification of diagrams into structural/behavioral is the correct organizing scheme
- domain assumption Diagram type, task, technique, and evaluation categories can be reliably extracted from paper descriptions; borderline cases resolved by author judgment
- domain assumption Snowballing until no new papers appear yields a complete corpus
- ad hoc to paper 'GPT'/'ChatGPT' as explicit search terms operationalize 'research using large language models for diagrams'
invented entities (2)
-
Four-category task taxonomy (construction, transformation, evaluation, understanding & assistance)
-
Technical integration typology (standalone, multi-stage, agent-based, RAG, human-in-the-loop)
Cite this review
Pith. "Pith review of Large Language Models for Software Engineering Diagrams: A Systematic Review of UML and ER modelling." pith.science (2026). https://pith.science/paper/UOCP5XN5
@misc{pith2026260726100,
author = {Pith},
title = {Pith review of: Large Language Models for Software Engineering Diagrams: A Systematic Review of UML and ER modelling},
year = {2026},
howpublished = {\url{https://pith.science/paper/UOCP5XN5}},
note = {Machine review of arXiv:2607.26100}
}
read the original abstract
Large language models (LLMs) are increasingly applied to diagram-based software and data modelling. Among various modelling notations, UML and entity-relationship (ER) diagrams are the most widely adopted for software modelling and data modelling, respectively. Recent literature has investigated various applications of LLMs in diagram modelling; however, their effectiveness and limitations have not been extensively discussed. This systematic literature review analyses 64 studies published between 2023 and 2025, examining diagram coverage, modelling tasks, technical approaches, evaluation practices, and limitations. Our findings reveal significant concentration patterns and gaps. UML-based software modelling strongly dominates, with class diagrams receiving the most attention whilst behavioural diagrams and data modelling remain underrepresented. Diagram construction from natural language is the primary focus, with limited work on transformation, quality assurance, and consistency checking. GPT-based models are heavily prevalent, raising concerns about reproducibility and vendor dependence. Evaluation practices are heterogeneous, employing diverse metrics and custom datasets with limited benchmark reuse and inconsistent reporting of robustness and statistical significance. Common limitations include semantic inaccuracies, hallucinated diagram elements, sensitivity to prompt formulation, and reproducibility constraints. This survey provides the first systematic synthesis of LLM-based diagram modelling research, highlighting needs for standardised benchmarks, stronger evaluation protocols, broader diagram coverage, and techniques for improving semantic reliability and multi-view consistency.
Reference graph
Works this paper leans on
-
[61]
Information16(7), 565 (2025)
Al-Ahmad, B., Alsobeh, A., Meqdadi, O., Shaikh, N.: A student-centric evalua- tion survey to explore the impact of llms on uml modeling. Information16(7), 565 (2025)
2025
-
[79]
Information 16(7), 565 (2025) https://doi.org/10.3390/info16070565
Al-Ahmad, B., Alsobeh, A., Meqdadi, O., Shaikh, N.: A Student-Centric Eval- uation Survey to Explore the Impact of LLMs on UML Modeling. Information 16(7), 565 (2025) https://doi.org/10.3390/info16070565 . Accessed 2026-01-01
-
[48]
Wang, B., Wang, C., Liang, P., Li, B., Zeng, C.: How llms aid in uml mod- eling: An exploratory study with novice analysts. arxiv 2024. arXiv preprint arXiv:2404.17739
arXiv 2024
-
[42]
In: Proceedings of the ACM/IEEE 27th International Conference on Model Driven Engineering Languages and Systems, pp
Sultan, B., Apvrille, L.: Ai-driven consistency of sysml diagrams. In: Proceedings of the ACM/IEEE 27th International Conference on Model Driven Engineering Languages and Systems, pp. 149–159 (2024)
2024
-
[97]
Information and Software Technology190, 107955 (2026) https://doi.org/10.1016/j.infsof
Liu, X., Liu, Y., Zhuang, Y., Hou, W.: UCD-LLM: A use case diagram require- ment modeling multi-agent framework with large language model. Information and Software Technology190, 107955 (2026) https://doi.org/10.1016/j.infsof. 2025.107955 . Accessed 2026-02-09
arXiv 2026
-
[25]
arXiv preprint arXiv:2312.15223 (2023) 34
Zhang, Q., Fang, C., Xie, Y., Zhang, Y., Yang, Y., Sun, W., Yu, S., Chen, Z.: A survey on large language models for software engineering. arXiv preprint arXiv:2312.15223 (2023) 34
arXiv 2023
-
[1]
Addison-Wesley, Upper Saddle River, NJ (2005)
Booch, G., Rumbaugh, J., Jacobson, I.: The Unified Modeling Language User Guide, 2nd ed edn. Addison-Wesley, Upper Saddle River, NJ (2005)
2005
-
[2]
Computer 39(02), 25–31 (2006)
Schmidt, D.C.: Guest editor’s introduction: Model-driven engineering. Computer 39(02), 25–31 (2006). Publisher: IEEE Computer Society
2006
Show all 124 references
-
[3]
OMG http://www
SysML, O.: Systems modeling language. OMG http://www. sysmlomg. org (2006) 32
2006
-
[4]
Addison-Wesley [u.a.], Harlow (2006)
Jackson, M.A.: Problem Frames: Analysing and Structuring Software Develop- ment Problems, Transferred to digital print edn. Addison-Wesley [u.a.], Harlow (2006)
2006
-
[5]
America: Pearson Education Inc (2011)
Sommerville, I.: Software engineering (ed.). America: Pearson Education Inc (2011)
2011
-
[6]
In: 25th International Con- ference on Software Engineering, 2003
Clements, P., Garlan, D., Little, R., Nord, R., Stafford, J.: Document- ing software architectures: views and beyond. In: 25th International Con- ference on Software Engineering, 2003. Proceedings., pp. 740–741. IEEE, Portland, OR, USA (2003). https://doi.org/10.1109/ICSE.2003...
2003 arXiv
-
[7]
Addison-Wesley Object Technology Series
Fowler, M.: UML Distilled: A Brief Guide to the Standard Object Modeling Lan- guage. Addison-Wesley Object Technology Series. Pearson Education, Boston, MA (2018).https://books.google.co.uk/books?id=VTdtDwAAQBAJ
2018
-
[8]
print edn
Rumbaugh, J., Jacobson, I., Booch, G.: The Unified Modeling Language Ref- erence Manual, 2. print edn. The Addison-Wesley object technology series. Addison-Wesley, Boston (2006)
2006
-
[9]
In: 2013 35th International Conference on Software Engineering (ICSE), pp
Petre, M.: UML in practice. In: 2013 35th International Conference on Software Engineering (ICSE), pp. 722–731. IEEE, San Fran- cisco, CA, USA (2013). https://doi.org/10.1109/ICSE.2013.6606618 . http://ieeexplore.ieee.org/document/6606618/Accessed 2026-03-09
2013
-
[10]
In: Proceedings of the 33rd International Conference on Software Engineering, pp
Hutchinson, J., Whittle, J., Rouncefield, M., Kristoffersen, S.: Empirical assess- ment of MDE in industry. In: Proceedings of the 33rd International Conference on Software Engineering, pp. 471–480 (2011)
2011
-
[11]
arXiv preprint arXiv:2107.03374 (2021)
Chen, M.: Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374 (2021)
2021 arXiv
-
[12]
Empirical Software Engineering30(3), 65 (2025)
Tambon, F., Moradi-Dakhel, A., Nikanjam, A., Khomh, F., Desmarais, M.C., Antoniol, G.: Bugs in large language models generated code: An empirical study. Empirical Software Engineering30(3), 65 (2025). Publisher: Springer
2025
-
[13]
In: 2023 IEEE/ACM 45th Inter- national Conference on Software Engineering (ICSE), pp
Mastropaolo, A., Pascarella, L., Guglielmi, E., Ciniselli, M., Scalabrino, S., Oliveto, R., Bavota, G.: On the Robustness of Code Generation Techniques: An Empirical Study on GitHub Copilot. In: 2023 IEEE/ACM 45th Inter- national Conference on Software Engineering (ICSE), pp. ...
2023
-
[14]
The Addison-Wesley object technology series
Kleppe, A.G., Warmer, J.B., Bast, W.: MDA Explained: the Model Driven Architecture: Practice and Promise. The Addison-Wesley object technology series. Addison-Wesley, Boston (2003) 33
2003
-
[15]
Software & Systems Modeling11(4), 513–526 (2012)
Selic, B.: What will it take? A view on adoption of model-based methods in practice. Software & Systems Modeling11(4), 513–526 (2012). Publisher: Springer
2012
-
[16]
329–380 (2001)
Spanoudakis, G., Zisman, A.: INCONSISTENCY MANAGEMENT IN SOFT- W ARE ENGINEERING: SUR VEY AND OPEN RESEARCH ISSUES, pp. 329–380 (2001). https://doi.org/10.1142/9789812389718 0015 . http://www. worldscientific.com/doi/abs/10.1142/9789812389718 0015 Accessed 2026-03-22
2001 doi
-
[17]
Computer33(4), 24–29 (2002)
Nuseibeh, B., Easterbrook, S., Russo, A.: Leveraging inconsistency in software development. Computer33(4), 24–29 (2002). Publisher: IEEE
2002
-
[18]
Software Quality Journal22(1), 121–149 (2014)
Landh¨ außer, M., K¨ orner, S.J., Tichy, W.F.: From requirements to UML models and back: how automatic processing of text can support requirements engineer- ing. Software Quality Journal22(1), 121–149 (2014). Publisher: Springer
2014
-
[19]
Ilieva, M.G., Ormandjieva, O.: Automatic Transition of Natural Language Soft- ware Requirements Specification into Formal Presentation. In: Hutchison, D., Kanade, T., Kittler, J., Kleinberg, J.M., Mattern, F., Mitchell, J.C., Naor, M., Nierstrasz, O., Pandu Rangan, C., Steffen...
-
[20]
In: CS&P, vol
Gr¨ opler, R., Sudhi, V., Garc ´ ıa, E.J.C., Bergmann, A.: NLP-Based Requirements Formalization for Automatic Test Case Generation. In: CS&P, vol. 21, pp. 18–30 (2021)
2021
-
[21]
IEEE Transactions on Software Engineering50(9), 2269–2293 (2024)
Marc´ en, A.C., Iglesias, A., Lape˜ na, R., P´ erez, F., Cetina, C.: A systematic literature review of model-driven engineering using machine learning. IEEE Transactions on Software Engineering50(9), 2269–2293 (2024)
2024
-
[22]
ACM Transactions on Database Systems1(1), 9–36 (1976)
Chen, P.P.-S.: The entity-relationship model—toward a unified view of data. ACM Transactions on Database Systems1(1), 9–36 (1976)
1976
-
[23]
ISBN-10137035152, 18 (2011)
Sommerville, I.: Software engineering 9th edition. ISBN-10137035152, 18 (2011)
2011
-
[24]
ACM Transactions on Software Engineering and Methodology 33(8), 1–79 (2024)
Hou, X., Zhao, Y., Liu, Y., Yang, Z., Wang, K., Li, L., Luo, X., Lo, D., Grundy, J., Wang, H.: Large language models for software engineering: A systematic literature review. ACM Transactions on Software Engineering and Methodology 33(8), 1–79 (2024)
2024
-
[26]
In: 2023 IEEE/ACM International Conference on Software Engineering: Future of Software Engineering (ICSE-FoSE), pp
Fan, A., Gokkaya, B., Harman, M., Lyubarskiy, M., Sengupta, S., Yoo, S., Zhang, J.M.: Large language models for software engineering: Survey and open prob- lems. In: 2023 IEEE/ACM International Conference on Software Engineering: Future of Software Engineering (ICSE-FoSE), pp....
2023
-
[27]
arXiv preprint arXiv:2409.02977 (2024)
Liu, J., Wang, K., Chen, Y., Peng, X., Chen, Z., Zhang, L., Lou, Y.: Large language model-based agents for software engineering: A survey. arXiv preprint arXiv:2409.02977 (2024)
2024 arXiv
-
[28]
Information and Software Technology169, 107423 (2024)
Naveed, H., Arora, C., Khalajzadeh, H., Grundy, J., Haggag, O.: Model driven engineering for machine learning components: A systematic literature review. Information and Software Technology169, 107423 (2024)
2024
-
[29]
Software and Systems Modeling 24(2), 445–469 (2025)
R¨ adler, S., Berardinelli, L., Winter, K., Rahimi, A., Rinderle-Ma, S.: Bridging mde and ai: a systematic review of domain-specific languages and model-driven practices in ai software systems engineering. Software and Systems Modeling 24(2), 445–469 (2025)
2025
-
[30]
di rocco et al
Di Rocco, J., Di Ruscio, D., Di Sipio, C., Nguyen, P.T., Rubei, R.: On the use of large language models in model-driven engineering: J. di rocco et al. Software and Systems Modeling24(3), 923–948 (2025)
2025
-
[31]
ACM Transactions on Software Engineering and Methodology34(5), 1–25 (2025)
Burgue˜ no, L., Di Ruscio, D., Sahraoui, H., Wimmer, M.: Automation in model- driven engineering: A look back, and ahead. ACM Transactions on Software Engineering and Methodology34(5), 1–25 (2025)
2025
-
[32]
Frontiers in Computer Science7, 1519437 (2025)
Hemmat, A., Sharbaf, M., Kolahdouz-Rahimi, S., Lano, K., Tehrani, S.Y.: Research directions for using llm in software requirement engineering: A systematic review. Frontiers in Computer Science7, 1519437 (2025)
2025
-
[33]
arXiv preprint arXiv:2509.11446 (2025)
Zadenoori, M.A., Dabrowski, J., Alhoshan, W., Zhao, L., Ferrari, A.: Large language models (llms) for requirements engineering (re): A systematic literature review. arXiv preprint arXiv:2509.11446 (2025)
2025
-
[34]
Software: Practice and Experience56(2), 141–170 (2026)
Cheng, H., Husen, J.H., Lu, Y., Racharak, T., Yoshioka, N., Ubayashi, N., Washizaki, H.: Generative ai for requirements engineering: A systematic litera- ture review. Software: Practice and Experience56(2), 141–170 (2026)
2026
-
[35]
applica- tions, challenges, and future directions
Esposito, M., Li, X., Moreschini, S., Ahmad, N., Cerny, T., Vaidhyanathan, K., Lenarduzzi, V., Taibi, D.: Generative ai for software architecture. applica- tions, challenges, and future directions. Journal of Systems and Software, 112607 (2025)
2025
-
[36]
arXiv preprint arXiv:2505.16697 (2025) 35
Schmid, L., Hey, T., Armbruster, M., Corallo, S., Fuchß, D., Keim, J., Liu, H., Koziolek, A.: Software architecture meets llms: A systematic literature review. arXiv preprint arXiv:2505.16697 (2025) 35
2025 arXiv
-
[37]
Technical report, Technical report, ver
Keele, S., et al.: Guidelines for performing systematic literature reviews in soft- ware engineering. Technical report, Technical report, ver. 2.3 ebse technical report. ebse (2007)
2007
-
[38]
ACM Transactions on Software Engineering and Methodology33(5), 1–59 (2024)
Chen, Z., Zhang, J.M., Hort, M., Harman, M., Sarro, F.: Fairness testing: A comprehensive survey and analysis of trends. ACM Transactions on Software Engineering and Methodology33(5), 1–59 (2024). Publisher: ACM New York, NY
2024
-
[39]
In: International Conference on Advanced Information Systems Engineering, pp
Reinhartz-Berger, I., Ali, S.J., Bork, D.: Leveraging llms for domain modeling: The impact of granularity and strategy on quality. In: International Conference on Advanced Information Systems Engineering, pp. 3–19 (2025). Springer
2025
-
[40]
Procedia Computer Science246, 1346–1354 (2024)
Naimi, L., Jakimi, A., Saadane, R., Chehri, A.,et al.: Automating software doc- umentation: Employing llms for precise use case description. Procedia Computer Science246, 1346–1354 (2024)
2024
-
[41]
In: Proceedings of the ACM/IEEE 27th International Conference on Model Driven Engineering Languages and Systems, pp
Tabassum, M.R., Ritchie, M.J., Mustafiz, S., Kienzle, J.: Using llms for use case modelling of iot systems: An experience report. In: Proceedings of the ACM/IEEE 27th International Conference on Model Driven Engineering Languages and Systems, pp. 611–619 (2024)
2024
-
[43]
In: 37th International BCS Human-Computer Interaction Conference, pp
O’Neill, I.: Getting gpt to answer like me. In: 37th International BCS Human-Computer Interaction Conference, pp. 7–13 (2024). BCS Learning & Development
2024
-
[44]
In: Proceedings of the 33rd ACM International Conference on the Foundations of Software Engineering, pp
Klimek, R.: Re-oriented model development with llm support and deduction- based verification. In: Proceedings of the 33rd ACM International Conference on the Foundations of Software Engineering, pp. 1297–1304 (2025)
2025
-
[45]
In: SoutheastCon 2025, pp
Ramachandran, R.: Transforming software architecture design with intelligent assistants-a comparative analysis. In: SoutheastCon 2025, pp. 1446–1454 (2025). IEEE
2025
-
[46]
In: 2025 25th International Conference on Software Quality, Reliability and Security (QRS), pp
Lu, J., Sun, P., Chen, Y., Yin, G., Ye, P.: Aug: an interactive tool for clarify- ing and generating uml models based on large language models. In: 2025 25th International Conference on Software Quality, Reliability and Security (QRS), pp. 78–85 (2025). IEEE
2025
-
[47]
In: 2025 IEEE International Confer- ence on Software Analysis, Evolution and Reengineering (SANER), pp
Hassine, J.: Evaluating multi-modal llms for automatically recognizing semantic elements in uml use case diagram images. In: 2025 IEEE International Confer- ence on Software Analysis, Evolution and Reengineering (SANER), pp. 861–866 (2025). IEEE 36
2025
-
[49]
Machine Learning with Applications, 100660 (2025)
Bates, A., Vavricka, R., Carleton, S., Shao, R., Pan, C.: Unified modeling lan- guage code generation from diagram images using multimodal large language models. Machine Learning with Applications, 100660 (2025)
2025
-
[50]
In: Proceedings of the 2024 IEEE/ACM 46th International Conference on Software Engineering: Companion Proceedings, pp
Siala, H.A.: Enhancing model-driven reverse engineering using machine learn- ing. In: Proceedings of the 2024 IEEE/ACM 46th International Conference on Software Engineering: Companion Proceedings, pp. 173–175 (2024)
2024
-
[51]
In: Proceedings of the 30th ACM Conference on Innovation and Technology in Computer Science Education V
Ouh, E.L., Tan, K.W., Lo, S.L., Gan, B.K.S.: Evaluating chatgpt to answer multi-modal exercises in computer science education. In: Proceedings of the 30th ACM Conference on Innovation and Technology in Computer Science Education V. 1, pp. 58–64 (2025)
2025
-
[52]
a rule-based approach
Jahan, M., Hassan, M.M., Golpayegani, R., Ranjbaran, G., Roy, C., Roy, B., Schneider, K.: Automated derivation of uml sequence diagrams from user stories: Unleashing the power of generative ai vs. a rule-based approach. In: Proceedings of the ACM/IEEE 27th International Confer...
2024
-
[53]
In: 2024 36th International Con- ference on Software Engineering Education and Training (CSEE&T), pp
Speth, S., Meißner, N., Becker, S.: Chatgpt’s aptitude in utilizing uml diagrams for software engineering exercise generation. In: 2024 36th International Con- ference on Software Engineering Education and Training (CSEE&T), pp. 1–5 (2024). IEEE
2024
-
[54]
In: 2025 IEEE/ACM 47th International Conference on Software Engineering: Software Engineering in Practice (ICSE-SEIP), pp
Xiao, C., St ˚ ahl, D., Bosch, J.: Uml sequence diagram generation: A multi-model, multi-domain evaluation. In: 2025 IEEE/ACM 47th International Conference on Software Engineering: Software Engineering in Practice (ICSE-SEIP), pp. 272–283 (2025). IEEE
2025
-
[55]
In: 2024 IEEE 32nd International Requirements Engineering Conference Workshops (REW), pp
Ferrari, A., Abualhaija, S., Arora, C.: Model generation with llms: From requirements to uml sequence diagrams. In: 2024 IEEE 32nd International Requirements Engineering Conference Workshops (REW), pp. 291–300 (2024). IEEE
2024
-
[56]
Theoretical Computer Science1021, 114879 (2024)
Zhao, Z., Zhang, N., Yu, B., Duan, Z.: Generating java code pairing with chatgpt. Theoretical Computer Science1021, 114879 (2024)
2024
-
[57]
In: 2023 IEEE/ACM 45th Interna- tional Conference on Software Engineering: New Ideas and Emerging Results (ICSE-NIER), pp
Chaaben, M.B., Burgue˜ no, L., Sahraoui, H.: Towards using few-shot prompt learning for automating model completion. In: 2023 IEEE/ACM 45th Interna- tional Conference on Software Engineering: New Ideas and Emerging Results (ICSE-NIER), pp. 7–12 (2023). IEEE
2023
-
[58]
In: 37 Proceedings of the ACM/IEEE 27th International Conference on Model Driven Engineering Languages and Systems, pp
Ben Chaaben, M.: Software modeling assistance with large language models. In: 37 Proceedings of the ACM/IEEE 27th International Conference on Model Driven Engineering Languages and Systems, pp. 188–191 (2024)
2024
-
[59]
In: 2025 IEEE 22nd International Conference on Software Architecture Companion (ICSA-C), pp
Moezkarimi, Z., Eriksson, K., Johansson, A.A., Bucaioni, A., Sirjani, M.: Har- nessing chatgpt for model transformation in software architecture: From uml state diagrams to rebeca models for formal verification. In: 2025 IEEE 22nd International Conference on Software Architect...
2025
-
[60]
Software and Systems Modeling22(3), 781–793 (2023)
C´ amara, J., Troya, J., Burgue˜ no, L., Vallecillo, A.: On the assessment of gener- ative ai in modeling tasks: an experience report with chatgpt and uml. Software and Systems Modeling22(3), 781–793 (2023)
2023
-
[62]
Computer Applications in Engineering Education33(5), 70080 (2025)
Ib´ a˜ nez, M.B., Barr´ on-Estrada, M.L., Zatarain-Cabada, R.: Can multimodal large language models grade like an expert? a study on uml class diagram assess- ment accuracy. Computer Applications in Engineering Education33(5), 70080 (2025)
2025
-
[63]
In: Proceedings of the ACM/IEEE 27th International Conference on Model Driven Engineering Languages and Systems, pp
Yang, Y., Chen, B., Chen, K., Mussbacher, G., Varr´ o, D.: Multi-step iterative automated domain modeling with large language models. In: Proceedings of the ACM/IEEE 27th International Conference on Model Driven Engineering Languages and Systems, pp. 587–595 (2024)
2024
-
[64]
In: Proceedings of the 33rd ACM International Conference on the Foundations of Software Engineering, pp
Wang, T., Trimble, M., Brown, C.: Devcoach: Supporting students learning the software development life cycle with a generative ai powered multi-agent system. In: Proceedings of the 33rd ACM International Conference on the Foundations of Software Engineering, pp. 987–998 (2025)
2025
-
[65]
In: Proceed- ings of the 18th ACM/IEEE International Symposium on Empirical Software Engineering and Measurement, pp
De Bari, D., Garaccione, G., Coppola, R., Torchiano, M., Ardito, L.: Evaluating large language models in exercises of uml class diagram modeling. In: Proceed- ings of the 18th ACM/IEEE International Symposium on Empirical Software Engineering and Measurement, pp. 393–399 (2024)
2024
-
[66]
In: Proceedings of the ACM/IEEE 27th International Conference on Model Driven Engineering Lan- guages and Systems, pp
Ardimento, P., Bernardi, M.L., Cimitile, M., Scalera, M.: Enhancing soft- ware modeling learning with ai-powered scaffolding. In: Proceedings of the ACM/IEEE 27th International Conference on Model Driven Engineering Lan- guages and Systems, pp. 103–106 (2024)
2024
-
[67]
In: Proceedings of the ACM/IEEE 27th International Conference on Model Driven Engineering Languages and Systems, pp
Ardimento, P., Bernardi, M.L., Cimitile, M., Scalera, M.: A rag-based feedback tool to augment uml class diagram learning. In: Proceedings of the ACM/IEEE 27th International Conference on Model Driven Engineering Languages and Systems, pp. 26–30 (2024) 38
2024
-
[68]
In: Proceedings of the 1st International Workshop on Large Language Models for Code, pp
Antal, G., Voz´ ar, R., Ferenc, R.: Toward a new era of rapid development: Assess- ing gpt-4-vision’s capabilities in uml-based code generation. In: Proceedings of the 1st International Workshop on Large Language Models for Code, pp. 84–87 (2024)
2024
-
[69]
In: Proceedings of the 27th International Conference on Evaluation and Assessment in Software Engineering, pp
Ahmad, A., Waseem, M., Liang, P., Fahmideh, M., Aktar, M.S., Mikkonen, T.: Towards human-bot collaborative software architecting with chatgpt. In: Proceedings of the 27th International Conference on Evaluation and Assessment in Software Engineering, pp. 279–285 (2023)
2023
-
[70]
In: Proceedings of the 27th ACM Inter- national Systems and Software Product Line Conference-Volume B, pp
Acher, M., Martinez, J.: Generative ai for reengineering variants into software product lines: an experience report. In: Proceedings of the 27th ACM Inter- national Systems and Software Product Line Conference-Volume B, pp. 57–66 (2023)
2023
-
[71]
In: Proceedings of the 2024 IEEE/ACM First International Conference on AI Foundation Models and Software Engineer- ing, pp
Abukhalaf, S., Hamdaqa, M., Khomh, F.: Pathocl: path-based prompt augmen- tation for ocl generation with gpt-4. In: Proceedings of the 2024 IEEE/ACM First International Conference on AI Foundation Models and Software Engineer- ing, pp. 108–118 (2024)
2024
-
[72]
IEEE Access (2025)
Babaalla, Z., Jakimi, A., Oualla, M.: Llm-driven mda pipeline for generating uml class diagrams and code. IEEE Access (2025)
2025
-
[73]
In: Proceedings of the 46th International Conference on Software Engineering: Software Engineering Education and Training, pp
Xue, Y., Chen, H., Bai, G.R., Tairas, R., Huang, Y.: Does chatgpt help with introductory programming? an experiment of students using chatgpt in cs1. In: Proceedings of the 46th International Conference on Software Engineering: Software Engineering Education and Training, pp. ...
2024
-
[74]
In: 2024 IEEE 48th Annual Computers, Software, and Applications Conference (COMPSAC), pp
Li, Y., Keung, J., Ma, X., Chong, C.Y., Zhang, J., Liao, Y.: Llm-based class diagram derivation from user stories with chain-of-thought promptings. In: 2024 IEEE 48th Annual Computers, Software, and Applications Conference (COMPSAC), pp. 45–50 (2024). IEEE
2024
-
[75]
In: 2023 IEEE/ACM 20th International Conference on Mining Software Repositories (MSR), pp
Abukhalaf, S., Hamdaqa, M., Khomh, F.: On codex prompt engineering for ocl generation: an empirical study. In: 2023 IEEE/ACM 20th International Conference on Mining Software Repositories (MSR), pp. 148–157 (2023). IEEE
2023
-
[76]
In: 2024 International Joint Conference on Neural Networks (IJCNN), pp
Ardimento, P., Bernardi, M.L., Cimitile, M.: Teaching uml using a rag-based llm. In: 2024 International Joint Conference on Neural Networks (IJCNN), pp. 1–8 (2024). IEEE
2024
-
[77]
In: 2024 IEEE Working Conference on Software Visualization (VISSOFT), pp
Shehata, M., Lepore, B., Cummings, H., Parra, E.: Creating uml class dia- grams with general-purpose llms. In: 2024 IEEE Working Conference on Software Visualization (VISSOFT), pp. 157–158 (2024). IEEE
2024
-
[78]
In: 2025 8th International Conference on Software and 39 System Engineering (ICoSSE), pp
Siala, H.A., Lano, K.: Using large language models to extract uml class diagrams from java programs. In: 2025 8th International Conference on Software and 39 System Engineering (ICoSSE), pp. 70–74 (2025). IEEE
2025
-
[80]
In: 2025 International Conference on Next Generation Information System Engineering (NGISE), pp
Ojha, M., Gupta, S., Sharma, R.: Towards Class Diagram Generation from User Stories Using LLMs. In: 2025 International Conference on Next Generation Information System Engineering (NGISE), pp. 1–6. IEEE, Ghaziabad, Delhi (NCR), India (2025). https://doi.org/10.1109/NGISE64126....
2025
-
[81]
IEEE Access13, 211605–211619 (2025) https://doi.org/10.1109/ACCESS.2025
Nair, R.P., Thushara, M.G., Sugumaran, V.: Fine-Tuned LLMs Versus Rule- Based NLP for UML Diagram Generation: An Educational Evaluation. IEEE Access13, 211605–211619 (2025) https://doi.org/10.1109/ACCESS.2025. 3638372 . Accessed 2026-01-01
2025 doi
-
[82]
In: 2025 IEEE/ACM International Workshop on Nat- ural Language-Based Software Engineering (NLBSE), pp
Garaccione, G., Vega Carrazan, P.F., Coppola, R., Ardito, L.: Eval- uating Large Language Models in Exercises of UML Use Case Dia- grams Modeling. In: 2025 IEEE/ACM International Workshop on Nat- ural Language-Based Software Engineering (NLBSE), pp. 41–44. IEEE, Ottawa, ON, Ca...
2025
-
[83]
In: Proceedings of the 13th International Conference on Model-Based Software and Systems Engineering, pp
Kop, C.: Evaluating the Quality of Class Diagrams Created by a Genera- tive AI: Findings, Guidelines and Automation Options:. In: Proceedings of the 13th International Conference on Model-Based Software and Systems Engineering, pp. 150–157. SCITEPRESS - Science and Technology ...
2025 doi
-
[84]
In: Maass, W., Han, H., Yasar, H., Multari, N
Silva, J., Ma, Q., Cabot, J., Kelsen, P., Proper, H.A.: Application of the Tree-of- Thoughts Framework to LLM-Enabled Domain Modeling. In: Maass, W., Han, H., Yasar, H., Multari, N. (eds.) Conceptual Modeling vol. 15238, pp. 94–111. Springer, Cham (2025). https://doi.org/10.10...
2025 doi
-
[85]
In: Proceedings of the 17th International Conference on Computer Supported Education, pp
Bouali, N., Gerhold, M., Rehman, T., Ahmed, F.: Toward Automated UML Dia- gram Assessment: Comparing LLM-Generated Scores with Teaching Assistants:. In: Proceedings of the 17th International Conference on Computer Supported Education, pp. 158–169. SCITEPRESS - Science and Tech...
2025 doi
-
[86]
Enterprise Mod- elling and Information Systems Architectures (EMISAJ), 3–115 (2023) https: //doi.org/10.18417/EMISA.18.3
Fill, H.-G., Fettke, P., K¨ opke, J.: Conceptual Modeling and Large Language 40 Models: Impressions From First Experiments With ChatGPT. Enterprise Mod- elling and Information Systems Architectures (EMISAJ), 3–115 (2023) https: //doi.org/10.18417/EMISA.18.3 . Artwork Size: 3:1...
2023 doi
-
[87]
Journal of Systems and Software, 112709 (2025)
Wang, C., Wang, B., Liang, P., Liang, J.: Assessing uml diagrams by gpt: Implications for education. Journal of Systems and Software, 112709 (2025)
2025
-
[88]
In 2023 ACM/IEEE 26th International Conference on Model Driven Engineering Languages and Systems (MODELS)
Chen, K., Yang, Y., Chen, B., L´ opez, J.A.H., Mussbacher, G., Varr´ o, D.: Auto- mated Domain Modeling with Large Language Models: A Comparative Study. In 2023 ACM/IEEE 26th International Conference on Model Driven Engineering Languages and Systems (MODELS). 162–172 (2023)
2023
-
[89]
In: International Conference on Future Data and Security Engineering, pp
Nguyen, V.-V., Nguyen, H.-K., Nguyen, K.-S., Luong Thi, M.-H., Nguyen, T.-V., Vu, D.-Q.: Automated uml generation: A framework for class diagram synthesis and multimodal validation. In: International Conference on Future Data and Security Engineering, pp. 212–224 (2025). Springer
2025
-
[90]
In: Proceedings of the IEEE/ACM 47th International Conference on Software Engineering
Fuchß, D., Hey, T., Keim, J., Liu, H., Ewald, N., Thirolf, T., Koziolek, A.: Lissa: toward generic traceability link recovery through retrieval-augmented gen- eration. In: Proceedings of the IEEE/ACM 47th International Conference on Software Engineering. ICSE, vol. 25 (2025)
2025
-
[91]
In: 2025 IEEE 22nd Interna- tional Conference on Software Architecture Companion (ICSA-C), pp
Tagliaferro, A., Corboe, S., Guindani, B.: Leveraging llms to automate software architecture design from informal specifications. In: 2025 IEEE 22nd Interna- tional Conference on Software Architecture Companion (ICSA-C), pp. 291–299 (2025). IEEE
2025
-
[92]
In: 2023 IEEE 11th International Conference on Systems and Control (ICSC), pp
Omar, M.A.: Measurement of chatgpt performance in mapping natural lan- guage speficaction into an entity relationship diagram. In: 2023 IEEE 11th International Conference on Systems and Control (ICSC), pp. 530–535 (2023). IEEE
2023
-
[93]
Cogent Education12(1), 2590901 (2025) https://doi
Rahmanian, M., Sami, A., Yu, Y.: Challenges and feasibility of multimodal LLMs in ER diagram evaluation. Cogent Education12(1), 2590901 (2025) https://doi. org/10.1080/2331186X.2025.2590901 . Accessed 2026-01-02
2025
-
[94]
In: Proceedings of the 16th ACM International Conference on Future and Sustainable Energy Systems, pp
Kuchenbuch, R., Lehnhoff, S., Sauer, J.: Smart grid assistive ai in requirement engineering: Improving the modeling of use cases and architecture models with llms. In: Proceedings of the 16th ACM International Conference on Future and Sustainable Energy Systems, pp. 495–504 (2025)
2025
-
[95]
In: Proceedings of the 3rd International Conference on Futuristic Technology, pp
G S, N.K., S, A., Thushara, M.G.: Comparative Analysis of Large Lan- guage Models for Automated Use Case Diagram Generation:. In: Proceedings of the 3rd International Conference on Futuristic Technology, pp. 465–
-
[96]
In: Guizzardi, R., Pufahl, L., Sturm, A., Van Der Aa, H
Calamo, M., Mecella, M., Snoeck, M.: Assessing the Suitability of Large Lan- guage Models in Generating UML Class Diagrams as Conceptual Models. In: Guizzardi, R., Pufahl, L., Sturm, A., Van Der Aa, H. (eds.) Enterprise, Business-Process and Information Systems Modeling vol. 5...
2025 doi
-
[98]
In: 2025 40th IEEE/ACM International Conference on Automated Software Engineering Workshops (ASEW), pp
Eisenreich, T., Friedlaender, N., Wagner, S.: Leveraging large language mod- els for use case model generation from software requirements. In: 2025 40th IEEE/ACM International Conference on Automated Software Engineering Workshops (ASEW), pp. 221–227 (2025). IEEE
2025
-
[99]
In: International Conference on Conceptual Modeling, pp
Gavric, A., Bork, D., Proper, H.A.: How does uml look and sound? using ai to interpret uml diagrams through multimodal evidence. In: International Conference on Conceptual Modeling, pp. 187–197 (2024). Springer
2024
-
[100]
Big Data and Cognitive Computing 10(1), 2 (2025) https://doi.org/10.3390/bdcc10010002
Ramachandran, R., Vijayan, P., Anilkumar, A., Gangadharan, V.: AI Assisted System for Automated Evaluation of Entity-Relationship Diagram and Schema Diagram Using Large Language Models. Big Data and Cognitive Computing 10(1), 2 (2025) https://doi.org/10.3390/bdcc10010002 . Acc...
2025 doi
-
[101]
Information16(5), 368 (2025) https://doi.org/10.3390/info16050368
Avignone, A., Tierno, A., Fiori, A., Chiusano, S.: Exploring Large Language Models’ Ability to Describe Entity-Relationship Schema-Based Conceptual Data Models. Information16(5), 368 (2025) https://doi.org/10.3390/info16050368 . Accessed 2026-01-01
2025 doi
-
[102]
CAAI Transactions on Intelligence Technology9(1), 250–263 (2024)
Yang, Y., Liu, Y., Bao, T., Wang, W., Niu, N., Yin, Y.: Deepocl: A deep neu- ral network for object constraint language generation from unrestricted nature language. CAAI Transactions on Intelligence Technology9(1), 250–263 (2024)
2024
-
[103]
Pearson Higher Education, ??? (2004)
Rumbaugh, J., Jacobson, I., Booch, G.: Unified Modeling Language Reference Manual, The (2nd Edition). Pearson Higher Education, ??? (2004)
2004
-
[104]
Advances in Neural Information Processing Systems37, 50528–50652 (2024) 42
Yang, J., Jimenez, C.E., Wettig, A., Lieret, K., Yao, S., Narasimhan, K., Press, O.: Swe-agent: Agent-computer interfaces enable automated software engineer- ing. Advances in Neural Information Processing Systems37, 50528–50652 (2024) 42
2024
-
[105]
arXiv preprint arXiv:2207.10397 (2022)
Chen, B., Zhang, F., Nguyen, A., Zan, D., Lin, Z., Lou, J.-G., Chen, W.: Codet: Code generation with generated tests. arXiv preprint arXiv:2207.10397 (2022)
2022 arXiv
-
[106]
experience: Evalu- ating the usability of code generation tools powered by large language models
Vaithilingam, P., Zhang, T., Glassman, E.L.: Expectation vs. experience: Evalu- ating the usability of code generation tools powered by large language models. In: Chi Conference on Human Factors in Computing Systems Extended Abstracts, pp. 1–7 (2022)
2022
-
[107]
arXiv preprint arXiv:2108.07732 (2021)
Austin, J., Odena, A., Nye, M., Bosma, M., Michalewski, H., Dohan, D., Jiang, E., Cai, C., Terry, M., Le, Q., et al.: Program synthesis with large language models. arXiv preprint arXiv:2108.07732 (2021)
2021 arXiv
-
[108]
In: Future of Software Engineering (FOSE’07), pp
France, R., Rumpe, B.: Model-driven development of complex software: A research roadmap. In: Future of Software Engineering (FOSE’07), pp. 37–54 (2007). IEEE
2007
-
[109]
In: Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pp
Rajpurkar, P., Zhang, J., Lopyrev, K., Liang, P.: Squad: 100,000+ questions for machine comprehension of text. In: Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pp. 2383–2392 (2016)
2016
-
[110]
In: Proceedings of the 2018 EMNLP Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP, pp
Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., Bowman, S.: Glue: A multi-task benchmark and analysis platform for natural language understand- ing. In: Proceedings of the 2018 EMNLP Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP, pp. 353–355 (2018)
2018
-
[111]
Springer (2003)
Handbook, A.: From contract drafting to software specification: Linguistic sources of ambiguity. Springer (2003)
2003
-
[112]
ACM Transactions on Software Engineering and Methodology (TOSEM)27(3), 1–51 (2018)
Stol, K.-J., Fitzgerald, B.: The abc of software engineering research. ACM Transactions on Software Engineering and Methodology (TOSEM)27(3), 1–51 (2018)
2018
-
[113]
In: Findings of the Association for Computational Linguistics ACL 2024, pp
Dhuliawala, S., Komeili, M., Xu, J., Raileanu, R., Li, X., Celikyilmaz, A., Weston, J.: Chain-of-Verification Reduces Hallucination in Large Language Models. In: Findings of the Association for Computational Linguistics ACL 2024, pp. 3563–3578. Association for Computational Li...
2024 doi
-
[114]
ACM Computing Surveys55(12), 1–38 (2023) https://doi.org/10.1145/3571730
Ji, Z., Lee, N., Frieske, R., Yu, T., Su, D., Xu, Y., Ishii, E., Bang, Y.J., Madotto, A., Fung, P.: Survey of Hallucination in Natural Language Generation. ACM Computing Surveys55(12), 1–38 (2023) https://doi.org/10.1145/3571730 . Accessed 2026-03-05
2023 doi
- [115]
-
[116]
Iclr1(2), 3 (2022)
Hu, E.J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., Chen, W.,et al.: Lora: Low-rank adaptation of large language models. Iclr1(2), 3 (2022)
2022
-
[117]
Advances in neural information processing systems36, 10088–10115 (2023)
Dettmers, T., Pagnoni, A., Holtzman, A., Zettlemoyer, L.: Qlora: Efficient fine- tuning of quantized llms. Advances in neural information processing systems36, 10088–10115 (2023)
2023
- [118]
-
[119]
arXiv (2023)
Wang, X., Wei, J., Schuurmans, D., Le, Q., Chi, E., Narang, S., Chowdh- ery, A., Zhou, D.: Self-Consistency Improves Chain of Thought Reasoning in Language Models. arXiv (2023). https://doi.org/10.48550/arXiv.2203.11171 . http://arxiv.org/abs/2203.11171 Accessed 2026-03-05
-
[120]
In: Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pp
Shin, R., Lin, C., Thomson, S., Chen Jr, C., Roy, S., Platanios, E.A., Pauls, A., Klein, D., Eisner, J., Van Durme, B.: Constrained language models yield few-shot semantic parsers. In: Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pp. ...
2021
-
[121]
In: Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pp
Scholak, T., Schucher, N., Bahdanau, D.: Picard: Parsing incrementally for con- strained auto-regressive decoding from language models. In: Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pp. 9895–9901 (2021)
2021
-
[122]
In: Trec, vol
Voorhees, E.M.,et al.: The trec-8 question answering track report. In: Trec, vol. 99, pp. 77–82 (1999) 44
1999
-
[397]
https://doi.org/10.1007/11428817 45
Springer, Berlin, Heidelberg (2005). https://doi.org/10.1007/11428817 45 . Series Title: Lecture Notes in Computer Science
2005 doi
-
[471]
https://doi.org/10.5220/0013594700004664
SCITEPRESS - Science and Technology Publications, Hotel Crowne 41 Plaza pune, India (2025). https://doi.org/10.5220/0013594700004664 . https://www.scitepress.org/DigitalLibrary/Link.aspx?doi=10.5220/0013594700004664 Accessed 2026-01-01
2025 doi
Reviewed August 1, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.