REVIEW 5 major objections 5 minor 48 references
Using ChatGPT-4 for the Identification of Common UX Factors within a Pool of Measurement Items from Established UX Questionnaires
T0 review · 5 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read ChatGPT-4 can sort 408 UX questionnaire items into shared semantic topics, and can filter the pool for a concept like learnability.
desk verdict A transparent but under-evidenced demonstration that ChatGPT-4 can cluster UX items; the filtering claim is refuted by the paper's own appendix. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central mechanism is ChatGPT-4's embedding-based semantic-textual-similarity ability, steered by a fixed sequence of seven prompts that move from broad classification to detailed subcategorization, self-improvement, comparison with the literature, generalization, and finally a concept-specific filter. This prompt chain converts the model's continuous similarity judgments into an explicit topic hierarchy plus an item-selection procedure.
What would settle it
Have a panel of UX experts independently sort the same 408 items into topics and compare their categories to ChatGPT-4's clusters; if agreement is near chance, or if repeated ChatGPT-4 runs with the same prompts produce wildly different taxonomies, the central claim that these are meaningful topics collapses. A complementary test is to check whether items the model groups together show high correlations in real survey responses, since the paper itself notes that semantic similarity and empirical similarity can diverge.
Extended reading notes
Core claim
On the paper's own terms, the discovery is that ChatGPT-4, given only the texts of 408 measurement items, produces a two-level taxonomy of six main topics and 15 subtopics that aligns with an established consolidation of 16 UX quality aspects, and that the same model can act as an item finder for a specific UX concept. The authors report that the AI-generated topics capture both task-related and emotional aspects, that the functional topics are generated particularly well, and that some hedonic factors from the literature, such as Novelty and Identity, are poorly covered. They also report that asking the model to select items describing learnability returns a top-15 list whose entries match the concept well.
Load-bearing premise
The paper assumes that the authors' visual inspection of ChatGPT-4's output is enough to prove the topics are meaningful and the item fits correct, without any quantitative validation against human experts or empirical item correlations.
Editorial extensions
If this is right
- Researchers can cheaply explore the semantic structure of any large item pool without manual coding.
- Items from different questionnaires can be compared on a common semantic basis, making it easier to select measures for ad-hoc surveys.
- The AI-generated topics could serve as a scaffold for a new holistic UX questionnaire.
- The filtering capability lets practitioners quickly reuse existing items for a target UX concept instead of writing new ones.
- Because LLM output is non-deterministic, repeated runs can yield multiple alternative classifications, which the paper sees as an explorative advantage.
Reading between the lines
- The plausibility claim could be made quantitative by measuring agreement between ChatGPT-4's clusters and independent expert sortings of the same items; the paper does not report such a measurement.
- A stronger version of the central claim would require showing that items the model groups semantically also correlate in actual user ratings, a link the paper explicitly does not assert.
- The same seven-prompt chain may transfer to other item pools or neighboring domains such as service or content evaluation, but the paper provides no evidence of transfer.
- Whether the classifications become a shared research standard depends on run-to-run stability, which the paper notes is not guaranteed.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper explores the use of ChatGPT-4 to identify common UX factors by semantically clustering measurement items from established UX questionnaires. The authors collect 408 items from 19 questionnaires (after excluding semantic differential and divergent item formats), run a sequence of seven prompts in a single ChatGPT-4 session, and report the resulting topic hierarchies. They also demonstrate a use case of filtering items that match the UX concept of Learnability. The paper claims that ChatGPT-4 can classify items into meaningful topics and filter items for predefined concepts, based on the authors' visual inspection of the outputs.
Significance. The problem addressed—the lack of semantic common ground across UX questionnaires—is real and relevant to the HCI measurement community. Applying a modern LLM to item-level semantic analysis is a timely and potentially useful idea, and the paper provides detailed appendices with example outputs, which is helpful for transparency. However, the central claims rest entirely on subjective inspection without any quantitative validation, inter-rater reliability, baseline comparison, or stability analysis. If the claims were appropriately validated, the approach could offer a low-cost tool for exploratory semantic structuring of item pools. At present, the evidence is insufficient to support the strong claims made in the abstract.
major comments (5)
- [V-G / Appendix A4] The claim that ChatGPT-4 can filter items matching the predefined UX concept of Learnability is contradicted by the paper's own appendix. Of the 15 top-ranked items, only items 1, 11, 12, 13, and 14 directly address ease of learning or understanding how to use the system; items 2, 3, and 10 measure effectiveness or efficiency, items 5 and 6 measure error handling and recovery, items 7, 8, and 9 measure information clarity and findability, and item 4 measures comfort. The authors state that these items 'fitted quite well,' but the appendix does not support this. Since no precision/recall against human expert labels or the original questionnaire factor assignments is provided, this load-bearing demonstration of filtering capability is not established.
- [V (all subsections) / VI-A] The central claim that ChatGPT-4 can classify items into 'meaningful topics' is evaluated only through the authors' informal visual inspection of the generated topics and example items. No inter-rater reliability, no comparison against human expert labels, no quantitative accuracy measure, and no baseline (e.g., random clustering or a simpler embedding-based method) are reported. Because the abstract advertises the classification result as the main finding, the evaluation must include some formal validation, such as agreement statistics with human raters or recovery of the original questionnaire factor structure.
- [IV (prompt5) / V-E / V-F] The prompt sequence injects the authors' previously published 16 UX quality aspects (Table I, from reference [6], co-authored by the third author) into the ChatGPT-4 session via prompt5 before the generalization step in prompt6. This can prime the model to produce outputs aligned with that particular list, so the later 'good alignment' between AI-generated topics and existing UX concepts is not a neutral finding. The paper should include a control condition in which prompt6 is run without the prior insertion of the 16 factors, or it should explicitly discuss this priming risk and its implications.
- [I / VI-B] The abstract and introduction state that items from '40 established UX questionnaires' were analyzed, but the actual analysis includes only 19 questionnaires after excluding semantic differentials and divergent measurement concepts. This scoping is described in Section IV and in the limitations, but the abstract should be revised to reflect that the findings apply to the subset of 19 questionnaires and 408 items, otherwise the generalizability claim is overstated.
- [VI-A] The paper acknowledges that LLMs are non-deterministic and that repeated runs yield different classifications, but all results are based on a single OpenAI session. No stability analysis (e.g., repeated runs with different temperatures or seeds) is reported, so the reader cannot assess whether the presented topics are robust or arbitrary. The paper should report multiple runs and provide a measure of output stability, or temper the claims accordingly.
minor comments (5)
- [Abstract] Please qualify the statement about '40 established UX questionnaires' to indicate that 19 questionnaires with 408 items were analyzed after exclusions.
- [Table II] The right-hand column of Table II contains stray text ('AI-generated topics.', 'Design') and the rows for Novelty and Identity are incomplete; the table should be cleaned and completed.
- [V-B] The phrase 'more precious' should be 'more precise'.
- [V-F / Appendix A3] The symbols (+), (-), and (+-) used to annotate item fits in Appendix A3 are not defined in the main text; please define them in Section V-F or in the appendix.
- [IV] The paper does not report the ChatGPT-4 model version, temperature, or other sampling parameters; please provide these details for reproducibility.
Circularity Check
Alignment with 'existing UX concepts' is partly self-confirming: the authors' own 16-factor list from [6] is fed to ChatGPT in prompt5 before the final generalization, then reported as good alignment.
-
self citation load bearing
[Table I (Section III), Section IV prompts 5-6, Section V.E-F, Section VI.A]
"we consulted existing UX concepts (see Table I) developed by [6] and compared them to the AI-generated categories ... prompt5: "In literature, I can find such a list with 16 UX factors. —inserted the defined quality aspects (see Table I) —. Can you compare this list with your categorization and contrast these lists?" ... "Against this, we added a further prompt 'I would like you to take your categorization you have done earlier and improve this into more generalized, holistic topics'.""
The 'existing UX concepts' used as the comparison benchmark are Table I from [6], a paper co-authored by the present third author (Schrepp). This list is inserted into the ChatGPT session in prompt5, immediately before prompt6 asks the model to produce its final generalized topics. The later claim in Section VI.A that 'The AI-generated topics indicate a good alignment compared to existing UX concepts' is therefore not an independent confirmation against an external taxonomy: the model was shown the target framework and then asked to generalize in its presence. The alignment is partly manufactured by the prompt design, although the paper does report mismatches (e.g., Novelty and Identity are not covered), so the circularity is partial rather than total.
full rationale
The core exploratory classification (prompts 1-4) is not circular in the derivation-chain sense: ChatGPT-4 receives 408 item texts and produces topic lists, and the paper reports those lists. Whether the resulting topics are 'meaningful' is a subjective-validity question, not a case of output being identical to input by construction. The circularity enters at the comparison stage. The paper validates its AI-generated structure by 'good alignment' with the consolidated UX factors of Table I, but that table is taken from [6], a paper co-authored by the third author, and the same list is inserted into the ChatGPT session in prompt5 immediately before prompt6 requests the final generalized topics. Thus the alignment reported in Section VI.A is partly an artifact of feeding the target framework to the model; the model was not tested blind against an independent external taxonomy. The paper itself concedes partial mismatch, so this is not a fully forced result. The Learnability-filter claim (prompt7) is validated only by the authors' visual inspection of the model's own output, without inter-rater agreement or precision/recall against source-questionnaire factor labels; that is a correctness and validation weakness (and Appendix A4 arguably contradicts the claim), but it is not a self-citation or definitional circularity in the strict sense. Overall score 4: some self-citation is load-bearing for the alignment claim, while the central classification demonstration retains independent content.
Assumptions & free parameters
assumptions (4)
- domain assumption The 40 UX questionnaires listed in [7] are the established set for UX measurement.
- domain assumption The exclusion of semantic differentials and concrete measurement concepts yields a comparable and analyzable set of 408 items.
- domain assumption Semantic similarity as computed by ChatGPT-4 is a valid proxy for the conceptual similarity of UX items.
- ad hoc to paper The authors' qualitative evaluation is sufficient evidence of the 'meaningfulness' of the generated topics.
Cite this review
Pith. "Pith review of Using ChatGPT-4 for the Identification of Common UX Factors within a Pool of Measurement Items from Established UX Questionnaires." pith.science (2026). https://pith.science/paper/IRPAZPUH
@misc{pith2026241113118,
author = {Pith},
title = {Pith review of: Using ChatGPT-4 for the Identification of Common UX Factors within a Pool of Measurement Items from Established UX Questionnaires},
year = {2026},
howpublished = {\url{https://pith.science/paper/IRPAZPUH}},
note = {Machine review of arXiv:2411.13118}
}
read the original abstract
Measuring User Experience (UX) with standardized questionnaires is a widely used method. A questionnaire is based on different scales that represent UX factors and items. However, the questionnaires have no common ground concerning naming different factors and the items used to measure them. This study aims to identify general UX factors based on the formulation of the measurement items. Items from a set of 40 established UX questionnaires were analyzed by Generative AI (GenAI) to identify semantically similar items and to cluster similar topics. We used the LLM ChatGPT-4 for this analysis. Results show that ChatGPT-4 can classify items into meaningful topics and thus help to create a deeper understanding of the structure of the UX research field. In addition, we show that ChatGPT-4 can filter items related to a predefined UX concept out of a pool of UX items.
Reference graph
Works this paper leans on
-
[6]
On the importance of ux quality aspects for different product categories,
M. Schrepp et al. , “On the importance of ux quality aspects for different product categories,” International Journal of In - teractive Multimedia and Artificial Intelligence, vol. In Press, pp. 232–246, Jun. 2023. DOI: 10.9781/ijimai.2023.03.001
-
[1]
I. O. for Standardization 9241-210:2019, Ergonomics of human-system interaction — Part 210: Human-centred design for interactive systems . ISO - International Organization for Standardization, 2019
work page 2019
-
[2]
Efficient measurement of the user expe - rience of interactive products. how to use the user experience questionnaire (ueq).example: Spanish language version,
M. Rauschenberger, M. Schrepp, M. P. Cota, S. Olschner, and J. Thomaschewski, “Efficient measurement of the user expe - rience of interactive products. how to use the user experience questionnaire (ueq).example: Spanish language version,” Int. J. Interact. Multim. Artif. Intell., vol. 2, pp. 39–45, 2013
2013
-
[3]
W. B. Albert and T. T. Tullis, Measuring the User Experience. Collecting, Analyzing, and Presenting UX Metrics . Morgan Kaufmann, 2022
2022
-
[4]
Standardized usability questionnaires: Features and quality focus,
A. Assila, K. M. de Oliveira, and H. Ezzedine, “Standardized usability questionnaires: Features and quality focus,” Computer Science and Information Technology, vol. 6, pp. 15–31, 2016
2016
-
[5]
Applicability of user experience and usability questionnaires,
A. Hinderks, D. Winter, M. Schrepp, and J. Thomaschewski, “Applicability of user experience and usability questionnaires,” J. Univers. Comput. Sci., vol. 25, pp. 1717–1735, 2019
2019
-
[7]
A comparison of ux questionnaires - what is their underlying concept of user experience?
M. Schrepp, “A comparison of ux questionnaires - what is their underlying concept of user experience?” In Mensch und Computer 2020 - Workshopband, C. Hansen, A. Nu¨rnberger, and B. Preim, Eds., Bonn: Gesellschaft fu¨r Informatik e.V.,
work page 2020
-
[8]
From usability to user experience,
H. M. Hassan and G. H. Galal-Edeen, “From usability to user experience,” in 2017 International Conference on Intel - ligent Informatics and Biomedical Sciences (ICIIBMS), 2017, pp. 216–222. DOI: 10.1109/ICIIBMS.2017.8279761
arXiv 2017
Show all 48 references
-
[9]
The thing and i: Understanding the relation - ship between user and product,
M. Hassenzahl, “The thing and i: Understanding the relation - ship between user and product,” in Funology: From Usability to Enjoyment, M. A. Blythe, K. Overbeeke, A. F. Monk, and P. C. Wright, Eds. Dordrecht: Springer Netherlands, 2004, pp. 31 –42, ISBN: 978 -1-4020-2967-7. D...
2004 doi
-
[10]
Mikolov, I
T. Mikolov, I. Sutskever, K. Chen, G. Corrado, and J. Dean, Distributed representations of words and phrases and their compositionality, retrieved: 10/2023, 2013. eprint: 1310.4546. [Online]. Available: https://arxiv.org/abs/1310.4546
2023 arXiv
-
[11]
Kenter, A
T. Kenter, A. Borisov, and M. de Rijke, Siamese cbow: Op - timizing word embeddings for sentence representations , 2016. eprint: 1606.04640
2016 arXiv
-
[12]
Conneau, D
A. Conneau, D. Kiela, H. Schwenk, L. Barrault, and A. Bordes, Supervised learning of universal sentence representations from natural language inference data, 2018. eprint: 1705.02364
2018 arXiv
-
[13]
Sentence similarity based on semantic nets and corpus statis - tics,
Y. Li, D. McLean, Z. A. Bandar, J. D. O’shea, and K. Crockett, “Sentence similarity based on semantic nets and corpus statis - tics,” IEEE transactions on knowledge and data engineering , vol. 18, no. 8, pp. 1138–1150, 2006
2006
-
[14]
A statistical approach to mechanized encoding and searching of literary information,
H. P. Luhn, “A statistical approach to mechanized encoding and searching of literary information,” IBM Journal of research and development, vol. 1, no. 4, pp. 309–317, 1957
1957
-
[15]
A statistical interpretation of term specificity and its application in retrieval,
K. Spa¨rck Jones, “A statistical interpretation of term specificity and its application in retrieval,” Journal of documentation , vol. 60, no. 5, pp. 493–502, 2004
2004
-
[16]
Okapi at trec-3,
S. E. Robertson, S. Walker, S. Jones, M. M. Hancock-Beaulieu, and M. Gatford, “Okapi at trec-3,” Nist Special Publication Sp, vol. 109, pp. 109–126, 1995
1995
-
[17]
Indexing by latent semantic analysis,
S. Deerwester, S. T. Dumais, G. W. Furnas, T. K. Landauer, and R. Harshman, “Indexing by latent semantic analysis,” Journal of the American society for information science, vol. 41, no. 6, pp. 391–407, 1990
1990
-
[18]
Distributed representations of sen - tences and documents,
Q. Le and T. Mikolov, “Distributed representations of sen - tences and documents,” in International conference on machine learning, PMLR, 2014, pp. 1188–1196
2014
-
[19]
Sentence-bert: Sentence embed- dings using siamese bert -networks,
N. Reimers and I. Gurevych, “Sentence-bert: Sentence embed- dings using siamese bert -networks,” in Conference on Empir - ical Methods in Natural Language Processing, 2019
2019
-
[20]
Augmented sbert: Data augmentation method for improving bi-encoders for pairwise sentence scoring tasks,
N. Thakur, N. Reimers, J. Daxenberger, and I. Gurevych, “Augmented sbert: Data augmentation method for improving bi-encoders for pairwise sentence scoring tasks,” arXiv preprint arXiv:2010.08240, Oct. 2020
2010 arXiv
-
[21]
Sentence Similarity Based on Contexts,
X. Sun et al. , “Sentence Similarity Based on Contexts,” Transactions of the Association for Computational Linguistics, vol. 10, pp. 573–588, 2022, ISSN: 2307-387X. DOI: 10.1162/ tacl a 00477
2022
-
[22]
Empirical simi - larity,
I. Gilboa, O. Lieberman, and D. Schmeidler, “Empirical simi - larity,” The Review of Economics and Statistics, vol. 88, no. 3, pp. 433–444, 2006
2006
-
[23]
What causes the dependency between perceived aesthetics and perceived usability?,
M. Schrepp, R. Otten, K. Blum, and J. Thomaschewski, “What causes the dependency between perceived aesthetics and perceived usability?,” pp. 78–85, 2021
2021
-
[24]
Apparent usability vs. inherent usability, chi’95 conference companion,
M. Kuroso and K. Kashimura, “Apparent usability vs. inherent usability, chi’95 conference companion,” in Conference on human factors in computing systems, Denver, Colorado, 1995, pp. 292–293
1995
-
[25]
Aesthetics and apparent usability: Empirically assessing cultural and methodological issues,
N. Tractinsky, “Aesthetics and apparent usability: Empirically assessing cultural and methodological issues,” in Proceedings of the ACM SIGCHI Conference on Human factors in comput - ing systems, 1997, pp. 115–122
1997
-
[26]
Cognitive processes causing the relationship between aesthetics and usability,
W. Ilmberger, M. Schrepp, and T. Held, “Cognitive processes causing the relationship between aesthetics and usability,” in HCI and Usability for Education and Work: 4th Sym - posium of the Workgroup Human -Computer Interaction and Usability Engineering of the Austrian Computer...
2008
-
[27]
The role of visual complexity and prototypi - cality regarding first impression of websites: Working towards understanding aesthetic judgments,
A. N. Tuch, E. E. Presslaber, M. Sto¨cklin, K. Opwis, and J. A. Bargas-Avila, “The role of visual complexity and prototypi - cality regarding first impression of websites: Working towards understanding aesthetic judgments,” International journal of human-computer studies, vol....
2012
-
[28]
A test of the context dependency of three causal models of halo rater error.,
C. E. Lance, J. A. LaPointe, and A. M. Stewart, “A test of the context dependency of three causal models of halo rater error.,” Journal of Applied Psychology , vol. 79, no. 3, pp. 332 –340, 1994
1994
-
[29]
Inferential beliefs in con- sumer evaluations: An assessment of alternative processing strategies,
G. T. Ford and R. A. Smith, “Inferential beliefs in con- sumer evaluations: An assessment of alternative processing strategies,” Journal of consumer research, vol. 14, no. 3, pp. 363–371, 1987
1987
-
[30]
D. A. Norman, Emotional design: Why we love (or hate) everyday things. Civitas Books, 2004
2004
-
[31]
Formalising guidelines for the design of screen layouts,
D. C. L. Ngo, L. S. Teo, and J. G. Byrne, “Formalising guidelines for the design of screen layouts,” Displays, vol. 21, no. 1, pp. 3–15, 2000
2000
-
[32]
A method of quantifying order in typographic design,
G. Bonsiepe, “A method of quantifying order in typographic design,” Visible Language, vol. 2, no. 3, pp. 203–220, 1968
1968
-
[33]
Stan - dardized questionnaires for user experience evaluation: A systematic literature review,
I. D´ıaz-Oreiro, G. Lo´pez, L. Quesada, and Guerrero, “Stan - dardized questionnaires for user experience evaluation: A systematic literature review,” Proceedings, vol. 31, pp. 14 –26, Nov. 2019. DOI: 10.3390/proceedings2019031014
2019 doi
-
[34]
Construction and evaluation of a user experience questionnaire,
B. Laugwitz, T. Held, and M. Schrepp, “Construction and evaluation of a user experience questionnaire,” in HCI and Usability for Education and Work , A. Holzinger, Ed., Berlin, Heidelberg: Springer Berlin Heidelberg, 2008, pp. 63 –76, ISBN: 978-3-540-89350-9
2008
-
[35]
Ueq user experience questionnaire,
T. UEQ, “Ueq user experience questionnaire,” 2018, retrieved: 10/2023. [Online]. Available: https://www.ueq-online.org/
2018
-
[36]
A proposed index of usability: A method for comparing the relative usability of different software systems,
H. X. Lin, Y.-Y. Choong, and G. Salvendy, “A proposed index of usability: A method for comparing the relative usability of different software systems,” Behaviour & information technol- ogy, vol. 16, no. 4-5, pp. 267–277, 1997
1997
-
[37]
Schrepp, User Experience Questionnaires: How to use questionnaires to measure the user experience of your prod - ucts? KDP, ISBN-13: 979-8736459766, 2021
M. Schrepp, User Experience Questionnaires: How to use questionnaires to measure the user experience of your prod - ucts? KDP, ISBN-13: 979-8736459766, 2021
2021
-
[38]
Design and validation of a framework for the creation of user experience questionnaires,
M. Schrepp and J. Thomaschewski, “Design and validation of a framework for the creation of user experience questionnaires,” International Journal of Interactive Multimedia and Artificial Intelligence, vol. InPress, pp. 88–95, Dec. 2019. DOI: 10.9781/ ijimai.2019.06.006
2019
-
[39]
Ueq+ a modular extension of the user experience questionnaire,
M. Schrepp, “Ueq+ a modular extension of the user experience questionnaire,” 2019, retrieved: 10/2023. [Online]. Available: http://www.ueqplus.ueq-research.org/
2019
-
[40]
Faktoren der user experience: Systematische u¨bersicht u¨ber produktrelevante ux-qualita¨tsaspekte,
D. Winter, M. Schrepp, and J. Thomaschewski, “Faktoren der user experience: Systematische u¨bersicht u¨ber produktrelevante ux-qualita¨tsaspekte,” in Workshop, A. Endmann, H. Fischer, and M. Kro¨kel, Eds. Berlin, Mu¨nchen, Boston: De Gruyter, 2015, pp. 33–41, ISBN: 97831104438...
2015
-
[41]
Quantifying user experience through self-reporting questionnaires: A systematic analysis of sentence similarity between the items of the measurement approaches,
S. Graser and S. Bo¨hm, “Quantifying user experience through self-reporting questionnaires: A systematic analysis of sentence similarity between the items of the measurement approaches,” in Lecture Notes in Computer Science, LNCS, volume 14014 , Springer Nature, 2023
2023
-
[42]
Applying augmented sbert and bertopic in ux research: A sentence similarity and topic model- ing approach to analyzing items from multiple questionnaires,
S. Graser and S. Bo¨hm, “Applying augmented sbert and bertopic in ux research: A sentence similarity and topic model- ing approach to analyzing items from multiple questionnaires,” in Proceedings of the IWEMB 2023, Seventh International Workshop on Entrepreneurship, Electronic...
2023
-
[43]
Bertopic: Neural topic modeling with a class -based tf -idf procedure,
M. Grootendorst, “Bertopic: Neural topic modeling with a class -based tf -idf procedure,” arXiv preprint arXiv:2203.05794, 2022
2022 arXiv
-
[44]
A comprehensive survey of ai -generated con- tent (aigc): A history of generative ai from gan to chatgpt,
Y. Cao et al., “A comprehensive survey of ai -generated con- tent (aigc): A history of generative ai from gan to chatgpt,” arXiv:2303.04226, pp. 1 –44, 2023, retrieved: 10/2023. [On - line]. Available: https://arxiv.org/abs/2303.04226
2023 arXiv
-
[45]
The power of generative ai: A review of requirements, models, inputndash;output formats, evaluation metrics, and challenges,
A. Bandi, P. V. S. R. Adapa, and Y. E. V. P. K. Kuchi, “The power of generative ai: A review of requirements, models, inputndash;output formats, evaluation metrics, and challenges,” Future Internet, vol. 15, no. 8, pp. 260 –320, 2023, ISSN: 1999-
2023
-
[46]
Gpt-4 technical report,
OpenAI, “Gpt-4 technical report,” ArXiv, vol. abs/2303.08774, 2023, retrieved: 10/2023. [Online]. Available: https://arxiv.org/ abs/2303.08774
2023 arXiv
-
[2020]
DOI: 10.18420/muc2020-ws105-236
-
[5903]
DOI: 10.3390/fi15080260
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.