Pith. sign in

REVIEW 3 major objections 6 minor 61 references

Identifying Semantic Similarity for UX Items from Established Questionnaires Using ChatGPT-4

T0 review · 3 major / 6 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read This paper claims that ChatGPT-4 can reconstruct UX factors, filter questionnaire items for predefined UX concepts, and uncover semantic overlaps between UX concepts, and that all three research questions are confirmed.

desk verdict A useful extension of the authors' prior ChatGPT-4 UX work, but the load-bearing claim that the classifications are 'useful' rests on qualitative inspection and needs a human-rater baseline to be a real result. read the letter →

arxiv 2411.13616 v1 pith:MZJCXKTG submitted 2024-11-20 cs.HC

classification cs.HC
keywords userexperienceUXquestionnairessemanticsimilarityChatGPT-4largelanguagemodelstextualfactorsgenerativeAI
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper asks whether ChatGPT-4 can take over the labor-intensive job of understanding the semantic structure of UX questionnaires. On a pool of 408 items from 19 established questionnaires, it runs six prompts to reconstruct UX factors, a generic prompt adapted per concept to filter items for Learnability, Efficiency, Usefulness, Dependability, and Stimulation, and a second pool of 135 standardized "I perceive the product as ..." items to map overlaps between eleven common UX concepts. The authors report that all three investigations produced plausible, useful results and confirm the three research questions. If the claim holds, researchers could explore semantic dependencies across questionnaires quickly and cheaply, without manually comparing hundreds of item formulations.

What carries the argument

The machinery is a four-step procedure in which ChatGPT-4 supplies every semantic judgment. Step one collects 408 items from 19 of 40 established UX questionnaires, excluding semantic differentials and items tied to a specific interface or product type. Step two applies six successive prompts to the first item pool, moving from a broad classification to detailed subtopics, an improved categorization, a comparison with the 16 consolidated UX quality aspects, and finally a generalized holistic topic structure. Step three uses a generic prompt with a replaceable concept definition to filter the best-matching items from the same pool. Step four builds a second pool of 135 artificial items of the form "I perceive the product as <adjective>", asks ChatGPT-4 in three separate runs which adjectives belong to each of eleven UX concepts, and keeps only adjectives assigned consistently across all runs; the resulting concept–adjective links form the semantic similarity network. The named outputs are the classification hierarchies, the filtered item lists, and the overlap network.

What would settle it

Give the same 408-item and 135-adjective pools to several UX experts, have them independently produce the classifications, item lists, and concept assignments, and compute agreement with ChatGPT-4's outputs; low agreement would show the model's judgments do not track expert semantics. A simpler stability check would be to rerun the prompts with reworded formulations and see whether the concept-overlap network, including the Aesthetics–Value overlap and the Clarity-mediated link to usability, survives.

Watch

Extended reading notes

Core claim

The central claim is that a large language model, ChatGPT-4, can perform meaningful semantic analysis of UX questionnaire items, and that this is useful for UX research. In the first investigation, iterative prompting turned 408 items into a hierarchy of topics and subtopics that the authors judge largely coherent with a published consolidation of 16 UX quality aspects, although some hedonic aspects such as Novelty and Identity are not well captured. In the second, prompt-based filtering retrieved top-10 items for predefined concepts, with strong alignment for classical usability concepts and weaker alignment for Stimulation, which the authors attribute to the usability-oriented item pool. In the third, adjectives standardized into "I perceive the product as X" statements produced a concept-connection network that reproduces known empirical findings, including the strong Aesthetics–Value overlap and an indirect link from Aesthetics to usability through Clarity. The paper concludes that applying ChatGPT-4 was useful for all three tasks and that the three research questions are confirmed.

Load-bearing premise

The paper assumes that ChatGPT-4's semantic judgments are a valid stand-in for how UX experts understand the meaning of questionnaire items, without measuring agreement against expert human ratings.

Editorial extensions

If this is right

  • UX researchers could use ChatGPT-4 as a quick first pass to map the semantic landscape of a large item pool before selecting or constructing a questionnaire.
  • Ad-hoc surveys could be assembled by retrieving existing, validated items from a pool instead of writing new ones, with human review reserved for the few misfits the model produces.
  • The semantic concept network provides a bridge between semantic and empirical views of UX, since known empirical correlations such as the aesthetics–usability dependency reappear in the purely semantic map.
  • The near match between semantic assignments and the empirically constructed UEQ/UEQ+ scales suggests that a questionnaire built on semantic similarity alone is a realistic future artifact.
  • Because the method is cheap and quick, classifications can be run repeatedly and treated as explorative hypotheses that guide, rather than replace, expert judgment.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The authors do not measure run-to-run reliability; a direct extension is to run each prompt several times and report agreement across runs as a stability score.
  • The absence of a human-rater baseline means the paper's positive conclusion rests on the model's agreement with expert semantics; a testable extension is to compare ChatGPT-4's assignments with expert sortings of the same items using an agreement index such as Cohen's kappa.
  • The standardized adjective format could be exported to other construct domains, for example trust in automation, to map concept overlaps there; that extrapolation goes beyond what the paper claims.
  • The Clarity-mediated link between Aesthetics and usability, visible in the concept network, yields a concrete hypothesis for user studies: systematically varying layout clarity should shift perceived usability more than perceived beauty.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper explores whether ChatGPT-4 can be used to analyze semantic similarities among UX questionnaire items. In the first investigation, ChatGPT-4 is asked to (re)construct UX factors by classifying 408 items from 19 established questionnaires into topics and subtopics through six successively refined prompts. In the second, a generic prompt is adapted to filter items representing predefined UX concepts (Learnability, Efficiency, Usefulness, Dependability, Stimulation) from the same item pool. In the third, 135 artificial standardized items of the form "I perceive the product as X" are used to map semantic connections among 11 common UX concepts, with each prompt run three times and only items consistently assigned in all three runs retained. The authors report that the generated classifications and item assignments are plausible and align well with existing UX knowledge, concluding in Section VIII-A that "applying ChatGPT was useful for conducting all three tasks" and that the three research questions can be confirmed.

Significance. If the central claim were quantitatively supported, the paper would provide a low-cost, scalable method for exploring the semantic structure of UX questionnaires and for supporting questionnaire selection and ad-hoc item generation. The third investigation is the most methodologically promising part: the artificial standardization of items, the three-run consistency filter, and the explicit comparison with empirically constructed UEQ+ scales are sensible design choices. The authors also clearly articulate the distinction between semantic and empirical similarity in Section II-B, which is an important conceptual contribution. However, the paper currently offers no quantitative validation of the core assertion that ChatGPT-4's outputs are meaningful, and it therefore reads more as a demonstration than as a validated research result.

major comments (3)
  1. [Section VIII-A, with Sections V and VI] The central conclusion that "applying ChatGPT was useful for conducting all three tasks" and that RQ1 and RQ2 are "confirmed" is not supported by the evidence presented. In Sections V and VI, each prompt was run only once, and the evaluations (e.g., "plausible" in V-B1, "fit well" in VI-B, and the discussion of misclassifications) are the authors' qualitative judgments. No inter-rater reliability, no comparison against expert human raters, and no quantitative metric such as precision, recall, or agreement are reported. Since the authors themselves note in Section VIII-A that ChatGPT-4 is non-deterministic, a single run cannot establish that the observed outputs reflect stable semantic competence rather than one arbitrary draw.
  2. [Section VI-B] The item-filtering results are vulnerable to lexical cueing. The prompt for Efficiency explicitly contains the word "efficiently," and the concept descriptions for the other dimensions contain the target words (e.g., "control," "secure," "stimulating"). The reported top-10 lists include items that contain these exact words or near-synonyms; for example, item 7 under Efficiency, "The processing times of the software are easy for me to estimate," contains "processing times" and is acknowledged by the authors to be a Dependability item. Without a control condition (e.g., prompts that avoid using the target words) or a human baseline, the results do not allow the reader to distinguish semantic understanding from keyword matching, which is a load-bearing issue for RQ2.
  3. [Section VII-B] The external validation against empirically constructed UEQ+ scales is qualitative only. The statement that the correspondence is "remarkably close" is not accompanied by any counts, confusion matrix, or statistical measure. The three-run consistency filter in Section VII-A is a useful reliability step, but it documents within-model consistency under a fixed prompting condition; it does not measure agreement with expert judgment or with empirically derived scale assignments, so it cannot by itself establish the validity of the semantic assignments.
minor comments (6)
  1. [Section V-B2] The phrase "it is more precious" should presumably be "it is more precise."
  2. [Appendix A3] There is a typo in the Stimulation topic: "I continued to use thr application out of curiosity" should read "the application."
  3. [Section IV] The list of the 19 included questionnaires is not provided, so the reader cannot verify the exclusion criteria or reproduce the 408-item pool; adding an appendix with the questionnaire names and item list would improve reproducibility.
  4. [Figure 3] Figure 3 is difficult to read because of the small font size and dense connection lines; a supplementary table listing the shared items for each pair of UX concepts would improve clarity.
  5. [Section VII] The transformation of statement items into adjectives ("we removed all other parts of the item and kept only the positive adjective") is described informally; a more explicit rule for handling negations and multi-word phrases would reduce ambiguity.
  6. [Section II-B] The claim that LLMs "use word embeddings ... to calculate semantic similarity" is stated as a fact about ChatGPT-4's internal mechanism; since that mechanism is not directly observable, the wording should be more cautious.

Circularity Check

1 steps flagged · score 2.0 of 10

No central circularity; one minor self-referential validation occurs when the authors' own consolidated UX factor list is fed into prompt5 and then used as the benchmark to declare the AI-generated topics logical.

  1. self citation load bearing [Section V-B-5 ('Prompt5: Comparison Towards Existing Consolidation'), Table II and its concluding paragraph]
    "In particular, we inserted the existing UX quality aspects and formulated the prompt as follows: 'In literature, I can find such a list with 16 UX factors. —inserted the defined quality aspects (see Table I) [13] —. Can you compare this list with your categorization and contrast these lists?' ... Nevertheless, there are many similarities between the two consolidations, and thus, the AI-generated topics by ChatGPT can be considered logical."

    The reference standard used to declare the AI-generated topics 'logical' is the consolidated UX factor list from [13], a prior work sharing an author with the present paper, and this same list is inserted into the ChatGPT prompt immediately before the comparison. The model is asked to relate its earlier categorization to a list supplied by the experimenters, so the observed alignment is partly a consequence of the prompt content rather than an independent rediscovery of the factors. This makes the prompt5-based validation self-referential. The circularity is limited because prompts1-4 and prompt6 generate categories without receiving the factor list, and RQ2 and RQ3 rely on other checks, so the central derivation is not forced by this step.

full rationale

The paper's main derivation chain is not circular. For RQ1, prompts1-4 and prompt6 ask ChatGPT to classify 408 item texts into semantic topics without providing the target factor names; the resulting topics are therefore not defined by the inputs. For RQ2, the generic prompt describes a UX concept and asks the model to select fitting items from the same raw item pool, and the paper assesses the selections against expert expectations—this is an empirical use-case, not a fitted parameter renamed as a prediction. For RQ3, the artificial items are constructed from adjectives in existing questionnaires and the model is prompted with concept descriptions; the final check against empirically built UEQ+ scales is an external benchmark, even though the item pool is derived from related questionnaires. No equation-level reduction, fitted input, or uniqueness import is present. The one notable self-referential element is the prompt5 comparison: the authors' own consolidated factor list from [13] is fed into the prompt and then used as the reference for saying the AI-generated topics are 'logical.' This is a minor self-citation that partially contaminates one validation comparison, but it is not load-bearing for the overall conclusion that GenAI is useful for the three tasks, which also rests on qualitative inspection and the external UEQ+ comparison. Concerns about non-determinism and the lack of a human-rater baseline are substantive validity issues, but they are not circularity.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The paper contributes a qualitative demonstration rather than a derivation; no fitted parameters or new theoretical entities are introduced. The load-bearing assumptions are domain assumptions about the validity of LLM judgments, the representativeness of the item pools, and the completeness of the chosen concept list.

assumptions (4)
  • domain assumption LLM semantic judgments are a valid proxy for human semantic similarity of UX items.
    The whole method assumes ChatGPT-4's clustering and filtering reflect meaningful semantic relations; the paper only checks plausibility by inspection, not against human raters.
  • domain assumption The selected item pool (408 items from 19 questionnaires) and the 135 standardized adjectives are representative of the UX questionnaire space.
    Section IV; exclusions of semantic differentials and product-specific items shape results, and the 135 items are artificially constructed from positive adjectives only.
  • domain assumption The 11 UX concepts used as a reference frame (Learnability, Efficiency, etc.) are the relevant dimensions for UX.
    These concepts come from the authors' own prior work [13] and UEQ+ [44]; the semantic overlap analysis cannot detect relations outside this frame.
  • domain assumption Considering an adjective related to a concept only if it was assigned in all three ChatGPT runs is a sufficient reliability rule.
    Section VII-A; this threshold reduces noise but is a heuristic with no justification.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Identifying Semantic Similarity for UX Items from Established Questionnaires Using ChatGPT-4." pith.science (2026). https://pith.science/paper/MZJCXKTG

@misc{pith2026241113616,
  author       = {Pith},
  title        = {Pith review of: Identifying Semantic Similarity for UX Items from Established Questionnaires Using ChatGPT-4},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/MZJCXKTG}},
  note         = {Machine review of arXiv:2411.13616}
}
read the original abstract

Questionnaires are a widely used tool for measuring the user experience (UX) of products. There exists a huge number of such questionnaires that contain different items (questions) and scales representing distinct aspects of UX, such as efficiency, learnability, fun of use, or aesthetics. These items and scales are not independent; they often have semantic overlap. However, due to the large number of available items and scales in the UX f ield, analyzing and understanding these semantic dependencies can be challenging. Large language models (LLM) are powerful tools to categorize texts, including UX items. We explore how ChatGPT-4 can be utilized to analyze the semantic structure of sets of UX items. This paper investigates three different use cases. In the first investigation, ChatGPT-4 is used to generate a semantic classification of UX items extracted from 40 UX questionnaires. The results demonstrate that ChatGPT-4 can effectively classify items into meaningful topics. The second investigation demonstrates ChatGPT-4's ability to filter items related to a predefined UX concept from a pool of UX items. In the third investigation, a second set of more abstract items is used to describe another classification task. The outcome of this investigation helps to determine semantic similarities between common UX concepts and enhances our understanding of the concept of UX. Overall, it is considered useful to apply GenAI in UX research

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

61 extracted references · 46 canonical work pages

  1. [1]

    Using ChatGPT -4 for the identification of common ux factors within a pool of measurement items from established ux questionnaires,

    S. Graser, S. Bo¨hm, and M. Schrepp, “Using ChatGPT -4 for the identification of common ux factors within a pool of measurement items from established ux questionnaires,” in CENTRIC 2023: The Sixteenth International Conference on Advances in Human -oriented and Personalized Mechanisms, Technologies, and Services, 2023, pp. 19–28

  2. [2]

    I. O. for Standardization 9241 -210:2019, Ergonomics of human-system interaction — Part 210: Human-centred design for interactive systems . ISO - International Organization for Standardization, 2019

  3. [3]

    Efficient measurement of the user expe - rience of interactive products. how to use the user experience questionnaire (ueq).example: Spanish language version,

    M. Rauschenberger, M. Schrepp, M. P. Cota, S. Olschner, and J. Thomaschewski, “Efficient measurement of the user expe - rience of interactive products. how to use the user experience questionnaire (ueq).example: Spanish language version,” Int. J. Interact. Multim. Artif. Intell., vol. 2, pp. 39–45, 2013

  4. [4]

    W. B. Albert and T. T. Tullis, Measuring the User Experience. Collecting, Analyzing, and Presenting UX Metrics . Morgan Kaufmann, 2022

  5. [5]

    Standardized usability questionnaires: Features and quality focus,

    A. Assila, K. M. de Oliveira, and H. Ezzedine, “Standardized usability questionnaires: Features and quality focus,” Computer Science and Information Technology, vol. 6, pp. 15–31, 2016

  6. [6]

    Applicability of user experience and usability questionnaires,

    A. Hinderks, D. Winter, M. Schrepp, and J. Thomaschewski, “Applicability of user experience and usability questionnaires,” J. Univers. Comput. Sci., vol. 25, pp. 1717–1735, 2019

  7. [7]

    A review of post-study and post- task subjective questionnaires to guide assessment of system usability,

    A. Hodrien and T. Fernando, “A review of post-study and post- task subjective questionnaires to guide assessment of system usability,” Journal of Usability Studies, vol. 16(3), no. 3, pp. 203–232, 2021

  8. [8]

    Schrepp, User Experience Questionnaires: How to use questionnaires to measure the user experience of your prod - ucts? KDP, ISBN-13: 979-8736459766, 2021

    M. Schrepp, User Experience Questionnaires: How to use questionnaires to measure the user experience of your prod - ucts? KDP, ISBN-13: 979-8736459766, 2021

Show all 61 references
  1. [9]

    A comparison of UX questionnaires - what is their underlying concept of user experience?

    M. Schrepp, “A comparison of UX questionnaires - what is their underlying concept of user experience?” In Mensch und Computer 2020 - Workshopband, C. Hansen, A. Nu¨rnberger, and B. Preim, Eds., Bonn: Gesellschaft fu¨r Informatik e.V.,

  2. [10]

    From usability to user experience,

    H. M. Hassan and G. H. Galal-Edeen, “From usability to user experience,” in 2017 International Conference on Intel - ligent Informatics and Biomedical Sciences (ICIIBMS), 2017, pp. 216–222. DOI: 10.1109/ICIIBMS.2017.8279761

  3. [11]

    Preece, Y

    J. Preece, Y. Rogers, and H. Sharp, Interaction Design: Beyond Human-Computer Interaction . Wiley John + Sons, ISBN -13 978-1119020752, 2015

  4. [12]

    The thing and I: Understanding the relation - ship between user and product,

    M. Hassenzahl, “The thing and I: Understanding the relation - ship between user and product,” in Funology: From Usability to Enjoyment, M. A. Blythe, K. Overbeeke, A. F. Monk, and P. C. Wright, Eds. Dordrecht: Springer Netherlands, 2004, pp. 31 –42, ISBN: 978 -1-4020-2967-7. D...

  5. [13]

    On the importance of UX quality aspects for different product categories,

    M. Schrepp et al., “On the importance of UX quality aspects for different product categories,” International Journal of In - teractive Multimedia and Artificial Intelligence, vol. In Press, pp. 232–246, Jun. 2023. DOI: 10.9781/ijimai.2023.03.001

  6. [14]

    Mikolov, I

    T. Mikolov, I. Sutskever, K. Chen, G. Corrado, and J. Dean, Distributed representations of words and phrases and their compositionality, retrieved: 10/2023, 2013. eprint: 1310.4546. [Online]. Available: https://arxiv.org/abs/1310.4546

  7. [15]

    Kenter, A

    T. Kenter, A. Borisov, and M. de Rijke, Siamese cbow: Op - timizing word embeddings for sentence representations , 2016. eprint: 1606.04640

  8. [16]

    Conneau, D

    A. Conneau, D. Kiela, H. Schwenk, L. Barrault, and A. Bordes, Supervised learning of universal sentence representations from natural language inference data, 2018. eprint: 1705.02364

  9. [17]

    Sentence similarity based on semantic nets and corpus statis - tics,

    Y. Li, D. McLean, Z. A. Bandar, J. D. O’shea, and K. Crockett, “Sentence similarity based on semantic nets and corpus statis - tics,” IEEE Transactions on Knowledge and Data Engineering, vol. 18, no. 8, pp. 1138–1150, 2006

  10. [18]

    A statistical approach to mechanized encoding and searching of literary information,

    H. P. Luhn, “A statistical approach to mechanized encoding and searching of literary information,” IBM Journal of Re - search and Development, vol. 1, no. 4, pp. 309–317, 1957

  11. [19]

    A statistical interpretation of term specificity and its application in retrieval,

    K. Spa¨rck Jones, “A statistical interpretation of term specificity and its application in retrieval,” Journal of Documentation , vol. 60, no. 5, pp. 493–502, 2004

  12. [20]

    Okapi at trec-3,

    S. E. Robertson, S. Walker, S. Jones, M. M. Hancock-Beaulieu, and M. Gatford, “Okapi at trec-3,” Nist Special Publication Sp, vol. 109, pp. 109–126, 1995

  13. [21]

    Indexing by latent semantic analysis,

    S. Deerwester, S. T. Dumais, G. W. Furnas, T. K. Landauer, and R. Harshman, “Indexing by latent semantic analysis,” Journal of the American Society for Information Science, vol. 41, no. 6, pp. 391–407, 1990

  14. [22]

    Distributed representations of sen - tences and documents,

    Q. Le and T. Mikolov, “Distributed representations of sen - tences and documents,” in International Conference on Ma - chine Learning, PMLR, 2014, pp. 1188–1196

  15. [23]

    Sentence -BERT: Sentence em - beddings using siamese BERT -networks,

    N. Reimers and I. Gurevych, “Sentence -BERT: Sentence em - beddings using siamese BERT -networks,” in Conference on Empirical Methods in Natural Language Processing, 2019

  16. [24]

    Augmented SBERT: Data augmentation method for improv - ing bi -encoders for pairwise sentence scoring tasks,

    N. Thakur, N. Reimers, J. Daxenberger, and I. Gurevych, “Augmented SBERT: Data augmentation method for improv - ing bi -encoders for pairwise sentence scoring tasks,” arXiv preprint arXiv:2010.08240, Oct. 2020

  17. [25]

    Sentence similarity based on contexts,

    X. Sun et al., “Sentence similarity based on contexts,” Trans- actions of the Association for Computational Linguistics, vol. 10, pp. 573–588, 2022, ISSN: 2307-387X. DOI: 10.1162/ tacl a 00477

  18. [26]

    Empirical simi - larity,

    I. Gilboa, O. Lieberman, and D. Schmeidler, “Empirical simi - larity,” The Review of Economics and Statistics, vol. 88, no. 3, pp. 433–444, 2006

  19. [27]

    What causes the dependency between perceived aesthetics and perceived usability?,

    M. Schrepp, R. Otten, K. Blum, and J. Thomaschewski, “What causes the dependency between perceived aesthetics and perceived usability?,” pp. 78–85, 2021

  20. [28]

    Apparent usability vs. inherent usability, chi’95 conference companion,

    M. Kuroso and K. Kashimura, “Apparent usability vs. inherent usability, chi’95 conference companion,” in Conference on human factors in computing systems, Denver, Colorado, 1995, pp. 292–293

  21. [29]

    Aesthetics and apparent usability: Empirically assessing cultural and methodological issues,

    N. Tractinsky, “Aesthetics and apparent usability: Empirically assessing cultural and methodological issues,” in Proceedings of the ACM SIGCHI Conference on Human Factors in Com - puting Systems, 1997, pp. 115–122

  22. [30]

    Cognitive processes causing the relationship between aesthetics and usability,

    W. Ilmberger, M. Schrepp, and T. Held, “Cognitive processes causing the relationship between aesthetics and usability,” in HCI and Usability for Education and Work: 4th Sym - posium of the Workgroup Human -Computer Interaction and Usability Engineering of the Austrian Computer...

  23. [31]

    The role of visual complexity and prototypi - cality regarding first impression of websites: Working towards understanding aesthetic judgments,

    A. N. Tuch, E. E. Presslaber, M. Sto¨cklin, K. Opwis, and J. A. Bargas-Avila, “The role of visual complexity and prototypi - cality regarding first impression of websites: Working towards understanding aesthetic judgments,” International Journal of Human-Computer Studies, vol....

  24. [32]

    A test of the context dependency of three causal models of halo rater error.,

    C. E. Lance, J. A. LaPointe, and A. M. Stewart, “A test of the context dependency of three causal models of halo rater error.,” Journal of Applied Psychology , vol. 79, no. 3, pp. 332 –340, 1994

  25. [33]

    Inferential beliefs in con- sumer evaluations: An assessment of alternative processing strategies,

    G. T. Ford and R. A. Smith, “Inferential beliefs in con- sumer evaluations: An assessment of alternative processing strategies,” Journal of Consumer Research, vol. 14, no. 3, pp. 363–371, 1987

  26. [34]

    D. A. Norman, Emotional design: Why we love (or hate) everyday things. Civitas Books, 2004

  27. [35]

    Formalising guidelines for the design of screen layouts,

    D. C. L. Ngo, L. S. Teo, and J. G. Byrne, “Formalising guidelines for the design of screen layouts,” Displays, vol. 21, no. 1, pp. 3–15, 2000

  28. [36]

    A method of quantifying order in typographic design,

    G. Bonsiepe, “A method of quantifying order in typographic design,” Visible Language, vol. 2, no. 3, pp. 203–220, 1968

  29. [37]

    Facets of visual aesthetics.,

    M. Moshagen and M. Thielsch, “Facets of visual aesthetics.,” International Journal of Human -Computer Studies, 25 (13), 1717-1735., no. 68(10), pp. 689–709, 2010

  30. [38]

    Stan - dardized questionnaires for user experience evaluation: A systematic literature review,

    I. D´ıaz-Oreiro, G. Lo´pez, L. Quesada, and Guerrero, “Stan - dardized questionnaires for user experience evaluation: A systematic literature review,” Proceedings, vol. 31, pp. 14 –26, Nov. 2019. DOI: 10.3390/proceedings2019031014

  31. [39]

    Construction and evaluation of a user experience questionnaire,

    B. Laugwitz, T. Held, and M. Schrepp, “Construction and evaluation of a user experience questionnaire,” in HCI and Usability for Education and Work , A. Holzinger, Ed., Berlin, Heidelberg: Springer Berlin Heidelberg, 2008, pp. 63 –76, ISBN: 978-3-540-89350-9

  32. [40]

    UEQ User Experience Questionnaire,

    M. Schrepp, J. Thomaschewski, and A. Hinderks, “UEQ User Experience Questionnaire,” 2018, retrieved: 10/2023. [Online]. Available: https://www.ueq-online.org/

  33. [41]

    A proposed index of usability: A method for comparing the relative usability of different software systems,

    H. X. Lin, Y.-Y. Choong, and G. Salvendy, “A proposed index of usability: A method for comparing the relative usability of different software systems,” Behaviour & Information Tech - nology, vol. 16, no. 4-5, pp. 267–277, 1997

  34. [42]

    Sus: A “quick and dirty’usability,

    J. Brooke, “Sus: A “quick and dirty’usability,” Usability Eval- uation in Industry, vol. 189, no. 3, pp. 189–194, 1996

  35. [43]

    The isomet- rics usability inventory: An operationalization of iso 9241- 10 supporting summative and formative evaluation of software systems,

    G. Gediga, K.-C. Hamborg, and I. Du¨ntsch, “The isomet- rics usability inventory: An operationalization of iso 9241- 10 supporting summative and formative evaluation of software systems,” Behaviour & Information Technology, vol. 18, no. 3, pp. 151–164, 1999

  36. [44]

    Design and validation of a framework for the creation of user experience questionnaires,

    M. Schrepp and J. Thomaschewski, “Design and validation of a framework for the creation of user experience questionnaires,” International Journal of Interactive Multimedia and Artificial Intelligence, vol. InPress, pp. 88–95, Dec. 2019. DOI: 10.9781/ ijimai.2019.06.006

  37. [45]

    UEQ+ a modular exten - sion of the user experience questionnaire,

    M. Schrepp and J. Thomaschewski, “UEQ+ a modular exten - sion of the user experience questionnaire,” 2019, retrieved: 10/2023. [Online]. Available: http : / / www . ueqplus . ueq - research.org/

  38. [46]

    Faktoren der User Experience: Systematische U ¨ bersicht u¨ber produk- trelevante UX-Qualita¨tsaspekte,

    D. Winter, M. Schrepp, and J. Thomaschewski, “Faktoren der User Experience: Systematische U ¨ bersicht u¨ber produk- trelevante UX-Qualita¨tsaspekte,” in Workshop, A. Endmann, H. Fischer, and M. Kro¨kel, Eds. Berlin, Mu¨nchen, Boston: De Gruyter, 2015, pp. 33 –41, ISBN: 978311...

  39. [47]

    Quantifying user experience through self-reporting questionnaires: A systematic analysis of the sentence similarity between the items of the measurement ap - proaches,

    S. Graser and S. Bo¨hm, “Quantifying user experience through self-reporting questionnaires: A systematic analysis of the sentence similarity between the items of the measurement ap - proaches,” in HCI International 2023 – Late Breaking Posters, C. Stephanidis, M. Antona, S. Nt...

  40. [48]

    Applying augmented sbert and bertopic in ux research: A sentence similarity and topic model- ing approach to analyzing items from multiple questionnaires,

    S. Graser and S. Bo¨hm, “Applying augmented sbert and bertopic in ux research: A sentence similarity and topic model- ing approach to analyzing items from multiple questionnaires,” in Proceedings of the IWEMB 2023, Seventh International Workshop on Entrepreneurship, Electronic...

  41. [49]

    Bertopic: Neural topic modeling with a class -based tf -idf procedure,

    M. Grootendorst, “Bertopic: Neural topic modeling with a class -based tf -idf procedure,” arXiv preprint arXiv:2203.05794, 2022

  42. [50]

    A comprehensive survey of ai -generated con - tent (aigc): A history of generative ai from gan to chatgpt,

    Y. Cao et al., “A comprehensive survey of ai -generated con - tent (aigc): A history of generative ai from gan to chatgpt,” arXiv:2303.04226, pp. 1 –44, 2023, retrieved: 10/2023. [On - line]. Available: https://arxiv.org/abs/2303.04226

  43. [51]

    The power of generative ai: A review of requirements, models, inputndash;output formats, evaluation metrics, and challenges,

    A. Bandi, P. V. S. R. Adapa, and Y. E. V. P. K. Kuchi, “The power of generative ai: A review of requirements, models, inputndash;output formats, evaluation metrics, and challenges,” Future Internet, vol. 15, no. 8, pp. 260 –320, 2023, ISSN: 1999-

  44. [52]

    Gpt-4 technical report,

    OpenAI, “Gpt-4 technical report,” ArXiv, vol. abs/2303.08774, 2023, retrieved: 10/2023. [Online]. Available: https://arxiv.org/ abs/2303.08774

  45. [53]

    Large language models: A survey,

    S. Minaee et al. , “Large language models: A survey,” arXiv preprint arXiv:2402.06196, 2024

  46. [54]

    Language models are few-shot learners,

    T. B. Brown et al., “Language models are few-shot learners,” ArXiv, vol. abs/2005.14165, 2020

  47. [55]

    Training language models to follow in - structions with human feedback,

    L. Ouyang et al. , “Training language models to follow in - structions with human feedback,” ArXiv, vol. abs/2203.02155, 2022

  48. [56]

    Hugginggpt: Solving ai tasks with chatgpt and its friends in hugging face,

    Y. Shen, K. Song, X. Tan, D. Li, W. Lu, and Y. Zhuang, “Hugginggpt: Solving ai tasks with chatgpt and its friends in hugging face,” Advances in Neural Information Processing Systems, vol. 36, 2024

  49. [57]

    A survey on evaluation of large language models,

    Y. Chang et al. , “A survey on evaluation of large language models,” ACM Transactions on Intelligent Systems and Tech - nology, vol. 15, no. 3, pp. 1–45, 2024

  50. [58]

    Large language models can rate news outlet credibility,

    K.-C. Yang and F. Menczer, “Large language models can rate news outlet credibility,” ArXiv, vol. abs/2304.00228, 2023

  51. [59]

    UX Fragebo¨gen und Wort- wolken,

    B. Rummel and M. S. Martin, “UX Fragebo¨gen und Wort- wolken,” Mensch und Computer 2019-Workshopband, 2019

  52. [2020]

    DOI: 10.18420/muc2020-ws105-236

  53. [5903]

    DOI: 10.3390/fi15080260

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.