{"id":"46b2fedd-a69e-4cf5-8e69-90cc0f4a520a","arxiv_id":"2411.13616","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"ChatGPT-4 can generate plausible semantic classifications of UX questionnaire items and reveal overlaps between UX concepts, though the results are judged only qualitatively.","lead":"The paper tests whether ChatGPT-4 can sort user experience questionnaire items into meaningful semantic groups and find overlapping concepts. It shows the model produces plausible categories and similarity maps, but without quantitative validation against human experts.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The claim that ChatGPT-4's semantic classifications are 'useful' for all three tasks is load-bearing but unmeasured: RQ1/RQ2 rest on single non-deterministic runs judged qualitatively by the authors, with no human-expert baseline or run-to-run stability check to rule out keyword matching or…","rationale":"The reader's weakest assumption—that ChatGPT's semantic judgments are a valid proxy for human semantic similarity—is indeed load-bearing, and I agree with the CONDITIONAL verdict. My stress-test highlights a second, partly independent condition: reproducibility. The paper openly acknowledges non-determinism but uses it to excuse instability rather than measuring it. For RQ1 and RQ2, single runs are presented as representative; only RQ3 uses three runs and filters to items that were consistently assigned. If repeated runs produce materially different classifications, the specific topics and item lists reported are not robust, and any of many outputs could have been selected to support the narrative. The qualitative comparison to UEQ+ scales in Section VII-B is the strongest positive evidence, since it provides an external anchor, but it is not quantified: no overlap counts, no chance baseline, and several exceptions are acknowledged. The proposed concrete test combines the two concerns: it measures ChatGPT-expert agreement against expert-expert agreement and also measures run-to-run overlap. If ChatGPT's agreement with experts is near chance or its outputs are unstable, the central claim in Section VIII-A is unsupported; if it passes, the paper's modest conclusions would be substantially strengthened. This is an addressable weakness, so the verdict remains CONDITIONAL rather than REJECT. I partially agree with the reader because the validity/baseline issue was already identified, but the stability issue is made more central here than in the reader's weakest_assumption.","tokens_in":21076,"tokens_out":6614,"duration_ms":68390,"concrete_test":"Have three UX experts independently assign a random sample of 100 items from the 408-item pool to the 16 UX quality aspects of [13], and run each of prompts1-7 three times. Compute Cohen's kappa between ChatGPT's modal assignment and the expert majority, and the mean pairwise Jaccard overlap across ChatGPT runs. If ChatGPT-expert kappa is not at least as high as expert-expert kappa, or if run-to-run overlap is low, the Section VIII-A usefulness conclusion fails. This test directly quantifies the missing validity and stability evidence.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim in Section VIII-A ('applying ChatGPT was useful for conducting all three tasks. The three research questions can be confirmed') requires that ChatGPT-4's outputs are both semantically valid and at least minimally stable. Neither condition is measured. For RQ1 and RQ2, each prompt was run only once (Sections V-A and VI-A), although Section VIII-A concedes the model is non-deterministic; the first two investigations therefore present one arbitrary draw as evidence. The reported classifications and top-10 item lists are assessed by the authors' own qualitative reading (e.g., 'the detected items fit well'), with no inter-rater agreement, no comparison against expert judgment, and no quantitative fit metric. The external check against UEQ+ scales in Section VII-B is also qualitative ('remarkably close'), and no agreement counts are reported. Because the prompts contain the target words (e.g., 'efficiently' in the Efficiency prompt), the observed matches could be explained by lexical cueing rather than semantic understanding. Thus the load-bearing condition—that ChatGPT's semantic assignments track expert semantic similarity closely enough to be useful—is asserted, not demonstrated.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper explores whether ChatGPT-4 can be used to analyze semantic similarities among UX questionnaire items. In the first investigation, ChatGPT-4 is asked to (re)construct UX factors by classifying 408 items from 19 established questionnaires into topics and subtopics through six successively refined prompts. In the second, a generic prompt is adapted to filter items representing predefined UX concepts (Learnability, Efficiency, Usefulness, Dependability, Stimulation) from the same item pool. In the third, 135 artificial standardized items of the form \"I perceive the product as X\" are used to map semantic connections among 11 common UX concepts, with each prompt run three times and only items consistently assigned in all three runs retained. The authors report that the generated classifications and item assignments are plausible and align well with existing UX knowledge, concluding in Section VIII-A that \"applying ChatGPT was useful for conducting all three tasks\" and that the three research questions can be confirmed.","tokens_in":21252,"tokens_out":3563,"duration_ms":35941,"significance":"If the central claim were quantitatively supported, the paper would provide a low-cost, scalable method for exploring the semantic structure of UX questionnaires and for supporting questionnaire selection and ad-hoc item generation. The third investigation is the most methodologically promising part: the artificial standardization of items, the three-run consistency filter, and the explicit comparison with empirically constructed UEQ+ scales are sensible design choices. The authors also clearly articulate the distinction between semantic and empirical similarity in Section II-B, which is an important conceptual contribution. However, the paper currently offers no quantitative validation of the core assertion that ChatGPT-4's outputs are meaningful, and it therefore reads more as a demonstration than as a validated research result.","major_comments":[{"comment":"The central conclusion that \"applying ChatGPT was useful for conducting all three tasks\" and that RQ1 and RQ2 are \"confirmed\" is not supported by the evidence presented. In Sections V and VI, each prompt was run only once, and the evaluations (e.g., \"plausible\" in V-B1, \"fit well\" in VI-B, and the discussion of misclassifications) are the authors' qualitative judgments. No inter-rater reliability, no comparison against expert human raters, and no quantitative metric such as precision, recall, or agreement are reported. Since the authors themselves note in Section VIII-A that ChatGPT-4 is non-deterministic, a single run cannot establish that the observed outputs reflect stable semantic competence rather than one arbitrary draw.","section":"Section VIII-A, with Sections V and VI"},{"comment":"The item-filtering results are vulnerable to lexical cueing. The prompt for Efficiency explicitly contains the word \"efficiently,\" and the concept descriptions for the other dimensions contain the target words (e.g., \"control,\" \"secure,\" \"stimulating\"). The reported top-10 lists include items that contain these exact words or near-synonyms; for example, item 7 under Efficiency, \"The processing times of the software are easy for me to estimate,\" contains \"processing times\" and is acknowledged by the authors to be a Dependability item. Without a control condition (e.g., prompts that avoid using the target words) or a human baseline, the results do not allow the reader to distinguish semantic understanding from keyword matching, which is a load-bearing issue for RQ2.","section":"Section VI-B"},{"comment":"The external validation against empirically constructed UEQ+ scales is qualitative only. The statement that the correspondence is \"remarkably close\" is not accompanied by any counts, confusion matrix, or statistical measure. The three-run consistency filter in Section VII-A is a useful reliability step, but it documents within-model consistency under a fixed prompting condition; it does not measure agreement with expert judgment or with empirically derived scale assignments, so it cannot by itself establish the validity of the semantic assignments.","section":"Section VII-B"}],"minor_comments":[{"comment":"The phrase \"it is more precious\" should presumably be \"it is more precise.\"","section":"Section V-B2"},{"comment":"There is a typo in the Stimulation topic: \"I continued to use thr application out of curiosity\" should read \"the application.\"","section":"Appendix A3"},{"comment":"The list of the 19 included questionnaires is not provided, so the reader cannot verify the exclusion criteria or reproduce the 408-item pool; adding an appendix with the questionnaire names and item list would improve reproducibility.","section":"Section IV"},{"comment":"Figure 3 is difficult to read because of the small font size and dense connection lines; a supplementary table listing the shared items for each pair of UX concepts would improve clarity.","section":"Figure 3"},{"comment":"The transformation of statement items into adjectives (\"we removed all other parts of the item and kept only the positive adjective\") is described informally; a more explicit rule for handling negations and multi-word phrases would reduce ambiguity.","section":"Section VII"},{"comment":"The claim that LLMs \"use word embeddings ... to calculate semantic similarity\" is stated as a fact about ChatGPT-4's internal mechanism; since that mechanism is not directly observable, the wording should be more cautious.","section":"Section II-B"}],"recommendation":"major_revision","confidential_remarks":"The paper is an exploratory demonstration of using ChatGPT-4 for semantic structuring of UX items. The main gap is the absence of any quantitative validation or human-expert baseline for the central claim. A revision that adds a human-rater study on a subset of items, multiple runs for the RQ1 and RQ2 prompts, and a quantitative measure of agreement with existing UX knowledge would substantially strengthen the contribution. The paper's third investigation is the most promising and could serve as the core of a revised, more focused manuscript."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Stefan, short take: this is a modest, honest, exploratory paper about using ChatGPT-4 to structure UX questionnaire items. What's genuinely new are the item-filtering use case (Section VI) and the adjective-based semantic similarity graph (Section VII). The graph is a nice visual artifact, and the authors are careful to say the output depends on prompt wording and LLM version. The external comparison to empirically built UEQ+ scales is a good idea, and the limitations are stated clearly.\n\nWhat the paper does well: It explicitly distinguishes semantic similarity from empirical similarity (Section II-B) and uses that to explain why a semantic map won't just reproduce empirically derived scales. That framing keeps the claims appropriately modest. Also, the three-run consistency filter in Section VII is a real, if small, improvement over single-run prompting.\n\nThe soft spots are where the reader's stress-test lands. For RQ1 and RQ2, each prompt was run once, and the results are assessed by the authors' own reading ('fit well', 'plausible'). No inter-rater agreement, no human-expert baseline, no quantitative fit metric. The prompts contain the target words, so lexical cueing can explain some matches. The UEQ+ comparison is qualitative ('remarkably close') with no agreement counts. In RQ3, the artificial items are a transformation (adjectives extracted), so the graph says something about the transformed item set, not the original questionnaires. None of these flaws are hidden; the paper concedes non-determinism and the need for human review. But the central conclusion that ChatGPT-4 was 'useful' for all three tasks is asserted on the strength of selected examples.\n\nThis is an exploratory methods paper, not a theory breakthrough. For UX researchers wanting a cheap way to map semantic overlap between questionnaire items or concepts, it's a legitimate pilot and a reasonable starting point. I would send it to a serious referee, but with a clear request for a human expert baseline, agreement metrics, and the item pools and outputs published. The revision path is feasible and would turn 'plausible' into 'shown'.\n\nRecommendation: engage with it as a borderline revise-and-resubmit. Don't desk-reject.","headline":"A useful extension of the authors' prior ChatGPT-4 UX work, but the load-bearing claim that the classifications are 'useful' rests on qualitative inspection and needs a human-rater baseline to be a real result.","tokens_in":21809,"tokens_out":4897,"would_cite":false,"duration_ms":45029,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that ChatGPT-4 can reconstruct UX factors, filter questionnaire items for predefined UX concepts, and uncover semantic overlaps between UX concepts, and that all three research questions are confirmed.","keywords":["user experience","UX questionnaires","semantic similarity","ChatGPT-4","large language models","semantic textual similarity","UX factors","generative AI"],"falsifier":"Give the same 408-item and 135-adjective pools to several UX experts, have them independently produce the classifications, item lists, and concept assignments, and compute agreement with ChatGPT-4's outputs; low agreement would show the model's judgments do not track expert semantics. A simpler stability check would be to rerun the prompts with reworded formulations and see whether the concept-overlap network, including the Aesthetics–Value overlap and the Clarity-mediated link to usability, survives.","tokens_in":20843,"feed_emoji":"🤖","tokens_out":7118,"duration_ms":64376,"temperature":0.7,"pith_summary":"This paper asks whether ChatGPT-4 can take over the labor-intensive job of understanding the semantic structure of UX questionnaires. On a pool of 408 items from 19 established questionnaires, it runs six prompts to reconstruct UX factors, a generic prompt adapted per concept to filter items for Learnability, Efficiency, Usefulness, Dependability, and Stimulation, and a second pool of 135 standardized \"I perceive the product as ...\" items to map overlaps between eleven common UX concepts. The authors report that all three investigations produced plausible, useful results and confirm the three research questions. If the claim holds, researchers could explore semantic dependencies across questionnaires quickly and cheaply, without manually comparing hundreds of item formulations.","feed_headline":"ChatGPT-4 groups UX questionnaire items into meaningful topics","feed_subtitle":"ChatGPT-4 also filters items for chosen UX concepts and maps the overlaps between them, the authors report.","key_machinery":"The machinery is a four-step procedure in which ChatGPT-4 supplies every semantic judgment. Step one collects 408 items from 19 of 40 established UX questionnaires, excluding semantic differentials and items tied to a specific interface or product type. Step two applies six successive prompts to the first item pool, moving from a broad classification to detailed subtopics, an improved categorization, a comparison with the 16 consolidated UX quality aspects, and finally a generalized holistic topic structure. Step three uses a generic prompt with a replaceable concept definition to filter the best-matching items from the same pool. Step four builds a second pool of 135 artificial items of the form \"I perceive the product as <adjective>\", asks ChatGPT-4 in three separate runs which adjectives belong to each of eleven UX concepts, and keeps only adjectives assigned consistently across all runs; the resulting concept–adjective links form the semantic similarity network. The named outputs are the classification hierarchies, the filtered item lists, and the overlap network.","core_discovery":"The central claim is that a large language model, ChatGPT-4, can perform meaningful semantic analysis of UX questionnaire items, and that this is useful for UX research. In the first investigation, iterative prompting turned 408 items into a hierarchy of topics and subtopics that the authors judge largely coherent with a published consolidation of 16 UX quality aspects, although some hedonic aspects such as Novelty and Identity are not well captured. In the second, prompt-based filtering retrieved top-10 items for predefined concepts, with strong alignment for classical usability concepts and weaker alignment for Stimulation, which the authors attribute to the usability-oriented item pool. In the third, adjectives standardized into \"I perceive the product as X\" statements produced a concept-connection network that reproduces known empirical findings, including the strong Aesthetics–Value overlap and an indirect link from Aesthetics to usability through Clarity. The paper concludes that applying ChatGPT-4 was useful for all three tasks and that the three research questions are confirmed.","pith_inferences":["The authors do not measure run-to-run reliability; a direct extension is to run each prompt several times and report agreement across runs as a stability score.","The absence of a human-rater baseline means the paper's positive conclusion rests on the model's agreement with expert semantics; a testable extension is to compare ChatGPT-4's assignments with expert sortings of the same items using an agreement index such as Cohen's kappa.","The standardized adjective format could be exported to other construct domains, for example trust in automation, to map concept overlaps there; that extrapolation goes beyond what the paper claims.","The Clarity-mediated link between Aesthetics and usability, visible in the concept network, yields a concrete hypothesis for user studies: systematically varying layout clarity should shift perceived usability more than perceived beauty."],"forward_implications":["UX researchers could use ChatGPT-4 as a quick first pass to map the semantic landscape of a large item pool before selecting or constructing a questionnaire.","Ad-hoc surveys could be assembled by retrieving existing, validated items from a pool instead of writing new ones, with human review reserved for the few misfits the model produces.","The semantic concept network provides a bridge between semantic and empirical views of UX, since known empirical correlations such as the aesthetics–usability dependency reappear in the purely semantic map.","The near match between semantic assignments and the empirically constructed UEQ/UEQ+ scales suggests that a questionnaire built on semantic similarity alone is a realistic future artifact.","Because the method is cheap and quick, classifications can be run repeatedly and treated as explorative hypotheses that guide, rather than replace, expert judgment."],"supporting_citations":[{"why":"The prior ChatGPT-4 approach this paper extends; supplies the prompting-based method for identifying UX factors.","marker":"[1]"},{"why":"Source of the 40 established UX questionnaires from which the 19-questionnaire, 408-item pool is derived.","marker":"[9]"},{"why":"Supplies the consolidated 16 UX quality aspects used as the comparison target in prompt5 and as the concepts in investigations two and three.","marker":"[13]"},{"why":"Provides the empirical finding on aesthetics–usability dependency mediated by Clarity that the semantic network is checked against.","marker":"[27]"},{"why":"Documents the construction of the UEQ scales, the empirical scales used to assess how close semantic assignments come to empirically built scales.","marker":"[39]"},{"why":"Describes the UEQ+ modular framework and its scale-construction method, another empirical benchmark for the semantic assignments.","marker":"[44]"},{"why":"The GPT-4 technical report that documents the model whose judgments are being tested.","marker":"[52]"}],"fun_headline_variants":["ChatGPT-4 maps UX questionnaire items into themes","AI sorts UX survey items into meaningful groups","ChatGPT-4 reveals links between UX concepts","LLM clusters UX items and filters for concepts","ChatGPT-4 decodes semantics of UX questionnaires"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper assumes that ChatGPT-4's semantic judgments are a valid stand-in for how UX experts understand the meaning of questionnaire items, without measuring agreement against expert human ratings.","fun_headline_variants_meta":{"raw":{"variants":["ChatGPT-4 maps UX questionnaire items into themes","AI sorts UX survey items into meaningful groups","ChatGPT-4 reveals links between UX concepts","LLM clusters UX items and filters for concepts","ChatGPT-4 decodes semantics of UX questionnaires"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000168,"raw_usage":{"total_tokens":1277,"prompt_tokens":979,"completion_tokens":298,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":595,"completion_tokens_details":{"reasoning_tokens":226}},"tokens_in":595,"tokens_out":298,"duration_ms":3785,"temperature":1.0,"reasoning_tokens":226,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T16:48:13.282409+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Give the same 408-item and 135-adjective pools to several UX experts, have them independently produce the classifications, item lists, and concept assignments, and compute agreement with ChatGPT-4's outputs; low agreement would show the model's judgments do not track expert semantics. A simpler stability check would be to rerun the prompts with reworded formulations and see whether the concept-overlap network, including the Aesthetics–Value overlap and the Clarity-mediated link to usability, survives.","supporting_citations":[{"cited_title":"Using ChatGPT -4 for the identification of common ux factors within a pool of measurement items from established ux questionnaires,","cited_arxiv_id":null,"evidence_quote":"The prior ChatGPT-4 approach this paper extends; supplies the prompting-based method for identifying UX factors."},{"cited_title":"A comparison of UX questionnaires - what is their underlying concept of user experience?","cited_arxiv_id":null,"evidence_quote":"Source of the 40 established UX questionnaires from which the 19-questionnaire, 408-item pool is derived."},{"cited_title":"On the importance of UX quality aspects for different product categories,","cited_arxiv_id":null,"evidence_quote":"Supplies the consolidated 16 UX quality aspects used as the comparison target in prompt5 and as the concepts in investigations two and three."},{"cited_title":"What causes the dependency between perceived aesthetics and perceived usability?,","cited_arxiv_id":null,"evidence_quote":"Provides the empirical finding on aesthetics–usability dependency mediated by Clarity that the semantic network is checked against."},{"cited_title":"Construction and evaluation of a user experience questionnaire,","cited_arxiv_id":null,"evidence_quote":"Documents the construction of the UEQ scales, the empirical scales used to assess how close semantic assignments come to empirically built scales."},{"cited_title":"Design and validation of a framework for the creation of user experience questionnaires,","cited_arxiv_id":null,"evidence_quote":"Describes the UEQ+ modular framework and its scale-construction method, another empirical benchmark for the semantic assignments."}],"review_version":1}