{"id":"4edf6a31-417c-4d66-969b-bfbb45215ae0","arxiv_id":"2411.13118","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"ChatGPT-4 produced plausible but unvalidated topic clusters from 408 UX questionnaire items, based on a qualitative demonstration without quantitative evaluation.","lead":"This study tested whether ChatGPT-4 could group 408 questionnaire items from 19 published UX surveys into shared topic clusters. The authors report that the model produced plausible categories such as Usability, Trust, and Aesthetics, and could filter items matching the concept of Learnability.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Appendix A4 contradicts the Learnability-filter claim: most of the top-15 items concern effectiveness, efficiency, error recovery, or information clarity rather than ease of learning.","rationale":"The reader's weakest assumption identified a general lack of quantitative validation behind the authors' visual inspection. My concern is more specific: the paper's own Appendix A4 provides a concrete example where that visual inspection is demonstrably wrong, because the top-15 items selected for Learnability include many items that measure effectiveness, efficiency, error recovery, or information clarity rather than ease of learning. This makes the filtering claim internally inconsistent, not merely unvalidated. However, the overall verdict remains CONDITIONAL because the paper's main classification result is exploratory and the filtering failure is fixable by re-evaluating prompt7 with formal metrics and expert judgment. The condition should be sharpened: the authors must provide quantitative evidence for the filtering claim or explicitly qualify/remove it. I agree partially with the reader: they correctly pointed to subjective visual inspection as the weak link, but the appendix shows the problem is not just missing rigor; the demonstrated output fails on its face. The proposed test would settle whether the 'good detection' statement is defensible.","tokens_in":14050,"tokens_out":7605,"duration_ms":77116,"concrete_test":"Have three independent UX researchers (not authors) label each of the 15 items in Appendix A4 as 'directly about ease of learning/understanding how to use the system' or 'not,' and compute precision@15 and Fleiss' kappa. Also map each item to the factor label of its original questionnaire (e.g., SUS Learnability, CSUQ, UEQ Perspicuity). If precision@15 is below 0.8, kappa is below 0.6, or fewer than 60% of the items originate from scales explicitly labeled as Learnability/Perspicuity/Intuitive Use, the 'good detection' claim fails and Section V-G's conclusion must be revised.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Appendix A4 is the only evidence for the paper's claim that ChatGPT-4 can filter items matching a predefined UX concept (Learnability). The authors rate this output as a 'good detection ... fitted quite well,' but direct inspection shows the opposite. Of the 15 items, only #1, #11, #12, and #14 are directly about ease of learning or understanding how to use the system; #13 is about familiar interaction (borderline intuitive use). Items #2 and #10 measure effectiveness, #3 measures efficiency, #5 and #6 measure error handling/recovery, #7 and #8 measure information clarity/findability, and #4 measures comfort. Thus a large fraction of the 'best' matches are off-target. Because the authors provide no inter-rater agreement, no precision/recall against the source questionnaires' factor labels, and no baseline, the filtering capability claim rests entirely on a visual inspection that the paper's own appendix refutes. Furthermore, prompt7 is run after prompt5 inserted the authors' 16 UX-factor list, so the model was not tested blind for the target concept. This is load-bearing because the abstract explicitly advertises the filtering result as a second demonstration of ChatGPT-4's utility.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper explores the use of ChatGPT-4 to identify common UX factors by semantically clustering measurement items from established UX questionnaires. The authors collect 408 items from 19 questionnaires (after excluding semantic differential and divergent item formats), run a sequence of seven prompts in a single ChatGPT-4 session, and report the resulting topic hierarchies. They also demonstrate a use case of filtering items that match the UX concept of Learnability. The paper claims that ChatGPT-4 can classify items into meaningful topics and filter items for predefined concepts, based on the authors' visual inspection of the outputs.","tokens_in":14267,"tokens_out":4034,"duration_ms":39308,"significance":"The problem addressed—the lack of semantic common ground across UX questionnaires—is real and relevant to the HCI measurement community. Applying a modern LLM to item-level semantic analysis is a timely and potentially useful idea, and the paper provides detailed appendices with example outputs, which is helpful for transparency. However, the central claims rest entirely on subjective inspection without any quantitative validation, inter-rater reliability, baseline comparison, or stability analysis. If the claims were appropriately validated, the approach could offer a low-cost tool for exploratory semantic structuring of item pools. At present, the evidence is insufficient to support the strong claims made in the abstract.","major_comments":[{"comment":"The claim that ChatGPT-4 can filter items matching the predefined UX concept of Learnability is contradicted by the paper's own appendix. Of the 15 top-ranked items, only items 1, 11, 12, 13, and 14 directly address ease of learning or understanding how to use the system; items 2, 3, and 10 measure effectiveness or efficiency, items 5 and 6 measure error handling and recovery, items 7, 8, and 9 measure information clarity and findability, and item 4 measures comfort. The authors state that these items 'fitted quite well,' but the appendix does not support this. Since no precision/recall against human expert labels or the original questionnaire factor assignments is provided, this load-bearing demonstration of filtering capability is not established.","section":"V-G / Appendix A4"},{"comment":"The central claim that ChatGPT-4 can classify items into 'meaningful topics' is evaluated only through the authors' informal visual inspection of the generated topics and example items. No inter-rater reliability, no comparison against human expert labels, no quantitative accuracy measure, and no baseline (e.g., random clustering or a simpler embedding-based method) are reported. Because the abstract advertises the classification result as the main finding, the evaluation must include some formal validation, such as agreement statistics with human raters or recovery of the original questionnaire factor structure.","section":"V (all subsections) / VI-A"},{"comment":"The prompt sequence injects the authors' previously published 16 UX quality aspects (Table I, from reference [6], co-authored by the third author) into the ChatGPT-4 session via prompt5 before the generalization step in prompt6. This can prime the model to produce outputs aligned with that particular list, so the later 'good alignment' between AI-generated topics and existing UX concepts is not a neutral finding. The paper should include a control condition in which prompt6 is run without the prior insertion of the 16 factors, or it should explicitly discuss this priming risk and its implications.","section":"IV (prompt5) / V-E / V-F"},{"comment":"The abstract and introduction state that items from '40 established UX questionnaires' were analyzed, but the actual analysis includes only 19 questionnaires after excluding semantic differentials and divergent measurement concepts. This scoping is described in Section IV and in the limitations, but the abstract should be revised to reflect that the findings apply to the subset of 19 questionnaires and 408 items, otherwise the generalizability claim is overstated.","section":"I / VI-B"},{"comment":"The paper acknowledges that LLMs are non-deterministic and that repeated runs yield different classifications, but all results are based on a single OpenAI session. No stability analysis (e.g., repeated runs with different temperatures or seeds) is reported, so the reader cannot assess whether the presented topics are robust or arbitrary. The paper should report multiple runs and provide a measure of output stability, or temper the claims accordingly.","section":"VI-A"}],"minor_comments":[{"comment":"Please qualify the statement about '40 established UX questionnaires' to indicate that 19 questionnaires with 408 items were analyzed after exclusions.","section":"Abstract"},{"comment":"The right-hand column of Table II contains stray text ('AI-generated topics.', 'Design') and the rows for Novelty and Identity are incomplete; the table should be cleaned and completed.","section":"Table II"},{"comment":"The phrase 'more precious' should be 'more precise'.","section":"V-B"},{"comment":"The symbols (+), (-), and (+-) used to annotate item fits in Appendix A3 are not defined in the main text; please define them in Section V-F or in the appendix.","section":"V-F / Appendix A3"},{"comment":"The paper does not report the ChatGPT-4 model version, temperature, or other sampling parameters; please provide these details for reproducibility.","section":"IV"}],"recommendation":"major_revision","confidential_remarks":"The paper is an exploratory proof-of-concept that would benefit from reframing as such. The injection of the authors' own 16-factor list into the prompt sequence, combined with the third author being a co-author of the cited framework, raises a methodological concern that should be addressed with a control condition or sensitivity analysis. The current abstract overstates the evidence. With added quantitative validation and tempered claims, the paper could make a useful contribution to the UX measurement community."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a modest exploratory case study, not a validated method. The authors show ChatGPT-4 can produce plausible-looking topic clusters from 408 UX questionnaire items, and they are transparent about their prompts and outputs. But the abstract's second claim — that ChatGPT-4 can filter items matching a predefined UX concept — is directly undermined by their own Appendix A4. Most of the top-15 \"Learnability\" items are about effectiveness, efficiency, error recovery, or information clarity, not ease of learning. So the \"good detection\" they report doesn't survive inspection.\n\nWhat's genuinely useful here: the prompt chain is fully documented, the appendices let readers see every assigned item, and the authors are upfront about the exploratory nature of the work. They also acknowledge the obvious limitation that LLM outputs are non-deterministic and that semantic differentials had to be excluded. That is more than many such papers do. Compared to their earlier sentence-embedding work, this is a new application, but it is a routine use of a general-purpose LLM — no new method.\n\nThe soft spots are real and load-bearing. The evaluation of topic quality is entirely subjective: no inter-rater agreement, no comparison against human expert sorting, no quantitative accuracy measure, no test–retest stability. There's also a contamination risk: prompt5 introduces the authors' own previously published 16 UX quality aspects into the session before prompt6 asks for \"generalized\" topics and before prompt7 runs the Learnability filter. That primes the model with their framework, so the later alignment is partly built in.\n\nThe paper's own evidence in Appendix A3 shows some classifications are loose — they mark several items with (+-) for partial fit, and the \"Novelty\" and \"Identity\" topics have only one clearly on-target item each. That's fine for an exploratory exercise, but it undercuts the conclusion that the topics are \"meaningful\" in any validated sense.\n\nFor whom: this is for a reader who wants a cheap, quick way to get a first-pass semantic structure over a large item pool, and who is willing to treat the output as a hypothesis generator, not a measurement result. That reader gets value from the documented prompt design and the assembled item list.\n\nBottom line: I would not cite this as evidence that LLM classification is reliable in UX research, but I would send it to a serious referee. The claim is modest, the transparency is good, and the weaknesses are addressable with inter-rater reliability, stability checks, and either removing the Learnability demonstration or reporting precision/recall against human labels. If an editor wants a publishable version, major revision is needed.","headline":"A transparent but under-evidenced demonstration that ChatGPT-4 can cluster UX items; the filtering claim is refuted by the paper's own appendix.","tokens_in":14764,"tokens_out":2657,"would_cite":false,"duration_ms":26377,"reading_group":"no","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"ChatGPT-4 can sort 408 UX questionnaire items into shared semantic topics, and can filter the pool for a concept like learnability.","keywords":["user experience","UX measurement","UX factors","semantic textual similarity","large language models","ChatGPT","questionnaire items","item clustering"],"falsifier":"Have a panel of UX experts independently sort the same 408 items into topics and compare their categories to ChatGPT-4's clusters; if agreement is near chance, or if repeated ChatGPT-4 runs with the same prompts produce wildly different taxonomies, the central claim that these are meaningful topics collapses. A complementary test is to check whether items the model groups together show high correlations in real survey responses, since the paper itself notes that semantic similarity and empirical similarity can diverge.","tokens_in":13847,"feed_emoji":"🧩","tokens_out":5341,"duration_ms":47461,"temperature":0.7,"pith_summary":"The paper aims to show that a large language model can read the wording of user-experience questionnaire items and group them into clusters that reflect common underlying UX factors. The authors fed 408 items from 19 established questionnaires to ChatGPT-4 through a chain of seven prompts, then inspected the resulting topics and item assignments. They conclude that the model produces meaningful, mostly coherent categories that overlap with an existing list of 16 UX quality aspects, and that it can retrieve items matching a predefined concept such as learnability. If this holds, UX researchers gain a fast and cheap way to compare items across questionnaires, build common-ground structures, and assemble ad-hoc surveys from existing item pools.","feed_headline":"ChatGPT-4 can group 408 UX items into shared topics","feed_subtitle":"A seven-prompt test shows the model's clusters align with known UX factors and can pick out learnability items.","key_machinery":"The central mechanism is ChatGPT-4's embedding-based semantic-textual-similarity ability, steered by a fixed sequence of seven prompts that move from broad classification to detailed subcategorization, self-improvement, comparison with the literature, generalization, and finally a concept-specific filter. This prompt chain converts the model's continuous similarity judgments into an explicit topic hierarchy plus an item-selection procedure.","core_discovery":"On the paper's own terms, the discovery is that ChatGPT-4, given only the texts of 408 measurement items, produces a two-level taxonomy of six main topics and 15 subtopics that aligns with an established consolidation of 16 UX quality aspects, and that the same model can act as an item finder for a specific UX concept. The authors report that the AI-generated topics capture both task-related and emotional aspects, that the functional topics are generated particularly well, and that some hedonic factors from the literature, such as Novelty and Identity, are poorly covered. They also report that asking the model to select items describing learnability returns a top-15 list whose entries match the concept well.","pith_inferences":["The plausibility claim could be made quantitative by measuring agreement between ChatGPT-4's clusters and independent expert sortings of the same items; the paper does not report such a measurement.","A stronger version of the central claim would require showing that items the model groups semantically also correlate in actual user ratings, a link the paper explicitly does not assert.","The same seven-prompt chain may transfer to other item pools or neighboring domains such as service or content evaluation, but the paper provides no evidence of transfer.","Whether the classifications become a shared research standard depends on run-to-run stability, which the paper notes is not guaranteed."],"forward_implications":["Researchers can cheaply explore the semantic structure of any large item pool without manual coding.","Items from different questionnaires can be compared on a common semantic basis, making it easier to select measures for ad-hoc surveys.","The AI-generated topics could serve as a scaffold for a new holistic UX questionnaire.","The filtering capability lets practitioners quickly reuse existing items for a target UX concept instead of writing new ones.","Because LLM output is non-deterministic, repeated runs can yield multiple alternative classifications, which the paper sees as an explorative advantage."],"supporting_citations":[{"why":"Supplies the set of 40 established UX questionnaires from which the 19 questionnaires and 408 items were drawn.","marker":"[7]"},{"why":"Provides the consolidated list of 16 UX quality aspects used as the comparison baseline in prompt5.","marker":"[6]"},{"why":"Prior work applying a Sentence Transformer to measure similarity between UX items; the approach this paper extends to an LLM.","marker":"[41]"},{"why":"Prior topic-modeling study on the same item class that the paper positions against.","marker":"[42]"},{"why":"Documents the GPT-4 model used as the analysis tool.","marker":"[46]"}],"fun_headline_variants":["ChatGPT-4 distills 408 UX items into six core topics","AI clusters 408 UX questionnaire items into shared factors","ChatGPT-4 aligns UX item groups with established quality aspects","LLM identifies common UX factors from 408 measurement items","ChatGPT-4 groups UX items into topics that match known factors"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper assumes that the authors' visual inspection of ChatGPT-4's output is enough to prove the topics are meaningful and the item fits correct, without any quantitative validation against human experts or empirical item correlations.","fun_headline_variants_meta":{"raw":{"variants":["ChatGPT-4 distills 408 UX items into six core topics","AI clusters 408 UX questionnaire items into shared factors","ChatGPT-4 aligns UX item groups with established quality aspects","LLM identifies common UX factors from 408 measurement items","ChatGPT-4 groups UX items into topics that match known factors"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000651,"raw_usage":{"total_tokens":2926,"prompt_tokens":828,"completion_tokens":2098,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":444,"completion_tokens_details":{"reasoning_tokens":2012}},"tokens_in":444,"tokens_out":2098,"duration_ms":16021,"temperature":1.0,"reasoning_tokens":2012,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T16:48:30.109632+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have a panel of UX experts independently sort the same 408 items into topics and compare their categories to ChatGPT-4's clusters; if agreement is near chance, or if repeated ChatGPT-4 runs with the same prompts produce wildly different taxonomies, the central claim that these are meaningful topics collapses. A complementary test is to check whether items the model groups together show high correlations in real survey responses, since the paper itself notes that semantic similarity and empirical similarity can diverge.","supporting_citations":[{"cited_title":"A comparison of ux questionnaires - what is their underlying concept of user experience?","cited_arxiv_id":null,"evidence_quote":"Supplies the set of 40 established UX questionnaires from which the 19 questionnaires and 408 items were drawn."},{"cited_title":"Quantifying user experience through self-reporting questionnaires: A systematic analysis of sentence similarity between the items of the measurement approaches,","cited_arxiv_id":null,"evidence_quote":"Prior work applying a Sentence Transformer to measure similarity between UX items; the approach this paper extends to an LLM."},{"cited_title":"Applying augmented sbert and bertopic in ux research: A sentence similarity and topic model- ing approach to analyzing items from multiple questionnaires,","cited_arxiv_id":null,"evidence_quote":"Prior topic-modeling study on the same item class that the paper positions against."}],"review_version":1}