{"id":"d83c3817-077b-4930-a8fc-c93cb916ffa5","arxiv_id":"2605.29456","paper_version":2,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":5.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"Multimodal LLMs applied to 16 real-world configurators using 18 synthesized criteria can identify usability issues and generate actionable suggestions, with human review confirming reliability.","lead":"The paper tests whether multimodal LLMs can scan screenshots of real configurator interfaces and rate them against 18 literature-derived usability criteria, producing severity scores and improvement suggestions. A human review of the outputs found the AI mostly reliable, suggesting this could cut the manual effort needed to check configurator usability.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.3","headline":"Qualitative human review of MLLM outputs provides no objective measure of reliability or criterion completeness","rationale":"The reader's weakest_assumption directly identifies the same methodological gap. With full text available the concern remains unchanged: the paper contains no stronger evidence than the qualitative review already flagged. No machine-checked elements, parameter-free derivations, or falsifiable quantitative predictions appear that would alter this assessment.","tokens_in":1652,"tokens_out":353,"duration_ms":12008,"concrete_test":"Recruit 3 independent usability experts; have them apply the exact 18 criteria to the same 16 configurator screenshots and compute (a) pairwise agreement with each other and (b) agreement with the MLLM outputs using Cohen's kappa on severity ratings plus overlap on identified issues; if mean kappa < 0.6 or expert-identified issues exceed MLLM issues by >30% on average, the reliability claim is unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim rests on the statement that 'a review of the results confirms that MLLMs can reliably identify...' (abstract and §4/§5). This review is described as qualitative inspection of generated severity ratings and suggestions for 16 configurators against the 18 synthesized criteria. No quantitative validation is reported: no expert baseline, no inter-rater agreement (Cohen/Fleiss kappa), no precision/recall against ground-truth issues, and no ablation on criterion overlap or omission. The 18 criteria themselves are synthesized from literature without reported coverage analysis or redundancy check. Because 'reliably' is asserted solely via this subjective pass, the claim reduces to an untested assertion that the authors' inspection found the outputs plausible.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper proposes using multimodal large language models (MLLMs) for semi-automated usability analysis of configurator user interfaces. It synthesizes 18 configurator-specific criteria from the literature, applies them via MLLMs to evaluate 16 real-world configurators (producing per-criterion severity ratings and improvement suggestions), and asserts that a qualitative review of the outputs confirms the MLLMs can reliably detect issues and offer domain-aware recommendations (while still requiring human validation).","tokens_in":1809,"tokens_out":465,"duration_ms":17406,"significance":"If the reliability claim were supported by objective evidence, the work would represent a useful empirical demonstration of applying existing MLLMs to a specialized software-engineering domain where general heuristics are insufficient. The synthesis of 18 domain criteria and the concrete application to 16 configurators could lower the effort for systematic UI analysis, provided the criteria are shown to be non-redundant and the MLLM outputs are validated against expert baselines.","major_comments":[{"comment":"Abstract, §4, and §5: The assertion that MLLMs 'can reliably identify configurator-specific usability issues' rests entirely on an unspecified qualitative review of the generated severity ratings and suggestions. No quantitative validation is reported (no inter-rater agreement metrics such as Cohen/Fleiss kappa, no precision/recall against ground-truth issues identified by experts, no ablation on criterion overlap/omission, and no baseline comparison). Because this review is the sole support for the central claim of reliability, the claim is unsupported by objective evidence.","section":"Abstract, §4, and §5"}],"minor_comments":[{"comment":"The paper should clarify the exact procedure used for the 'review of the results' (who performed it, how many configurators were inspected in detail, and what rubric was applied) to allow readers to assess its scope.","section":"§5"},{"comment":"Details on the literature synthesis process for the 18 criteria (search strategy, inclusion criteria, and any redundancy or coverage checks) would strengthen the methodological transparency.","section":"§3"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive feedback. We agree that the reliability claim in the abstract, §4, and §5 rests on qualitative review alone and lacks objective quantitative support, which is a limitation of the current exploratory study. We will revise the manuscript accordingly.","responses":[{"response":"We acknowledge that the central claim relies solely on an author-conducted qualitative review of the MLLM outputs without quantitative metrics, inter-rater agreement, expert baselines, or ablation studies. This is a genuine limitation of the work, which is positioned as an initial demonstration rather than a validated method. In the revision we will: (1) tone down the language in the abstract, §4, and §5 to state that the outputs 'appeared consistent with domain expectations upon qualitative review' instead of claiming 'reliability'; (2) add an explicit description of the review process (authors with configurator expertise examined a sample of outputs for plausibility and actionability); and (3) expand the limitations and future-work sections to highlight the absence of quantitative validation and the need for expert ground-truth comparisons. We cannot retroactively add the requested quantitative analyses to this study but will treat the point as a clear direction for follow-on research.","revision_made":"yes","referee_comment":"[Abstract, §4, and §5] Abstract, §4, and §5: The assertion that MLLMs 'can reliably identify configurator-specific usability issues' rests entirely on an unspecified qualitative review of the generated severity ratings and suggestions. No quantitative validation is reported (no inter-rater agreement metrics such as Cohen/Fleiss kappa, no precision/recall against ground-truth issues identified by experts, no ablation on criterion overlap/omission, and no baseline comparison). Because this review is the sole support for the central claim of reliability, the claim is unsupported by objective evidence."}],"tokens_in":1302,"tokens_out":402,"duration_ms":21506,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"This paper takes 18 usability criteria synthesized from prior work on configurators and feeds screenshots from 16 real-world examples into multimodal LLMs. For each criterion the models produce a severity rating plus improvement suggestions. The authors state that their review of those outputs confirms the models identify issues reliably and give domain-aware advice.\n\nThe concrete application to configurators is the main new element. Configurators have particular demands around option dependencies and visual feedback that general heuristics often miss, and the paper shows a workable per-criterion workflow on actual interfaces.\n\nThe limitation is the validation. The reliability claim rests on the authors saying they inspected the generated ratings and suggestions and found them plausible. No inter-rater numbers, no expert baseline comparison, no count of missed or over-reported issues, and no check on whether the 18 criteria overlap or leave gaps. That makes the central assertion hard to weigh from the text alone.\n\nThe rest of the work is an empirical demonstration that follows existing criteria and prompting patterns. The criteria list and the example outputs could still be useful as a reference for people who build or maintain configurators in e-commerce or product customization.\n\nI would send it to peer review. The idea is practical and the examples are already collected; referees can push for quantitative checks without requiring a full rewrite.","headline":"They run MLLMs on 16 configurator screenshots against 18 literature-derived criteria and conclude from their own review that the outputs are reliable.","tokens_in":2351,"tokens_out":339,"would_cite":false,"duration_ms":19684,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Multimodal large language models can reliably detect configurator-specific usability issues and suggest improvements when using 18 domain criteria.","keywords":["configurator usability","multimodal large language models","user interface analysis","usability criteria","software configuration","semi-automated usability evaluation","domain-specific heuristics"],"falsifier":"Independent expert evaluations of the same 16 configurators using the 18 criteria, measuring the level of agreement with MLLM severity ratings and recommendation quality.","tokens_in":2570,"feed_emoji":"🤖","tokens_out":560,"duration_ms":20670,"temperature":0.7,"pith_summary":"This paper examines whether multimodal large language models can perform usability analysis on configurator user interfaces. It synthesizes 18 specific criteria from existing literature and tests the approach on 16 real-world configurators. For each criterion, the models assign severity levels to issues and propose fixes. The results indicate that these models produce assessments that align with human judgment in identifying problems and offering relevant recommendations. This method could make evaluating and improving configurator usability more efficient by handling much of the initial analysis automatically.","feed_headline":"MLLMs detect configurator usability issues reliably","feed_subtitle":"Evaluation of 16 real-world examples with 18 criteria shows domain-aware severity ratings and improvement suggestions, reducing manual effor","key_machinery":"The set of 18 configurator-specific usability criteria, each evaluated separately by the MLLM to produce severity ratings and suggestions.","core_discovery":"By applying 18 configurator-specific usability criteria to screenshots or descriptions of 16 real-world configurators, multimodal large language models generate individual severity ratings and actionable improvement recommendations for each criterion, and a subsequent review finds these outputs to be reliable and domain-aware.","pith_inferences":["Integration with existing UI design tools could automate parts of the review process.","Expanding the criteria list might cover additional aspects of configurator interaction.","Similar techniques could apply to usability analysis in other specialized software domains.","Larger-scale studies with more configurators would test consistency across different MLLM versions."],"forward_implications":["Analysis effort for configurator usability decreases because MLLMs handle initial assessments.","Improvement suggestions are tailored to the configurator domain.","The approach works across multiple real-world examples.","Human oversight is still required for final validation."],"fun_headline_variants":["MLLMs evaluate configurator usability in 16 examples","18 criteria used by MLLMs on real configurator UIs","Reliable MLLM ratings for configurator interface issues","MLLMs suggest improvements for configurator UI problems","Analysis of configurators via multimodal LLMs and criteria"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"A qualitative human review of the MLLM outputs sufficiently establishes reliability, and the 18 criteria capture configurator usability without major omissions or overlaps.","fun_headline_variants_meta":{"raw":{"variants":["MLLMs evaluate configurator usability in 16 examples","18 criteria used by MLLMs on real configurator UIs","Reliable MLLM ratings for configurator interface issues","MLLMs suggest improvements for configurator UI problems","Analysis of configurators via multimodal LLMs and criteria"]},"model":"grok-4.3","cost_usd":0.006914,"raw_usage":{"total_tokens":3169,"prompt_tokens":592,"num_sources_used":0,"completion_tokens":79,"cost_in_usd_ticks":69137000,"prompt_tokens_details":{"text_tokens":592,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":2498,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":592,"tokens_out":79,"duration_ms":18133,"temperature":1.0,"reasoning_tokens":2498,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-29T06:55:13.473243+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Independent expert evaluations of the same 16 configurators using the 18 criteria, measuring the level of agreement with MLLM severity ratings and recommendation quality.","supporting_citations":[],"review_version":1}