{"id":"a604d1ce-e42e-4a0c-8f35-ef386aaae10d","arxiv_id":"2510.16435","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A new 1,893-question dataset and 12-category taxonomy reveal that household-robot users most want answers to safety- and error-related \"what if\" questions, while why-questions are rated lower in importance.","lead":"This paper contributes a dataset of 1,893 real user questions for household robots, collected from 100 participants and sorted into 12 categories and 70 subcategories. It also measures how important users think each type of question is, finding that safety- and error-related \"what if\" questions rank highest and why-questions rank near the bottom.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Internal count inconsistencies undermine the dataset-size claim: 2,052−143=1,909, Figure 3 category n's sum to 1,948, but the abstract/summary says 1,893; the public artifacts should be checked before the dataset is treated as canonical.","rationale":"I read the paper as a serious empirical contribution: the authors provide a public dataset, report inter-annotator agreement, use an appropriate mixed-effects model, and explicitly discuss limitations. The reader's identified weakest assumption — ecological validity — is real but is a standard limitation of stimulus-based elicitation studies and is acknowledged in Section 5.4. However, the more load-bearing and immediately checkable concern is internal numerical inconsistency. The paper gives three different totals for the final dataset: 1,893 (abstract/intro), 1,909 (2,052−143 in Section 4), and 1,948 (sum of Figure 3's n's, which match Figure 2's subcategory counts). If the true row count differs from 1,893, the central claim and all derived percentages and importance-model n's are uncertain. The test I propose — counting rows and category labels in the public data — would settle this definitively. If it confirms 1,893, the issue is a set of typos and the conditional verdict stands; if it confirms 1,909 or 1,948, the paper must be revised or rejected as a reliable resource. Since the reader already assigned CONDITIONAL, my concern does not move the verdict; it sharpens the condition: verify the dataset count and correct the internal discrepancies. I set agreement to 'partial' because I agree ecological validity is a limitation, but I do not regard it as the single most load-bearing issue; the count inconsistency is more fundamental and easier to resolve.","tokens_in":20618,"tokens_out":6054,"duration_ms":52048,"concrete_test":"Download the public dataset from https://github.com/lwachowiak/xai-questions-dataset (or the HuggingFace copy), count the number of question rows in the final cleaned CSV, and sum the primary-category labels by main category. Compare these counts to 1,893, 1,909, and 1,948. Then recompute the Section 4.2 mixed-effects model using the actual rows to verify the reported n's. If the row count is 1,893 and category n's sum to 1,893, Figure 3's n's and the abstract/body percentages need correction but the headline dataset size is preserved. If the row count is 1,948 or 1,909, the abstract, Section 4 exclusion arithmetic, or the dataset itself must be corrected before the central claim is accepted.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is that it introduces a dataset of 1,893 user questions, but multiple internal counts disagree. Section 4 states that 2,052 questions were collected and 143 excluded, which yields 1,909, not 1,893. More tellingly, the per-category sample sizes in Figure 3 — 120+190+126+245+83+157+52+417+106+209+125+118 — sum to 1,948, and these n's exactly match the subcategory counts in Figure 2. If each question is assigned to a single category, as the hierarchical coding scheme implies, then the categorized dataset should contain 1,948 rows. A further inconsistency appears in the abstract/body discrepancy for top-category percentages (21.4/12.6/10.7 vs. 22.5/12.7/11.3) and the number of text stimuli (abstract says 7, Section 3.2 says six). These are not isolated typos: three different dataset totals appear in the same paper. Because the dataset's reuse value depends on knowing exactly which questions are in it and how they are categorized, this count discrepancy is more immediately load-bearing than the acknowledged ecological-validity limitation. It is directly checkable from the public GitHub/HuggingFace artifacts, so it should be settled before the dataset is used for benchmarking or design guidance.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces a dataset of 1,893 user questions for household robots, elicited from 100 Prolific participants in response to 15 video and 7 text stimuli depicting everyday robot tasks. The questions were coded inductively into a two-level hierarchy of 12 categories and 70 subcategories. The authors report importance ratings for each question, use a linear mixed-effects model to compare importance across categories, and examine how question types and importance scores relate to participants' robot experience and attitudes. They claim that potential-issue questions receive the highest importance ratings, that why-questions rank relatively low, and that robotics novices ask different question types than more experienced users.","tokens_in":20879,"tokens_out":6165,"duration_ms":51068,"significance":"If the dataset is internally consistent and the public artifacts match the manuscript, this would be a valuable empirical resource for XAI and HRI: it provides a structured, publicly available corpus with a codebook, inter-annotator agreement (Cohen's kappa = .77 on a 100-question subset), and a reproducible analysis pipeline. The finding that users prioritize hypothetical-scenario and correctness-assurance questions over the why-questions dominant in much XAI literature is practically useful for prioritizing logging and explanation functionality. However, the current manuscript contains multiple internal numeric inconsistencies in the central dataset-size claim and related descriptive statistics; these must be resolved before the dataset can be treated as a canonical resource.","major_comments":[{"comment":"The central claim that the dataset contains 1,893 questions is not supported by the manuscript's own numbers. Section 4 reports 2,052 collected and 143 excluded, which gives 1,909, not 1,893. The per-category sample sizes in Figure 3 sum to 1,948 (120+190+126+245+83+157+52+417+106+209+125+118), and these counts match the subcategory counts in Figure 2. If each question belongs to exactly one category, the categorized dataset should contain 1,948 rows. The abstract, Section 3.3, and Section 4 thus give three different totals. Please reconcile these numbers and verify against the released GitHub/HuggingFace dataset, since the reuse value of the dataset depends on knowing exactly which questions are included.","section":"§1, §3.3, §4, Figure 2, Figure 3"},{"comment":"The top-category percentages are inconsistent between the abstract and the introduction: the abstract reports execution-details 21.4%, capabilities 12.6%, performance 10.7%, while Section 1 reports 22.5%, 12.7%, and 11.3%. These do not correspond to the same denominator (e.g., the Figure 2/3 counts give 417/1948 = 21.4%, but 417/1893 = 22.0%). Additionally, the abstract says 7 text stimuli while Section 3.2 says 'we wrote six text reports'; Section 3.1 implies a pool of 22 stimuli (15 videos + 7 texts). Please correct these inconsistencies, as they affect both the summary statistics and the reproducibility of the stimulus set.","section":"Abstract vs. §1; §3.2"},{"comment":"The contribution statement claims that users with different robot experience 'ask different questions,' but this is supported only by descriptive percentage differences in Figure 5 with no inferential test. The text reports, for example, that low-experience participants asked execution-details questions 23.3% of the time versus 15.7% for high-experience participants, without a significance test or effect-size estimate. If this is a central finding of RQ3, please add an appropriate statistical analysis (e.g., a mixed-effects model or permutation test on category proportions) or explicitly label the observation as descriptive and temper the contribution claim.","section":"§4.3, Figure 5"}],"minor_comments":[{"comment":"The robot-experience variable is described in Section 4.3 as a 1–5 scale, but Figure 1 and Figure 5 describe it as a 1–7 Likert scale. Please make the scale consistent everywhere, including in the regression interpretation.","section":"§3.4, §4.3"},{"comment":"The reliability subsection says the second annotator re-annotated '100 sentences ... balanced across categories, i.e., contained 8 samples per category.' With 12 categories, 8 samples per category gives 96, not 100. Please clarify how the 100-question sample was constructed and whether the balance statement is exact.","section":"§3.4"},{"comment":"The x-axis label reads '0 (not at all important), 5 (extremely important),' but the text and Methods say the scale is 1 (Not at all important) to 5 (extremely important). Please correct the axis label.","section":"Figure 3"},{"comment":"The list of pairwise contrasts reports corrected p-values only. For transparency, please also report the raw p-values or state explicitly that the reported values are Holm–Bonferroni-adjusted.","section":"§4.2"},{"comment":"The comparison of text- vs. video-elicited questions (36% vs. 16% execution-details) is informative, but it would be helpful to report whether this difference is statistically tested or is purely descriptive.","section":"§5.1"}],"recommendation":"major_revision","confidential_remarks":"The dataset and code are publicly linked, so the count inconsistencies can be checked directly. In my view, the correct total among 1,893, 1,909, and 1,948 is a load-bearing fact, not a cosmetic issue; the authors should resolve it and make the public artifacts consistent before this paper is accepted."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague, this one is worth a look, but bring a calculator. The paper delivers the first sizable natural-language dataset of user questions for household robots: 1,893 questions (claimed) from 100 participants, coded into a 12-category, 70-subcategory hierarchy with per-question importance scores. That alone is a genuine contribution—prior corpora are far smaller. The finding that potential-issues questions rank highest while why-questions rank second-lowest is a useful challenge to the field's focus on causal explanation. The methods are transparent: attention checks, inter-annotator agreement at κ=.77 on 100 questions, linear mixed-effects models with participant and stimulus random intercepts, Holm-Bonferroni correction. The authors also openly discuss ecological validity and novelty effects.\n\nThe soft spot is the arithmetic. Section 4 says 2,052 collected and 143 excluded, which gives 1,909, not 1,893. The n's in Figure 3 sum to 1,948, matching the subcategory counts in Figure 2. The abstract gives different top-category percentages than the body, and says seven text stimuli while Section 3.2 says six. Three different dataset totals in one paper is not a typo problem; it's a data-integrity problem. The public GitHub/HuggingFace artifacts are linked, so this is checkable in an afternoon. My guess is the 1,893 figure is from an earlier filtering pass and the figures were updated later, but the authors need to reconcile.\n\nThe ecological validity issue is real but proportional: questions were prompted by videos and text reports, not observed in real interaction, and the authors acknowledge this. That's a design limitation, not a fatal flaw. The count inconsistencies are more pressing because the dataset's reuse value depends on knowing exactly which questions are in it and how they're categorized.\n\nFor a reader: if you work on explainable robotics, HRI, or robot QA, this is a useful resource and a good starting point. It deserves a serious referee—send it out—but the authors should be asked to fix the counts and standardize the abstract/body numbers before it's treated as canonical.","headline":"Valuable dataset with a solid design, but the reported numbers don't add up—fix the counts before it's treated as canonical.","tokens_in":21424,"tokens_out":2592,"would_cite":true,"duration_ms":20516,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A new dataset of 1,893 user questions maps what people want to ask household robots.","keywords":["user questions","explainable robotics","human-robot interaction","question-answering","question taxonomy","household robots","explainable AI","dataset"],"falsifier":"Run a longitudinal field study where participants live with an actual household robot for several weeks and their spontaneous questions are logged as they ask them, then compare the distribution and importance ratings to the 12-category taxonomy; if the category frequencies or the importance ordering diverge substantially for the same tasks, the dataset's transferability breaks down.","tokens_in":20458,"feed_emoji":"🤖","tokens_out":2174,"duration_ms":20236,"temperature":0.7,"pith_summary":"This paper introduces a dataset of 1,893 natural-language questions that people would ask a household robot, collected from 100 participants who watched videos or read text summaries of robots doing everyday chores. The questions are organized into 12 categories and 70 subcategories, revealing that most questions concern execution details, robot capabilities, and performance assessment. When users rate importance, questions about how the robot would handle potential issues and ensure correct behavior rank highest, while the why-questions that dominate explainable-robotics research rank near the bottom. The dataset is offered as an empirical foundation for deciding what a robot should log, what question-answering systems should be built, and how explanations should be matched to user expectations.","feed_headline":"1,893 user questions reveal what people want to ask robots","feed_subtitle":"Safety and hypothetical-scenario questions rank highest, while the why-questions favored in XAI rank low.","key_machinery":"The carry load of the argument is the dataset itself together with its two-level coding scheme: 1,893 user questions hierarchically organized into 12 main categories and 70 subcategories, each defined with examples. The collection protocol pairs video and text stimuli of robot household tasks with a structured prompt that asks participants to write questions the robot should be able to answer and to rate each question's importance on a 5-point scale. Statistical inference on importance rankings uses a linear mixed-effects model with participant and stimulus as random intercepts, allowing within-participant and within-stimulus correlation to be accounted for.","core_discovery":"The paper's central claim is that the space of questions users want household robots to answer is much broader than the why-questions typically studied in explainable AI, and that this space can be systematically mapped. By eliciting questions from 100 participants across 15 video and 7 text stimuli of real robot task executions, the authors construct a hierarchical taxonomy with 12 main categories and 70 subcategories. The most frequently asked categories are execution-details (21.4%), what-abilities (12.6%), and self/task-assessment (10.7%). However, importance ratings from the same participants rank potential-issues questions (hypothetical difficulties and correctness assurance) as most i","pith_inferences":["The importance ranking could be converted into a weighted evaluation metric for generative question-answering systems, penalizing failure on high-importance categories more heavily.","The taxonomy may transfer beyond household robots to other service or collaborative robots, but validating that transfer would require similar elicitation in those domains.","The smaller gap between video and text elicitation results hints that users ask different questions about a robot's past activity versus live action; this could inform whether a robot should give periodic summaries or only respond on request.","A natural next step would be to collect paired user answers alongside questions to disambiguate intent, which the authors themselves note as a limitation and future direction."],"forward_implications":["Robot developers can use the taxonomy to decide which data a robot must log during task execution, since answering questions about environment state, execution details, or capabilities each requires different recorded information.","Question-answering and explanation modules should prioritize questions about potential issues and correctness assurance over the why-questions that have dominated XAI research, based on the importance rankings.","The dataset can serve as a benchmark for evaluating robot question-answering systems, offering a ground-truth set of user-generated questions with importance weights.","The difference in question types between novices and experienced users implies that adaptive question-answering systems may need to tailor their scope and defaults to user expertise.","The divergence between question frequency and importance within categories suggests that rare but highly important questions may deserve proactive explanation features rather than being overlooked."],"fun_headline_variants":["Users want robots to answer more than why-questions","New dataset: 1,893 questions people ask household robots","Hypothetical-scenario questions top robot QA importance","Beyond why: mapping what users ask robots","Novices and experts ask robots different questions"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The questions people write while watching short video clips or reading text summaries are assumed to be essentially the same questions they would actually ask a real household robot during repeated, everyday interaction.","fun_headline_variants_meta":{"raw":{"variants":["Users want robots to answer more than why-questions","New dataset: 1,893 questions people ask household robots","Hypothetical-scenario questions top robot QA importance","Beyond why: mapping what users ask robots","Novices and experts ask robots different questions"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000194,"raw_usage":{"total_tokens":1244,"prompt_tokens":848,"completion_tokens":396,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":592,"completion_tokens_details":{"reasoning_tokens":322}},"tokens_in":592,"tokens_out":396,"duration_ms":4064,"temperature":1.0,"reasoning_tokens":322,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T09:13:23.721393+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a longitudinal field study where participants live with an actual household robot for several weeks and their spontaneous questions are logged as they ask them, then compare the distribution and importance ratings to the 12-category taxonomy; if the category frequencies or the importance ordering diverge substantially for the same tasks, the dataset's transferability breaks down.","supporting_citations":[],"review_version":1}