{"id":"586b0334-3736-4cc4-bc1c-5e7ef439affc","arxiv_id":"2501.11556","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"The paper proposes a taxonomy of data-expectation gap experiences, covering sources of evidence, mismatch types, contextual factors, and personal evaluation mechanisms, based on two qualitative studies.","lead":"Researchers analyzed 200 smartwatch product reviews and ran a three-week field study with 16 users to build a vocabulary for describing when wearable data disagree with what people expect or feel. The framework, called the data-expectation gap, maps the contexts, emotions, and personal factors that turn a mismatched number into a tense or untrusted experience.","discovery_kind":"unification","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Vocabulary validation is circular: Study 2 codes with Study 1 themes and the B.4 mismatch rule is tuned on the same field data, so the breadth and analytic-utility claims are not independently established.","rationale":"The reader's weakest assumption correctly identifies the tension-potential protocol as tuned to the field data, and that is a real problem. I agree with the CONDITIONAL verdict and do not propose moving it. However, I see the tension-potential issue as one symptom of a broader validation gap: the entire vocabulary, not just the tension-potential term, is derived and then 're-confirmed' within the same research process. Study 2's deductive coding is explicitly influenced by Study 1 themes, so its apparent corroboration is partly circular. The B.4 rule is a second instance of fitting to the data. There is also no inter-rater reliability check, which matters because the vocabulary is offered as a shared analytical tool. The paper is transparent about many limitations, and the qualitative descriptions are rich and plausible, so rejection is not warranted. But the strongest claim, that the vocabulary captures a broad and context-bound spectrum and can be used by others as a design and analysis tool, would need independent application to a new dataset with acceptable inter-coder agreement. Until that test is run, CONDITIONAL remains the right verdict, and my read does not change it.","tokens_in":31774,"tokens_out":3932,"duration_ms":47667,"concrete_test":"Perform an external validation on a holdout corpus not used in vocabulary development: e.g., 50 newly sampled Amazon reviews and 6 fresh semi-structured interviews with owners of different smartwatch models. Have two researchers who were not involved in the original analysis apply only the final vocabulary definitions in Section 6.1, with no access to the original codes or themes. Report coverage (proportion of mismatch episodes assignable to at least one category) and inter-coder agreement per category (e.g., Cohen's kappa). In addition, re-run the B.4 quartile protocol on the new field data without re-tuning the thresholds, and compare the resulting 'tension potential' rates with participants' recollections.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim (Section 6.1) is that the vocabulary captures the breadth and context-bound character of encounters with the data-expectation gap and can serve as a design and analytical tool. For that claim to hold, the vocabulary must be recognizable and applicable beyond the two datasets that generated it. That condition is currently unverified, and the in-paper evidence has a circular structure. Section 5.1.4 states that Study 2's deductive analysis was 'influenced by' the findings and themes of Study 1, so re-identifying those same themes in the interviews is partly an artifact of the coding instrument rather than independent confirmation. Section B.4 says the rule for identifying potential mismatches from field data was 'experimentally determined' on the same data, with quartile thresholds chosen to flag strong disagreement and repeated-experience inconsistency; the resulting 'tension potential' counts in Section 5.2.1 are therefore fitted to the data rather than a predictive or independently grounded measure. All coding was performed by the first author with no inter-rater reliability check, and Section 7.4 acknowledges a convenience sample and a single device but offers no external validation. None of this makes the vocabulary false, but it does mean the strongest claim about breadth and analytic utility rests on the very data that produced the categories.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper addresses the 'data-expectation gap'—the mismatch between data reported by smartwatches and users' expectations—and proposes a vocabulary of experiential qualities describing such mismatches. The vocabulary is built from two studies: a thematic analysis of 200 online product reviews of four smartwatch models (Study 1), and a three-week in-the-wild study with 16 participants using a Fitbit Inspire 3, experience sampling, and interviews (Study 2). The resulting taxonomy is organized into three tension development mechanisms—mismatch detection, contextualisation, and personal evaluation—and introduces the construct 'tension potential.' The paper claims that the vocabulary captures the breadth and context-bound character of encounters with the data-expectation gap and can serve both as a design tool and as an analytical framework for Human-Data Interaction.","tokens_in":31994,"tokens_out":5621,"duration_ms":50872,"significance":"If the vocabulary is treated as a descriptive synthesis, the paper offers a useful contribution to Personal Informatics and HDI: it consolidates fragmented prior findings, provides detailed empirical grounding with ample participant quotes and full appendix materials, and translates the taxonomy into concrete design guidelines. The authors are transparent about their analysis process and openly list limitations, including the convenience sample and single device. However, the paper's strongest claim—that Study 2 validates Study 1's themes and that tension potential is a meaningful empirical construct—is undermined by a circular coding path and a data-fitted mismatch rule. The contribution is therefore stronger as synthesis than as validated theory, and the load-bearing validation claims need reworking.","major_comments":[{"comment":"The deductive analysis of Study 2 interviews was explicitly 'influenced by the same themes as the deductive analysis of Study 1, and the findings of Study 1' (§5.1.4), yet §5.2.2 reports that 'The themes identified in the previous study were, therefore, also present in this study' and §5.3 concludes that 'The second study validated the parameters affecting mismatch perception found in the first study.' This is a circular validation: re-discovering themes with a coding scheme built from those same themes does not constitute independent confirmation. The paper should either present Study 2 as an extension or elaboration rather than a validation, or provide an independent confirmatory analysis (e.g., a second coder blind to Study 1 themes, or a pre-registered coding scheme).","section":"§5.1.4, §5.2.2, §5.3"},{"comment":"The rule for identifying 'potential mismatches' from field data was 'experimentally determined' on the same data (Appendix B.4), using quartile thresholds to flag opposing-quartile disagreements and repeated-experience inconsistency. The 'tension potential' counts reported in §5.2.1 (e.g., eight participants with possible repeated-experience inconsistencies in sleep, of whom only P03 and P14 recalled such experiences) are therefore fitted to the very data they are used to describe, not independent predictions. To be load-bearing, the construct of tension potential needs external validation on a hold-out dataset or an a-priori threshold; otherwise it should be framed as an exploratory descriptor.","section":"§5.2.1, Appendix B.4"},{"comment":"All coding of the online reviews was performed by the first author with no inter-rater reliability check, and the selection of the 200 reviews used the first 50 reviews per brand with contextual detail (Appendix A.1) rather than a random sample. Without a second coder or a random sample, the reported frequencies (e.g., data-logic mismatch 55%, overestimation 33%) should be treated as tentative, and the claim in §4.1.1 that the sample was 'sufficient for understanding the breadth of accuracy issues' is not fully supported.","section":"§4.1.2, §4.1.1, Appendix A.1"},{"comment":"The limitations section acknowledges the convenience sample, single device, and lack of external validation, but these limitations conflict with the strong claims in §5.3 and §6.1 that the vocabulary captures 'the breadth and context-bound character' of encounters and can serve as a design and analytical tool. The paper should either temper these claims to describe the vocabulary as grounded in two specific datasets, or add a genuine external check (e.g., coding a hold-out set of reviews or interviews from a different device or brand).","section":"§7.4, §5.3, §6.1"}],"minor_comments":[{"comment":"In the definition of within-parameter co-occurrence, 'overestimation and overestimation of the same parameter' should presumably be 'overestimation and underestimation of the same parameter.'","section":"§6.1"},{"comment":"Rooksby et al. is cited as [22] in the sentence 'Rooksby et al. [22] emphasise emotional, social, and temporal aspects...', but the reference list assigns [22] to Li et al. (2010) and [17] to Rooksby et al.; this appears to be a citation error.","section":"§2.1"},{"comment":"Participant labels are inconsistent: 'P3' should be 'P03' and 'P016' should be 'P16.'","section":"§5.2.2"},{"comment":"The full-text title contains a typo ('V ocabulary' with an internal space) and Appendix A.1 contains 'ommissions,' which should be 'omissions.'","section":"Title, Appendix A.1"},{"comment":"The application of the vocabulary to LLM interactions is a useful speculation, but it is not supported by data in this paper; the text should make clear that this is an extrapolation rather than a finding.","section":"§7.2"}],"recommendation":"major_revision","confidential_remarks":"The paper is likely to be of interest to the CHI/IJHCI audience. The main concern is the gap between the validation claims and the evidence: the deductive coding path and the data-fitted mismatch rule make the 'validation' of Study 1 themes and the 'tension potential' construct circular. I would encourage the editor to ask for a revision that either provides an independent check (e.g., blind second coding, hold-out data) or substantially rewrites the claims to acknowledge the exploratory nature of the findings. The vocabulary itself has merit and could be restructured as a descriptive framework pending external validation."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The short version: this paper gives the field a genuinely useful shared vocabulary for a problem that has been described in fragments for a decade. The data-expectation gap, the three sources of evidence (data-logic, data-measurement, data-feeling), and the tension development mechanisms are a real synthesis, not just a relabeling of existing terms. The empirical work is solid qualitative research, and the paper is unusually transparent about its methods and limits.\n\nWhat I found most useful: the domain mapping. Table 1 alone justifies the paper—it shows how scattered the prior factors are. The two studies complement each other: 200 Amazon reviews for breadth, 16 field participants with ESM and Fitbit logs for depth. The field study's notion of tension potential—that the data contain many possible mismatches but users perceive only a few depending on interaction and interest—is a nice contribution, even if it is not the centerpiece.\n\nNow the soft spots, and they are real but not fatal. The second study is not independent confirmation. Section 5.1.4 says the deductive analysis was \"influenced by\" Study 1's themes, and then Section 5.2.2 reports that those themes \"were, therefore, also present in this study.\" That is circular, and the authors should either present Study 2 as an extension rather than a validation, or bring in external data. Similarly, the B.4 protocol for flagging potential mismatches was \"experimentally determined\" on the same field data, so the tension potential counts are fitted, not predictive. The review sample is the first 50 with contextual detail per brand, which risks selection on vividness. With a single coder and no inter-rater reliability, the categories are the first author's reading, honestly disclosed but not independently checked.\n\nNone of this makes the vocabulary false. It does mean the strongest claim—breadth and analytic utility—rests on one dataset plus a deductively-coded second study. The paper would be stronger if it framed the vocabulary as a synthesis to be tested, rather than a validated instrument.\n\nWho should read it: personal informatics and HCI researchers studying trust, accuracy, and self-tracking, plus designers of wearable feedback. The Section 6.2 scenario walkthrough is practical and would help a design team. A serious referee should engage with this—it deserves peer review, and the revision requests should focus on validation framing rather than rejecting the contribution.","headline":"A genuinely useful vocabulary for a real gap in wearable HCI, but the second study confirms the first rather than validating it independently.","tokens_in":32546,"tokens_out":2624,"would_cite":true,"duration_ms":29121,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A vocabulary of experiential qualities for the data-expectation gap in smartwatches.","keywords":["data-expectation gap","smartwatch accuracy","human-data interaction","personal informatics","user experience","tension potential","experience sampling","wearable fitness trackers"],"falsifier":"A direct test would be to run the same three-week protocol with a different watch model and a broader sample, flag potential mismatches with the same quartile rule, and ask participants immediately whether they noticed each discrepancy; the claim would be undermined if flagged potential mismatches are not perceived substantially more often than unflagged data, or if a substantial share of perceived mismatches cannot be expressed in the vocabulary's categories.","tokens_in":31527,"feed_emoji":"⌚","tokens_out":8086,"duration_ms":75447,"temperature":0.7,"pith_summary":"Many smartwatch users have watched their device report steps they never took, floors that do not exist, or a good sleep score after a restless night. The paper argues that these encounters, grouped under the term data-expectation gap, are not a single accuracy problem but a spectrum of experiences shaped by how a mismatch is detected, where and when it happens, and how the user evaluates it. Drawing on 200 online product reviews and a three-week field study with 16 participants, the authors build a vocabulary with three tension development mechanisms: mismatch detection, contextualisation, and personal evaluation. The claim is that this vocabulary captures the breadth and context-bound character of the phenomenon and can be used both to design human-data interactions and to analyse user experiences in a structured way.","feed_headline":"Three mechanisms explain why smartwatch data mismatches sting","feed_subtitle":"A two-study vocabulary turns the data-expectation gap into a design problem instead of an accuracy problem.","key_machinery":"The central object is the data-expectation gap, defined as a mismatch between detected and expected values, with a single instance called a mismatch. The argument is carried by a vocabulary of experiential qualities organised into three tension development mechanisms—mismatch detection, contextualisation, and personal evaluation—whose building blocks can be combined to describe any encounter. A supporting construct is tension potential, which separates mismatches that merely exist in logged data from mismatches a user actually perceives, explaining why the same watch data produce friction for one person and not another.","core_discovery":"The central discovery is that the same underlying data error can produce very different experiences, and that the difference is not captured by accuracy metrics alone. The paper defines the data-expectation gap as a mismatch between detected and expected values related to behaviours, actions, sensations, beliefs, or feelings, with a single instance called a mismatch. From two studies it derives a vocabulary whose building blocks group into three tension development mechanisms: mismatch detection, covering the source of evidence (data-logic, data-measurement, or data-feeling) and the type of mismatch (classification and value-estimation issues); contextualisation, covering activity, location, temporality, emotional state, and co-experience; and personal evaluation, covering personal factors, explainability, and the history of encounters. The paper also introduces tension potential, the idea that field data contain many possible mismatches that most users never perceive, because perception depends on interaction patterns, interest, and context.","pith_inferences":["A natural extension is to turn tension potential into a measurement instrument: log all potential mismatches from sensor data, use experience sampling to see which ones users actually notice, and test whether the vocabulary predicts perceived friction.","The finding that some users feel invalidated when data contradict a strongly felt negative state suggests a testable boundary for the assumption that positive feedback improves emotion, and could inform how subjective metrics such as sleep and stress are framed.","Because prior encounters with better-performing devices raised disappointment, onboarding and marketing could manage expectations; the paper does not propose this, but it follows from the history-of-experiences mechanism.","If context determines the meaning of a mismatch, adding uncertainty or confidence annotations to data could shift a mismatch from the data-logic category to the data-measurement category and reduce disbelief; this is a design hypothesis the paper leaves untested."],"forward_implications":["Designers can use the vocabulary as a scenario generator, walking through each data type and mismatch type paired with contextual and personal factors to prototype mitigations for tension.","Researchers gain a shared terminology that unifies previously disparate descriptions such as data inaccuracy, mismatch with beliefs, and mismatch with feelings.","The work implies that improving sensor precision or algorithmic recall will not eliminate the data-expectation gap; at least some tension must be addressed through explanation, feedback, customisation, and reflection mechanisms.","The vocabulary can be transferred to other data-driven and AI systems, where the same three mechanisms can structure how user responses to errors are analysed."],"supporting_citations":[{"why":"It supplies the accuracy-assessment strategies (ad-hoc assessment and folk-testing) and the evidence that users often assess device accuracy poorly.","marker":"[12]"},{"why":"It supplies the experience-centred framework whose compositional, sensual, emotional, and spatio-temporal threads shape the vocabulary's contextual categories.","marker":"[15]"},{"why":"It provides the call for a new terminology beyond traditional error metrics for activity inference systems.","marker":"[49]"},{"why":"It supplies evidence on subjective stress tracking and the mismatch between data and feelings, grounding the data-feeling mismatch category.","marker":"[14]"},{"why":"It introduces the tension between sensory-bodily and metric knowledge that underlies the data-feeling mismatch and trust in the self.","marker":"[28]"},{"why":"It provides the data sensemaking perspective and the term mismatch between data and expectations in self-tracking.","marker":"[4]"},{"why":"It supplies the concept of situated objectivity, the context-dependent assessment of measurement accuracy and value.","marker":"[44]"},{"why":"It supplies evidence that mismatches between data and beliefs affect trust and engagement with personal informatics tools.","marker":"[13]"}],"fun_headline_variants":["Why smartwatch data errors feel different","A vocabulary for smartwatch data-expectation gaps","Three mechanisms turn data errors into user tension","Two studies, one vocabulary for wearable data gaps"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The central claim depends on the assumption that the moments flagged as possible mismatches in the field data really are moments users would experience as a gap, even though the flagging rule was tuned on those same data.","fun_headline_variants_meta":{"raw":{"variants":["Why smartwatch data errors feel different","A vocabulary for smartwatch data-expectation gaps","Three mechanisms turn data errors into user tension","Two studies, one vocabulary for wearable data gaps"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.0004,"raw_usage":{"total_tokens":2068,"prompt_tokens":899,"completion_tokens":1169,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":515,"completion_tokens_details":{"reasoning_tokens":1112}},"tokens_in":515,"tokens_out":1169,"duration_ms":9646,"temperature":1.0,"reasoning_tokens":1112,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T18:07:06.433175+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A direct test would be to run the same three-week protocol with a different watch model and a broader sample, flag potential mismatches with the same quartile rule, and ask participants immediately whether they noticed each discrepancy; the claim would be undermined if flagged potential mismatches are not perceived substantially more often than unflagged data, or if a substantial share of perceived mismatches cannot be expressed in the vocabulary's categories.","supporting_citations":[{"cited_title":"Newman, and Mark S","cited_arxiv_id":null,"evidence_quote":"It supplies the accuracy-assessment strategies (ad-hoc assessment and folk-testing) and the evidence that users often assess device accuracy poorly."},{"cited_title":"Technology as Experience","cited_arxiv_id":null,"evidence_quote":"It supplies the experience-centred framework whose compositional, sensual, emotional, and spatio-temporal threads shape the vocabulary's contextual categories."},{"cited_title":"McDonald, Tammy Toscos, Mike Y","cited_arxiv_id":null,"evidence_quote":"It provides the call for a new terminology beyond traditional error metrics for activity inference systems."},{"cited_title":"Data engagement reconsidered: A study of automatic stress tracking technology in use","cited_arxiv_id":null,"evidence_quote":"It supplies evidence on subjective stress tracking and the mismatch between data and feelings, grounding the data-feeling mismatch category."},{"cited_title":"The temporal flows of self-tracking: Checking in, moving on, staying hooked","cited_arxiv_id":null,"evidence_quote":"It introduces the tension between sensory-bodily and metric knowledge that underlies the data-feeling mismatch and trust in the self."},{"cited_title":"Data sensemaking in self-tracking: Towards a new generation of self-tracking tools","cited_arxiv_id":null,"evidence_quote":"It provides the data sensemaking perspective and the term mismatch between data and expectations in self-tracking."},{"cited_title":"Living the metrics: Self-tracking and situated objectivity","cited_arxiv_id":null,"evidence_quote":"It supplies the concept of situated objectivity, the context-dependent assessment of measurement accuracy and value."},{"cited_title":"Personal informatics for everyday life: How users without prior self-tracking experience engage with personal data","cited_arxiv_id":null,"evidence_quote":"It supplies evidence that mismatches between data and beliefs affect trust and engagement with personal informatics tools."}],"review_version":1}