Pith. sign in

REVIEW 4 major objections 5 minor 78 references

The Data-Expectation Gap: A Vocabulary Describing Experiential Qualities of Data Inaccuracies in Smartwatches

T0 review · 4 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read A vocabulary of experiential qualities for the data-expectation gap in smartwatches.

desk verdict A genuinely useful vocabulary for a real gap in wearable HCI, but the second study confirms the first rather than validating it independently. read the letter →

arxiv 2501.11556 v1 pith:PK3L2A54 submitted 2025-01-20 cs.HC

classification cs.HC
keywords data-expectationgapsmartwatchaccuracyhuman-datainteractionpersonalinformaticsuserexperiencetensionpotentialsamplingwearablefitnesstrackers
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Many smartwatch users have watched their device report steps they never took, floors that do not exist, or a good sleep score after a restless night. The paper argues that these encounters, grouped under the term data-expectation gap, are not a single accuracy problem but a spectrum of experiences shaped by how a mismatch is detected, where and when it happens, and how the user evaluates it. Drawing on 200 online product reviews and a three-week field study with 16 participants, the authors build a vocabulary with three tension development mechanisms: mismatch detection, contextualisation, and personal evaluation. The claim is that this vocabulary captures the breadth and context-bound character of the phenomenon and can be used both to design human-data interactions and to analyse user experiences in a structured way.

What carries the argument

The central object is the data-expectation gap, defined as a mismatch between detected and expected values, with a single instance called a mismatch. The argument is carried by a vocabulary of experiential qualities organised into three tension development mechanisms—mismatch detection, contextualisation, and personal evaluation—whose building blocks can be combined to describe any encounter. A supporting construct is tension potential, which separates mismatches that merely exist in logged data from mismatches a user actually perceives, explaining why the same watch data produce friction for one person and not another.

What would settle it

A direct test would be to run the same three-week protocol with a different watch model and a broader sample, flag potential mismatches with the same quartile rule, and ask participants immediately whether they noticed each discrepancy; the claim would be undermined if flagged potential mismatches are not perceived substantially more often than unflagged data, or if a substantial share of perceived mismatches cannot be expressed in the vocabulary's categories.

Watch

Extended reading notes

Core claim

The central discovery is that the same underlying data error can produce very different experiences, and that the difference is not captured by accuracy metrics alone. The paper defines the data-expectation gap as a mismatch between detected and expected values related to behaviours, actions, sensations, beliefs, or feelings, with a single instance called a mismatch. From two studies it derives a vocabulary whose building blocks group into three tension development mechanisms: mismatch detection, covering the source of evidence (data-logic, data-measurement, or data-feeling) and the type of mismatch (classification and value-estimation issues); contextualisation, covering activity, location, temporality, emotional state, and co-experience; and personal evaluation, covering personal factors, explainability, and the history of encounters. The paper also introduces tension potential, the idea that field data contain many possible mismatches that most users never perceive, because perception depends on interaction patterns, interest, and context.

Load-bearing premise

The central claim depends on the assumption that the moments flagged as possible mismatches in the field data really are moments users would experience as a gap, even though the flagging rule was tuned on those same data.

Editorial extensions

If this is right

  • Designers can use the vocabulary as a scenario generator, walking through each data type and mismatch type paired with contextual and personal factors to prototype mitigations for tension.
  • Researchers gain a shared terminology that unifies previously disparate descriptions such as data inaccuracy, mismatch with beliefs, and mismatch with feelings.
  • The work implies that improving sensor precision or algorithmic recall will not eliminate the data-expectation gap; at least some tension must be addressed through explanation, feedback, customisation, and reflection mechanisms.
  • The vocabulary can be transferred to other data-driven and AI systems, where the same three mechanisms can structure how user responses to errors are analysed.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A natural extension is to turn tension potential into a measurement instrument: log all potential mismatches from sensor data, use experience sampling to see which ones users actually notice, and test whether the vocabulary predicts perceived friction.
  • The finding that some users feel invalidated when data contradict a strongly felt negative state suggests a testable boundary for the assumption that positive feedback improves emotion, and could inform how subjective metrics such as sleep and stress are framed.
  • Because prior encounters with better-performing devices raised disappointment, onboarding and marketing could manage expectations; the paper does not propose this, but it follows from the history-of-experiences mechanism.
  • If context determines the meaning of a mismatch, adding uncertainty or confidence annotations to data could shift a mismatch from the data-logic category to the data-measurement category and reduce disbelief; this is a design hypothesis the paper leaves untested.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper addresses the 'data-expectation gap'—the mismatch between data reported by smartwatches and users' expectations—and proposes a vocabulary of experiential qualities describing such mismatches. The vocabulary is built from two studies: a thematic analysis of 200 online product reviews of four smartwatch models (Study 1), and a three-week in-the-wild study with 16 participants using a Fitbit Inspire 3, experience sampling, and interviews (Study 2). The resulting taxonomy is organized into three tension development mechanisms—mismatch detection, contextualisation, and personal evaluation—and introduces the construct 'tension potential.' The paper claims that the vocabulary captures the breadth and context-bound character of encounters with the data-expectation gap and can serve both as a design tool and as an analytical framework for Human-Data Interaction.

Significance. If the vocabulary is treated as a descriptive synthesis, the paper offers a useful contribution to Personal Informatics and HDI: it consolidates fragmented prior findings, provides detailed empirical grounding with ample participant quotes and full appendix materials, and translates the taxonomy into concrete design guidelines. The authors are transparent about their analysis process and openly list limitations, including the convenience sample and single device. However, the paper's strongest claim—that Study 2 validates Study 1's themes and that tension potential is a meaningful empirical construct—is undermined by a circular coding path and a data-fitted mismatch rule. The contribution is therefore stronger as synthesis than as validated theory, and the load-bearing validation claims need reworking.

major comments (4)
  1. [§5.1.4, §5.2.2, §5.3] The deductive analysis of Study 2 interviews was explicitly 'influenced by the same themes as the deductive analysis of Study 1, and the findings of Study 1' (§5.1.4), yet §5.2.2 reports that 'The themes identified in the previous study were, therefore, also present in this study' and §5.3 concludes that 'The second study validated the parameters affecting mismatch perception found in the first study.' This is a circular validation: re-discovering themes with a coding scheme built from those same themes does not constitute independent confirmation. The paper should either present Study 2 as an extension or elaboration rather than a validation, or provide an independent confirmatory analysis (e.g., a second coder blind to Study 1 themes, or a pre-registered coding scheme).
  2. [§5.2.1, Appendix B.4] The rule for identifying 'potential mismatches' from field data was 'experimentally determined' on the same data (Appendix B.4), using quartile thresholds to flag opposing-quartile disagreements and repeated-experience inconsistency. The 'tension potential' counts reported in §5.2.1 (e.g., eight participants with possible repeated-experience inconsistencies in sleep, of whom only P03 and P14 recalled such experiences) are therefore fitted to the very data they are used to describe, not independent predictions. To be load-bearing, the construct of tension potential needs external validation on a hold-out dataset or an a-priori threshold; otherwise it should be framed as an exploratory descriptor.
  3. [§4.1.2, §4.1.1, Appendix A.1] All coding of the online reviews was performed by the first author with no inter-rater reliability check, and the selection of the 200 reviews used the first 50 reviews per brand with contextual detail (Appendix A.1) rather than a random sample. Without a second coder or a random sample, the reported frequencies (e.g., data-logic mismatch 55%, overestimation 33%) should be treated as tentative, and the claim in §4.1.1 that the sample was 'sufficient for understanding the breadth of accuracy issues' is not fully supported.
  4. [§7.4, §5.3, §6.1] The limitations section acknowledges the convenience sample, single device, and lack of external validation, but these limitations conflict with the strong claims in §5.3 and §6.1 that the vocabulary captures 'the breadth and context-bound character' of encounters and can serve as a design and analytical tool. The paper should either temper these claims to describe the vocabulary as grounded in two specific datasets, or add a genuine external check (e.g., coding a hold-out set of reviews or interviews from a different device or brand).
minor comments (5)
  1. [§6.1] In the definition of within-parameter co-occurrence, 'overestimation and overestimation of the same parameter' should presumably be 'overestimation and underestimation of the same parameter.'
  2. [§2.1] Rooksby et al. is cited as [22] in the sentence 'Rooksby et al. [22] emphasise emotional, social, and temporal aspects...', but the reference list assigns [22] to Li et al. (2010) and [17] to Rooksby et al.; this appears to be a citation error.
  3. [§5.2.2] Participant labels are inconsistent: 'P3' should be 'P03' and 'P016' should be 'P16.'
  4. [Title, Appendix A.1] The full-text title contains a typo ('V ocabulary' with an internal space) and Appendix A.1 contains 'ommissions,' which should be 'omissions.'
  5. [§7.2] The application of the vocabulary to LLM interactions is a useful speculation, but it is not supported by data in this paper; the text should make clear that this is an extrapolation rather than a finding.

Circularity Check

2 steps flagged · score 5.0 of 10

The vocabulary is empirically grounded, but its validation path is partially circular: Study 2 re-imported Study 1 themes and the B.4 mismatch rule was tuned on the same field data it is used to corroborate.

  1. fitted input called prediction [Appendix B.4 and Section 5.2.1]
    "We experimentally determined this protocol, aiming to identify a) instances of strong agreement or disagreement, and b) instances of inconsistency in repeated experiences, which a few participants recalled in the interviews."

    The rule for flagging potential mismatches from field data was tuned on the same field data, with thresholds chosen to reproduce phenomena that a few participants had already recalled. Section 5.2.1 then reports 'Field data corroborated most recollections (Table 3)' and defines 'tension potential' using counts generated by that tuned rule. The corroboration and the tension-potential counts are therefore produced by a rule fitted to the outcome they are claimed to confirm, rather than being an independent check.

  2. self definitional [Sections 5.1.4 and 5.2.2]
    "Initial codes were generated using inductive analysis, followed by deductive analysis influenced by the same themes as the deductive analysis of Study 1, and the findings of Study 1. ... The themes identified in the previous study were, therefore, also present in this study."

    The paper presents Study 2 as confirming the applicability of Study 1's themes, but the coding instrument already contained those themes through deductive analysis 'influenced by' Study 1. The word 'therefore' exposes the logical dependence: the presence of the themes in Study 2 is an expected consequence of the coding lens, not an independent observation. This partially undermines the claimed breadth and validation of the vocabulary, although Study 2 also contributed inductively new themes, so the central derivation is not wholly circular.

full rationale

The core vocabulary is a qualitative taxonomy built from 200 online reviews and 16 interviews, and its main definitions are not predicted from fitted parameters; that part of the derivation is self-contained and not circular. However, two validation moves are partially circular. First, Study 2's deductive coding imported Study 1 themes, and Section 5.2.2 then states these themes 'were, therefore, also present in this study'—the presence is partly manufactured by the coding instrument rather than independently discovered. Second, the B.4 protocol for identifying potential mismatches was 'experimentally determined' on the same field data and aimed at inconsistency patterns 'which a few participants recalled in the interviews'; the resulting tension-potential counts and the statement that 'field data corroborated most recollections' are therefore generated by a rule fitted to the outcome they are said to corroborate. These issues weaken the independent confirmation of the vocabulary, but they do not make the vocabulary itself equivalent to its inputs: Study 2 produced inductively new themes (emotional state, trust, explainability), and the Section 6.2.2 re-application to own field incidents is explicitly illustrative rather than a prediction. There is no load-bearing self-citation in the derivation; the authors' own prior work appears only as a peripheral example in Section 7.2. Score 5 reflects partial circularity in the validation path, not in the central definitional derivation.

Assumptions & free parameters 1 free parameters · 3 assumptions · 2 invented entities

The paper's central claim rests on qualitative self-report data, an adopted analytic framework, and single-coder thematic analysis. The only fitted numeric input is the threshold rule for detecting potential mismatches in the field data, which is tuned to the same dataset. The invented entities are conceptual vocabulary terms rather than physical quantities, and they lack independent falsifiable handles.

free parameters (1)
  • Mismatch detection thresholds (Q1/Q3 quartile rule) = Participant-specific lower and upper quartiles of ESM and Fitbit scores
    In Appendix B.4, the protocol to identify potential mismatches was 'experimentally determined' by the authors to flag strong disagreement or repeated inconsistency, so the rule is tuned to their field dataset.
assumptions (3)
  • domain assumption Participants' self-reports in reviews, ESM responses, and interviews accurately reflect their experienced encounters with the data-expectation gap.
    The entire empirical basis is self-report data; the authors acknowledge recall bias in Section 7.4 but argue triangulation addresses it.
  • domain assumption The Technology as Experience (TaE) framework's threads (compositional, sensual, emotional, spatio-temporal) are appropriate analytic lenses for categorizing mismatches.
    Adopted in Section 3.3 to guide analysis; if the framework does not fit, the vocabulary's structure would differ.
  • domain assumption The inductive and deductive coding by the first author, refined with the second author, captures the phenomena without systematic bias.
    Sections 4.1.2 and 5.1.4 describe single-coder analysis with discussion, not independent coding or inter-rater reliability measures.
invented entities (2)
  • Data-expectation gap
    purpose: Umbrella construct to unify mismatches between detected and expected values across behaviors, feelings, and beliefs.
    Introduced by definition in Section 3.2; it is a conceptual frame, not a measured quantity.
  • Tension potential
    purpose: New construct describing the probability, based on data logs, that a user encounters a mismatch even if not perceived.
    Defined in Section 5.2.1 from field data comparisons; no external validation is provided.

how reviews work

0 comments
Cite this review

Pith. "Pith review of The Data-Expectation Gap: A Vocabulary Describing Experiential Qualities of Data Inaccuracies in Smartwatches." pith.science (2026). https://pith.science/paper/PK3L2A54

@misc{pith2026250111556,
  author       = {Pith},
  title        = {Pith review of: The Data-Expectation Gap: A Vocabulary Describing Experiential Qualities of Data Inaccuracies in Smartwatches},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/PK3L2A54}},
  note         = {Machine review of arXiv:2501.11556}
}
read the original abstract

Many users of wrist-worn wearable fitness trackers encounter the data-expectation gap - mismatches between data and expectations. While we know such discrepancies exist, we are no closer to designing technologies that can address their negative effects. This is largely because encounters with mismatches are typically treated unidimensionally, while they may differ in context and implications. This treatment does not allow the design of human-data interaction (HDI) mechanisms accounting for temporal, social, emotional, and other factors potentially influencing the perception of mismatches. To address this problem, we present a vocabulary that describes the breadth and context-bound character of encounters with the data-expectation gap, drawing from findings from two studies. Our work contributes to Personal Informatics research providing knowledge on how encounters with the data-expectation gap are embedded in people's daily lives, and a vocabulary encapsulating this knowledge, which can be used when designing HDI experiences in wearable fitness trackers.

Figures

Figures reproduced from arXiv: 2501.11556 by the authors.

Figure 1
Figure 1. The process we followed to construct the vocabulary. [PITH_FULL_IMAGE:figures/full_fig_p006_1.png] view at source ↗
Figure 2
Figure 2. The smartwatches that we considered in the detailed analysis of online reviews. Fitbit Sense, which is not [PITH_FULL_IMAGE:figures/full_fig_p007_2.png] view at source ↗
Figure 3
Figure 3. Fitbit Inspire 3, which was used in the field study (Figure retrieved by Fitbit [54]). Source of evidence Type of mismatch and Parameter Scenario F Sleep quality overestimation Bad sleep, good sleep score in device F Sleep quality underestimation Good sleep, bad sleep score in device F Stress intensity overestimation Device shows high stress while you are calm F Stress intensity underestimation Device shows low stre… view at source ↗
Figures from the paper (8 more)
Figure 5
Figure 5. Figure 5: Evidence of the data-expectation gap in SQ. [PITH_FULL_IMAGE:figures/full_fig_p016_5.png]
Figure 6
Figure 6. Figure 6: Potential mismatches between perceived and detected SQ during the field experiment, and related contextual [PITH_FULL_IMAGE:figures/full_fig_p017_6.png]
Figure 7
Figure 7. Figure 7: Comparative analysis of activities in ESM and Fitbit logs. [PITH_FULL_IMAGE:figures/full_fig_p018_7.png]
Figure 8
Figure 8. Figure 8: Alignment between ESM and Fitbit activity logs, and related contextual data. The data are visualised as a [PITH_FULL_IMAGE:figures/full_fig_p019_8.png]
Figure 9
Figure 9. Figure 9: Alignment between P01’s ESM and Fitbit activity logs, represented following the same scheme as Figure 8. [PITH_FULL_IMAGE:figures/full_fig_p020_9.png]
Figure 10
Figure 10. Figure 10: P01’s negative experience of a mismatch in SQ data. The figure shows a snapshot of P01’s Fitbit and ESM [PITH_FULL_IMAGE:figures/full_fig_p025_10.png]
Figure 11
Figure 11. Figure 11: P04’s negative experience of a mismatch in SQ data, visualised following the same approach as Figure 10. [PITH_FULL_IMAGE:figures/full_fig_p026_11.png]
Figure 12
Figure 12. Figure 12: The ESM protocol. – (+)When were you taking off the device? – (+)Did you notice any recommendations related to sleep or activity? If yes, what did you do? • ESM data collection – What did you think of the m-Path app? Was there anything that you found interesting, usef…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

78 extracted references · 74 canonical work pages

  1. [1]

    Michaelis, Michael A

    Jessica R. Michaelis, Michael A. Rupp, James Kozachuk, Baotran Ho, Daniela Zapata-Ocampo, Daniel S. McConnell, and Janan A. Smither. Describing the user experience of wearable fitness technology through online product reviews. Proceedings of the Human Factors and Ergonomics Society Annual Meeting, 60(1):1073–1077, 2016

  2. [2]

    Wellbeing in the making: peoples’ experiences with wearable activity trackers

    Evangelos Karapanos, Rúben Gouveia, Marc Hassenzahl, and Jodi Forlizzi. Wellbeing in the making: peoples’ experiences with wearable activity trackers. Psychology of well-being, 6:1–17, 2016

  3. [3]

    A multi-level classification approach for sleep stage prediction with processed data derived from consumer wearable activity trackers

    Zilu Liang and Mario Alberto Chapa-Martell. A multi-level classification approach for sleep stage prediction with processed data derived from consumer wearable activity trackers. Frontiers in Digital Health, 3, 2021

  4. [4]

    Data sensemaking in self-tracking: Towards a new generation of self-tracking tools

    Aykut Co¸ skun and Arma˘gan Karahano ˘glu. Data sensemaking in self-tracking: Towards a new generation of self-tracking tools. International Journal of Human–Computer Interaction, 39(12):2339–2360, 2023

  5. [5]

    Users’ experiences of wearable activity trackers: a cross-sectional study

    Carol Maher, Jillian Ryan, Christina Ambrosi, and Sarah Edney. Users’ experiences of wearable activity trackers: a cross-sectional study. BMC public health, 17:1–8, 2017. 32 The Data-Expectation Gap: A V ocabulary of Experiential Qualities

  6. [6]

    Making lifelogging usable: Design guidelines for activity trackers

    Jochen Meyer, Jutta Fortmann, Merlin Wasmann, and Wilko Heuten. Making lifelogging usable: Design guidelines for activity trackers. In MultiMedia Modeling: 21st International Conference, MMM 2015, Sydney, NSW, Australia, January 5-7, 2015, Proceedings, Part II 21, pages 323–334. Springer, 2015

  7. [7]

    User satisfaction with wearables

    Raquel Benbunan-Fich. User satisfaction with wearables. AIS Transactions on Human-Computer Interaction, 12(1):1–27, 2020

  8. [8]

    Abandonment of personal quantification: A review and empirical study investigating reasons for wearable activity tracking attrition

    Christiane Attig and Thomas Franke. Abandonment of personal quantification: A review and empirical study investigating reasons for wearable activity tracking attrition. Computers in Human Behavior, 102:223–237, 2020

Show all 78 references
  1. [9]

    Activity tracking: Barriers, workarounds and customisation

    Daniel Harrison, Paul Marshall, Nadia Bianchi-Berthouze, and Jon Bird. Activity tracking: Barriers, workarounds and customisation. In Proceedings of the 2015 ACM International Joint Conference on Pervasive and Ubiquitous Computing, UbiComp ’15, page 617–621, New York, NY , USA...

  2. [10]

    The rise and fall of wearable fitness trackers

    Coorevits, Lynn and Coenen, Tanguy. The rise and fall of wearable fitness trackers. In Academy of Management, page 24, 2016

  3. [11]

    Amanda Lazar, Christian Koehler, Theresa Jean Tanenbaum, and David H. Nguyen. Why we use and abandon smart devices. In Proceedings of the 2015 ACM International Joint Conference on Pervasive and Ubiquitous Computing, UbiComp ’15, page 635–646, New York, NY , USA, 2015. Associa...

  4. [12]

    Newman, and Mark S

    Rayoung Yang, Eunice Shin, Mark W. Newman, and Mark S. Ackerman. When fitness trackers don’t ’fit’: End-user difficulties in the assessment of personal tracking device accuracy. In Proceedings of the 2015 ACM International Joint Conference on Pervasive and Ubiquitous Computing...

  5. [13]

    Personal informatics for everyday life: How users without prior self-tracking experience engage with personal data

    Amon Rapp and Federica Cena. Personal informatics for everyday life: How users without prior self-tracking experience engage with personal data. International Journal of Human-Computer Studies, 94:1–17, 2016

  6. [14]

    Data engagement reconsidered: A study of automatic stress tracking technology in use

    Xianghua (Sharon) Ding, Shuhan Wei, Xinning Gui, Ning Gu, and Peng Zhang. Data engagement reconsidered: A study of automatic stress tracking technology in use. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems, CHI ’21, New York, NY , USA, 2021. A...

  7. [15]

    Technology as Experience

    John McCarthy and Peter Wright. Technology as Experience. MIT Press, 2007

  8. [16]

    Making sense of experience

    Peter Wright, John McCarthy, and Lisa Meekison. Making sense of experience. In Funology: From usability to enjoyment, pages 43–53. Springer, 2003

  9. [17]

    Personal tracking as lived informatics

    John Rooksby, Mattias Rost, Alistair Morrison, and Matthew Chalmers. Personal tracking as lived informatics. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems, CHI ’14, page 1163–1172, New York, NY , USA, 2014. Association for Computing Machinery

  10. [18]

    Huang, Gail C

    Thomas Fritz, Elaine M. Huang, Gail C. Murphy, and Thomas Zimmermann. Persuasive technology in the real world: A study of long-term use of activity sensing devices for fitness. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems, CHI ’14, page 487–496...

  11. [19]

    Understanding the most satisfying and unsatisfying user experiences: Emotions, psychological needs, and context

    Timo Partala and Aleksi Kallinen. Understanding the most satisfying and unsatisfying user experiences: Emotions, psychological needs, and context. Interacting with Computers, 24(1):25–34, 10 2011

  12. [20]

    Human-data interaction: The human face of the data-driven society

    Richard Mortier, Hamed Haddadi, Tristan Henderson, Derek McAuley, and Jon Crowcroft. Human-data interaction: The human face of the data-driven society. arXiv preprint arXiv:1412.6159, 2014

  13. [21]

    Fitbit sense 2

    Fitbit. Fitbit sense 2. https://www.fitbit.com/global/us/products/smartwatches/sense2, 2023. Accessed: 2023-09-13

  14. [22]

    A stage-based model of personal informatics systems

    Ian Li, Anind Dey, and Jodi Forlizzi. A stage-based model of personal informatics systems. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems, CHI ’10, page 557–566, New York, NY , USA,

  15. [23]

    Lee, Bongshin Lee, Wanda Pratt, and Julie A

    Eun Kyoung Choe, Nicole B. Lee, Bongshin Lee, Wanda Pratt, and Julie A. Kientz. Understanding quantified- selfers’ practices in collecting and exploring personal data. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems, CHI ’14, page 1143–1152, New Y...

  16. [24]

    Wo´ zniak

    Jasmin Niess and Paweł W. Wo´ zniak. Supporting meaningful personal fitness: the tracker goal evolution model. In Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems, CHI ’18, page 1–12, New York, NY , USA, 2018. Association for Computing Machinery

  17. [25]

    Understanding information systems continuance: An expectation-confirmation model

    Anol Bhattacherjee. Understanding information systems continuance: An expectation-confirmation model. MIS quarterly, pages 351–370, 2001. 33 The Data-Expectation Gap: A V ocabulary of Experiential Qualities

  18. [26]

    Reynolds, Sean Victory, Kai Zheng, and Yunan Chen

    Mayara Costa Figueiredo, Clara Caldeira, Tera L. Reynolds, Sean Victory, Kai Zheng, and Yunan Chen. Self- tracking for fertility care: Collaborative support for a highly personalized problem. Proc. ACM Hum.-Comput. Interact., 1(CSCW), dec 2017

  19. [27]

    Personal data contexts, data sense, and self-tracking cycling

    Deborah Lupton, Sarah Pink, Christine Heyes LaBond, and Shanti Sumartojo. Personal data contexts, data sense, and self-tracking cycling. International Journal of Communication, 12:647–666, 2018

  20. [28]

    The temporal flows of self-tracking: Checking in, moving on, staying hooked

    Stine Lomborg, Nanna Bonde Thylstrup, and Julie Schwartz. The temporal flows of self-tracking: Checking in, moving on, staying hooked. New Media & Society, 20(12):4590–4607, 2018

  21. [29]

    Lee, Ashley Garrity, and Mark W

    Shriti Raj, Joyce M. Lee, Ashley Garrity, and Mark W. Newman. Clinical data in context: Towards sensemaking tools for interpreting personal health data. Proc. ACM Interact. Mob. Wearable Ubiquitous Technol., 3(1), mar 2019

  22. [30]

    Use and adoption challenges of wearable activity trackers

    Patrick C Shih, Kyungsik Han, Erika Shehan Poole, Mary Beth Rosson, and John M Carroll. Use and adoption challenges of wearable activity trackers. IConference 2015 proceedings, 2015

  23. [31]

    Phases of accuracy diagnosis:(in) visibility of system status in the fitbit

    Molly Zellweger Mackinlay. Phases of accuracy diagnosis:(in) visibility of system status in the fitbit. Intersect: The Stanford Journal of Science, Technology, and Society, 6(2), 2013

  24. [32]

    On being told how we feel: How algorithmic sensor feedback influences emotion perception

    Victoria Hollis, Alon Pekurovsky, Eunika Wu, and Steve Whittaker. On being told how we feel: How algorithmic sensor feedback influences emotion perception. Proc. ACM Interact. Mob. Wearable Ubiquitous Technol., 2(3), sep 2018

  25. [33]

    Health by numbers? exploring the practice and experience of datafied health

    Gavin JD Smith and Ben V onthethoff. Health by numbers? exploring the practice and experience of datafied health. In Self-Tracking, Health and Medicine, pages 6–21. Routledge, 2017

  26. [34]

    Self-tracking while doing sport: Comfort, motivation, attention and lifestyle of athletes using personal informatics tools

    Amon Rapp and Lia Tirabeni. Self-tracking while doing sport: Comfort, motivation, attention and lifestyle of athletes using personal informatics tools. International Journal of Human-Computer Studies, 140:102434, 2020

  27. [35]

    User experience: A concept without consensus? exploring practitioners’ perspectives through an international survey

    Carine Lallemand, Guillaume Gronier, and Vincent Koenig. User experience: A concept without consensus? exploring practitioners’ perspectives through an international survey. Computers in Human Behavior, 43:35–48, 2015

  28. [36]

    Beyond usability: Evaluating emotional response as an integral part of the user experience

    Anshu Agarwal and Andrew Meyer. Beyond usability: Evaluating emotional response as an integral part of the user experience. In CHI ’09 Extended Abstracts on Human Factors in Computing Systems, CHI EA ’09, page 2919–2930, New York, NY , USA, 2009. Association for Computing Machinery

  29. [37]

    From metrics to experiences: Investigating how sport data shapes the social context, self-determination and motivation of athletes

    Dees Postma, Dennis Reidsma, Robby van Delden, and Arma ˘gan Karahano˘glu. From metrics to experiences: Investigating how sport data shapes the social context, self-determination and motivation of athletes. Interacting with Computers, page iwae012, 2024

  30. [38]

    Anxious or empowered? a cross-sectional study exploring how wearable activity trackers make their owners feel

    Jillian Ryan, Sarah Edney, and Carol Maher. Anxious or empowered? a cross-sectional study exploring how wearable activity trackers make their owners feel. BMC psychology, 7:1–8, 2019

  31. [39]

    Impact of using wearable devices on psychological distress: Analysis of the health information national trends survey

    Avishek Choudhury and Onur Asan. Impact of using wearable devices on psychological distress: Analysis of the health information national trends survey. International Journal of Medical Informatics, 156:104612, 2021

  32. [40]

    User experience white paper: Bringing clarity to the concept of user experience

    V Roto, EL-C Law, APOS Vermeeren, and J Hoonhout. User experience white paper: Bringing clarity to the concept of user experience. s.n., 2011. geen ISBN Result from Dagstuhl seminar on demarcating user experience, sepember 15-18, 2010

  33. [41]

    Intimasea: Exploring shared stress display in close relationships

    Yanqi Jiang, Xianghua(Sharon) Ding, Xiaojuan Ma, Zhida Sun, and Ning Gu. Intimasea: Exploring shared stress display in close relationships. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems, CHI ’23, New York, NY , USA, 2023. Association for Compu...

  34. [42]

    Vermeeren, and Joke Kort

    Effie Lai-Chong Law, Virpi Roto, Marc Hassenzahl, Arnold P.O.S. Vermeeren, and Joke Kort. Understanding, scoping and defining user experience: A survey approach. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems, CHI ’09, page 719–728, New York, NY ...

  35. [43]

    User experience over time: An initial framework

    Evangelos Karapanos, John Zimmerman, Jodi Forlizzi, and Jean-Bernard Martens. User experience over time: An initial framework. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems, CHI ’09, page 729–738, New York, NY , USA, 2009. Association for Comput...

  36. [44]

    Living the metrics: Self-tracking and situated objectivity

    Mika Pantzar and Minna Ruckenstein. Living the metrics: Self-tracking and situated objectivity. DIGITAL HEALTH, 3:2055207617712590, 2017. PMID: 29942604

  37. [45]

    real you

    Jeffrey Warshaw, Tara Matthews, Steve Whittaker, Chris Kau, Mateo Bengualid, and Barton A. Smith. Can an algorithm know the "real you"? understanding people’s reactions to hyper-personal analytics systems. In Proceedings of the 33rd Annual ACM Conference on Human Factors in Co...

  38. [46]

    Chang, Emily Sun, Saeed Abdullah, and Geri Gay

    Jaime Snyder, Mark Matthews, Jacqueline Chien, Pamara F. Chang, Emily Sun, Saeed Abdullah, and Geri Gay. Moodlight: Exploring personal and social implications of ambient display of biosensor data. In Proceedings of the 18th ACM Conference on Computer Supported Cooperative Work...

  39. [47]

    Dice in the black box: User experiences with an inscrutable algorithm

    Aaron Springer, Victoria Hollis, and Steve Whittaker. Dice in the black box: User experiences with an inscrutable algorithm. In 2017 AAAI Spring Symposium Series, 2017

  40. [48]

    Adams, Malte F

    Jean Costa, Alexander T. Adams, Malte F. Jung, François Guimbretière, and Tanzeem Choudhury. Emotioncheck: leveraging bodily signals and false feedback to regulate our emotions. In Proceedings of the 2016 ACM Interna- tional Joint Conference on Pervasive and Ubiquitous Computi...

  41. [49]

    McDonald, Tammy Toscos, Mike Y

    Sunny Consolvo, David W. McDonald, Tammy Toscos, Mike Y . Chen, Jon Froehlich, Beverly Harrison, Predrag Klasnja, Anthony LaMarca, Louis LeGrand, Ryan Libby, Ian Smith, and James A. Landay. Activity sensing in the wild: A field trial of ubifit garden. In Proceedings of the SIG...

  42. [50]

    Engaging with health data: The interplay between self-tracking activities and emotions in fertility struggles

    Mayara Costa Figueiredo, Clara Caldeira, Elizabeth Victoria Eikey, Melissa Mazmanian, and Yunan Chen. Engaging with health data: The interplay between self-tracking activities and emotions in fertility struggles. Proc. ACM Hum.-Comput. Interact., 2(CSCW), nov 2018

  43. [51]

    Iso 5725-1:1994(en)

    ISO. Iso 5725-1:1994(en). https://www.iso.org/obp/ui/#iso:std:iso:5725:-1:ed-1:v1:en , 2024. Accessed: 2024-06-06

  44. [52]

    Interaction design gone wild: striving for wild theory

    Yvonne Rogers. Interaction design gone wild: striving for wild theory. interactions, 18(4):58–62, 2011

  45. [53]

    Researcher positionality–a consideration of its influence and place in qualitative research–a new researcher guide

    Andrew Gary Darwin Holmes. Researcher positionality–a consideration of its influence and place in qualitative research–a new researcher guide. Shanlax International Journal of Education, 8(4):1–10, 2020

  46. [54]

    Fitbit. Fitbit. https://www.fitbit.com/global/us/products, 2023. Accessed: 2023-08-21

  47. [55]

    Garmin. Garmin. https://www.garmin.com/en-US/, 2023. Accessed: 2023-09-11

  48. [56]

    Fitbit statistics 2024 by users, retail price, import and export

    Barry Elad. Fitbit statistics 2024 by users, retail price, import and export. https://www.fitbit.com/global/ us/products, 2024. Accessed: 2024-10-02

  49. [57]

    m-path: An easy-to-use and flexible platform for ecological momentary assessment and intervention in behavioral research and clinical practice

    Merijn Mestdagh, Stijn Verdonck, Maarten Piot, Koen Niemeijer, Peter Kuppens, Egon Dejonckheere, et al. m-path: An easy-to-use and flexible platform for ecological momentary assessment and intervention in behavioral research and clinical practice. 2022

  50. [58]

    Experience sampling method: Measuring the quality of everyday life

    Joel M Hektner, Jennifer Anne Schmidt, and Mihaly Csikszentmihalyi. Experience sampling method: Measuring the quality of everyday life. Sage, 2007

  51. [59]

    JONKER, KAHYUN SOPHIE KIM, MIKKEL KRENCHEL, MORGAN RAMSEY-ELLIOT, FRIEDERIKE SCHÜÜR, DA VID ZAX, and JOANNA ZHANG

    MARIA CURY , ERYN WHITWORTH, SEBASTIAN BARFORT, SÉRÉNA BOCHEREAU, JONATHAN BROW- DER, TANY A R. JONKER, KAHYUN SOPHIE KIM, MIKKEL KRENCHEL, MORGAN RAMSEY-ELLIOT, FRIEDERIKE SCHÜÜR, DA VID ZAX, and JOANNA ZHANG. Hybrid methodology: Combining ethnography, cognitive science, and ...

  52. [60]

    Reflecting on reflexive thematic analysis

    Virginia Braun and Victoria Clarke. Reflecting on reflexive thematic analysis. Qualitative Research in Sport, Exercise and Health, 11(4):589–597, 2019

  53. [61]

    A survey on trust modeling

    Jin-Hee Cho, Kevin Chan, and Sibel Adali. A survey on trust modeling. ACM Comput. Surv., 48(2), oct 2015

  54. [62]

    Great expectations and broken promises: Misleading claims, product failure, expectancy disconfirmation and consumer distrust

    Peter R Darke, Laurence Ashworth, and Kelley J Main. Great expectations and broken promises: Misleading claims, product failure, expectancy disconfirmation and consumer distrust. Journal of the Academy of Marketing Science, 38:347–362, 2010

  55. [63]

    Lim, and Mohan Kankanhalli

    Ashraf Abdul, Jo Vermeulen, Danding Wang, Brian Y . Lim, and Mohan Kankanhalli. Trends and trajectories for explainable, accountable and intelligible systems: An hci research agenda. In Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems , CHI ’18, pag...

  56. [64]

    Creating an understanding of data literacy for a data-driven society

    Annika Wolff, Daniel Gooch, Jose J Cavero Montaner, Umar Rashid, and Gerd Kortuem. Creating an understanding of data literacy for a data-driven society. The Journal of Community Informatics, 12(3), 2016

  57. [65]

    Beyond accuracy: The role of mental models in human-ai team performance

    Gagan Bansal, Besmira Nushi, Ece Kamar, Walter S Lasecki, Daniel S Weld, and Eric Horvitz. Beyond accuracy: The role of mental models in human-ai team performance. In Proceedings of the AAAI conference on human computation and crowdsourcing, volume 7, pages 2–11, 2019. 35 The ...

  58. [66]

    Epstein, Monica Caraway, Chuck Johnston, An Ping, James Fogarty, and Sean A

    Daniel A. Epstein, Monica Caraway, Chuck Johnston, An Ping, James Fogarty, and Sean A. Munson. Beyond abandonment to next steps: Understanding and designing for life after personal informatics tool use. InProceedings of the 2016 CHI Conference on Human Factors in Computing Sys...

  59. [67]

    Biosignals as social cues: Ambiguity and emotional interpretation in social displays of skin conductance

    Noura Howell, Laura Devendorf, Rundong (Kevin) Tian, Tomás Vega Galvez, Nan-Wei Gong, Ivan Poupyrev, Eric Paulos, and Kimiko Ryokai. Biosignals as social cues: Ambiguity and emotional interpretation in social displays of skin conductance. In Proceedings of the 2016 ACM Confere...

  60. [68]

    to click or not to click

    Hans Brombacher, Dimitra Dritsa, Steven V os, and Steven Houben. "to click or not to click": Back to basic for experience sampling for office well-being in shared office spaces. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems , CHI ’24, New York...

  61. [69]

    Do you understand the words that are comin outta my mouth? voice assistant comprehension of medication names

    Adam Palanica, Anirudh Thommandram, Andrew Lee, Michael Li, and Yan Fossat. Do you understand the words that are comin outta my mouth? voice assistant comprehension of medication names. NPJ digital medicine, 2(1):55, 2019

  62. [70]

    A review on evaluation metrics for data classification evaluations

    Mohammad Hossin and Md Nasir Sulaiman. A review on evaluation metrics for data classification evaluations. International journal of data mining & knowledge management process, 5(2):1, 2015

  63. [71]

    Re-examining whether, why, and how human-ai interaction is uniquely difficult to design

    Qian Yang, Aaron Steinfeld, Carolyn Rosé, and John Zimmerman. Re-examining whether, why, and how human-ai interaction is uniquely difficult to design. In Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems, CHI ’20, page 1–13, New York, NY , USA, 2020....

  64. [72]

    Understanding choice independence and error types in human-ai collaboration

    Alexander Erlei, Abhinav Sharma, and Ujwal Gadiraju. Understanding choice independence and error types in human-ai collaboration. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems, CHI ’24, New York, NY , USA, 2024. Association for Computing Machinery

  65. [73]

    When algorithms err: Differential impact of early vs

    Antino Kim, Mochen Yang, and Jingjing Zhang. When algorithms err: Differential impact of early vs. late errors on users’ reliance on algorithms. ACM Trans. Comput.-Hum. Interact., 30(1), March 2023

  66. [74]

    Emotional biosensing: Exploring critical alternatives

    Noura Howell, John Chuang, Abigail De Kosnik, Greg Niemeyer, and Kimiko Ryokai. Emotional biosensing: Exploring critical alternatives. Proc. ACM Hum.-Comput. Interact., 2(CSCW), nov 2018

  67. [75]

    Cox, Sandy J.J

    Anna L. Cox, Sandy J.J. Gould, Marta E. Cecchinato, Ioanna Iacovides, and Ian Renfree. Design frictions for mindful interactions: The case for microboundaries. In Proceedings of the 2016 CHI Conference Extended Abstracts on Human Factors in Computing Systems, CHI EA ’16, page ...

  68. [76]

    Analysing user experience of personal mobile products through contextual factors

    Hannu Korhonen, Juha Arrasvuori, and Kaisa Väänänen-Vainio-Mattila. Analysing user experience of personal mobile products through contextual factors. In Proceedings of the 9th International Conference on Mobile and Ubiquitous Multimedia, MUM ’10, New York, NY , USA, 2010. Asso...

  69. [77]

    Optimal number of response categories in rating scales: reliability, validity, discriminating power, and respondent preferences

    Carolyn C Preston and Andrew M Colman. Optimal number of response categories in rating scales: reliability, validity, discriminating power, and respondent preferences. Acta Psychologica, 104(1):1–15, 2000. 36

  70. [2010]

    Association for Computing Machinery

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.