{"id":"6b5cdd6f-ddd2-4114-9a22-2a2dff408903","arxiv_id":"2607.21527","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A 14-person interview study identifies five situated sources of inequity and 15 lifecycle fairness risks in wellbeing sensing, going beyond identity-based audits.","lead":"Interviews with 14 wellbeing-sensing researchers and practitioners surface five non-demographic sources of inequity and 15 fairness risks that span the full system lifecycle. The findings argue that auditing only model performance by race or gender misses most of what makes passive sensing unfair.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'systematically shape' claim overreaches a 14-interview, researcher-only, high-income-country sample; missing data-provider/end-user perspectives may make the five-source taxonomy incomplete.","rationale":"The paper is a well-conducted qualitative study: the interview instruments are included, coding procedures are transparent, and the authors are candid about limitations. The central insight—that fairness in passive sensing goes beyond post-hoc model audits—is plausible and supported by the data. However, the strongest claim ('systematically shape') depends on the completeness and representativeness of the five-source taxonomy. The reader's weakest assumption identifies exactly this: a small, network-recruited, academic-heavy, high-income-country sample, with no direct input from data providers or end users. I agree that this is the most load-bearing concern because the paper's proposed governance interventions (funding requirements, publication standards, IRB expansions) are premised on the taxonomy being broadly applicable. The proposed concrete test—interviewing data providers and end users, including from low- and middle-income countries—would directly test whether the taxonomy is exhaustive or whether additional sources of inequity exist. If the test reproduces the five sources, the claim gains strength; if it surfaces new sources, the paper's central generalization is unsupported. The reader's conditional verdict, with a request to temper the abstract and supplement with broader stakeholder interviews, is appropriate. No internal inconsistency or methodological flaw rises to the level of rejection; the concern is about external validity, not correctness. Hence, UNCHANGED relative to the reader's CONDITIONAL verdict.","tokens_in":20092,"tokens_out":4727,"duration_ms":51553,"concrete_test":"Run a second interview wave with a purposive sample of at least 15 data providers and end users from a mix of high- and low-income countries, using the same Part 1 and Part 2 protocol but omitting the Part 3 framework prompt, and independently code the transcripts for sources of inequity. If new sources emerge (e.g., disability-related access, social stigma, non-smartphone usage) or the five sources fail to reproduce, the 'systematically shape' wording must be tempered and the taxonomy revised.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract and conclusion claim that the five identified sources 'systematically shape fairness risks beyond identity-based attributes,' but the entire empirical basis is 14 semi-structured interviews with researchers/practitioners, 11 from academia, all working in five high-income countries (Recruitment & Participants; Table 2). The authors acknowledge in Limitations that their analysis 'relies solely on researcher accounts rather than the lived experiences of data providers, end users, or domain experts.' This is load-bearing because the taxonomy's completeness and the strength of the word 'systematically' are what motivate the governance recommendations (Discussion). A network-recruited, researcher-only sample from high-income countries cannot establish that these five sources are the recurring, systematic drivers of inequity in passive sensing more broadly. Data providers and end users—especially those with disabilities, from low-income settings, or under institutional surveillance—might identify additional or different sources (e.g., accessibility barriers, stigma, coercion). The absence of any low- or middle-income country also means device access, digital literacy, and cultural familiarity may have different manifestations. Without those perspectives, the empirical basis for the central claim is thinner than the wording suggests.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports semi-structured interviews with 14 passive-sensing researchers and practitioners across five countries (Australia, Canada, Japan, Korea, U.S.) to understand how fairness risks arise in wellbeing sensing beyond post-hoc identity-based model audits. It identifies five situated sources of inequity (comfort with monitoring/vulnerability, device/sensor access, digital/data literacy, cultural/linguistic familiarity, behavioral regularity), synthesizes 15 lifecycle-stage fairness risks with mitigation strategies (Table 1, Figure 1), and documents structural/institutional barriers that constrain fair practice. It concludes that fair passive sensing requires not only individual researcher effort but also ecosystem-level governance from funders, publication venues, IRBs, and deploying institutions. The paper includes the full interview protocol, participant demographics, and a detailed account of the thematic analysis process.","tokens_in":20307,"tokens_out":5317,"duration_ms":58191,"significance":"The contribution is potentially valuable: it extends fairness discourse beyond model-level audits to the full sensing lifecycle, grounds the discussion in practitioner accounts, and produces a concrete risk/mitigation table that could guide future work and governance. The study is transparent about its methods and limitations, and the thematic analysis uses multiple coding rounds with independent coding and discrepancy resolution, which supports credibility. The paper does not offer formal metrics or machine-checkable artifacts; its contribution is qualitative and interpretive. The main risk is that the strength of the central claim ('systematically shape fairness risks beyond identity-based attributes') exceeds what the evidence can support, given the small researcher-only, high-income-country sample and the absence of saturation or member-checking evidence.","major_comments":[{"comment":"The central claim that five situated sources 'systematically shape fairness risks beyond identity-based attributes' is stronger than the evidence presented. The empirical basis is 14 self-selected, network-recruited researchers/practitioners, 11 from academia, all in five high-income countries. The paper itself states that it 'relies solely on researcher accounts rather than the lived experiences of data providers, end users, or domain experts.' No saturation analysis or member checking is reported. A qualitative sample of this size and composition can identify plausible mechanisms and recurring concerns within the sample, but cannot establish that these five sources are the systematic drivers of inequity in passive sensing generally. This overreach matters because the governance implications in the Discussion rest on the systematicity of the identified sources. Please either temper the","section":"Abstract; Findings (Situated Sources of Inequity); Limitations"},{"comment":"The lifecycle framing may be partially an artifact of the research instrument. Participants were asked in Part 2, 'Which step(s) in the pipeline present potential fairness risks?' and in Part 3 were asked to annotate their pipeline diagram using a draft lifecycle-aware fairness framework adapted from the authors' prior work (Zhang et al. 2023). The four-phase structure in Table 1 and Figure 1 closely mirrors this scaffold. The paper should explain how the analysis distinguished participant-emergent themes from responses elicited by the lifecycle prompts, and should explicitly acknowledge this potential confirmatory bias. This does not invalidate the findings, but it is load-bearing for the claim that the lifecycle taxonomy is empirically grounded rather than imposed by the interview design.","section":"Procedure (Part 3); Table 1; Figure 1"},{"comment":"The 'five countries' framing overstates diversity: all five are high-income countries, and the cultural/linguistic examples are drawn from international students or workers within those countries, not from low- or middle-income settings. Missing perspectives from data providers with disabilities, low-income users, or people under institutional surveillance may yield additional sources of inequity (e.g., accessibility barriers, coercion, stigma) not captured here. At minimum, the paper should describe the sample as 'researcher perspectives from five high-income countries' and qualify claims about cultural drivers to the contexts actually studied. This is a scope limitation, but it directly affects the completeness of the proposed taxonomy.","section":"Recruitment & Participants; Limitations"}],"minor_comments":[{"comment":"The coding process is described in detail, but no inter-coder reliability statistic or saturation metric is reported. For a thematic analysis this is not mandatory, but one sentence on how the authors judged that themes were stable would strengthen the systematicity claim.","section":"Data Analysis"},{"comment":"The table is dense and mixes risks, mitigation strategies, and stakeholder actions. Consider separating the 'Potential Risks' and 'Mitigation Strategies' columns visually or using a full-page layout; the current format is hard to read in a two-column article.","section":"Table 1"},{"comment":"The reference to the authors' own prior framework (Zhang et al. 2023) is appropriate, but the degree to which that framework influenced the interview instrument and analysis should be stated in the main text, not only in the Procedure section.","section":"Discussion and Conclusion"},{"comment":"The screening survey asks whether participants have 'explicitly considered fairness-related concerns.' This could prime later responses about fairness awareness; this is worth a sentence in Limitations as a potential social-desirability bias.","section":"Screening Survey Questions"}],"recommendation":"major_revision","confidential_remarks":"For the editor: The paper makes a useful, domain-specific contribution and the qualitative methodology is largely sound, but the abstract and central framing overstate the strength of the evidence. The self-referential lifecycle instrument in Part 3 is the point I would most want the authors to address in revision, along with tempering 'systematically' or adding saturation/triangulation support. I do not see grounds for rejection; the manuscript is within the scope of the journal and the limitations are, in principle, fixable through careful rewriting and additional transparency."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth a read if you work on fairness in sensing or lifecycle audits. The contribution is a domain-specific, empirically grounded synthesis: five situated sources of inequity (comfort with monitoring, device access, literacy, cultural familiarity, behavioral regularity) and 15 risks/mitigations across the lifecycle. The interviews are analyzed with standard thematic analysis, the coding process is transparent, and the instruments are included. They are candid about limitations, noting the sample is researcher-only and not exhaustive.\n\nThe central argument—that post-hoc identity-based audits miss a lot in passive sensing—is plausible and supported by their data. The examples (GPS refusal, iOS-only exclusion, cultural unfamiliarity with stimuli) are concrete and ring true. The structural barriers section (incentives, publishing culture) is a useful reminder that awareness of fairness often doesn't translate into practice.\n\nThe soft spots are real but mostly about framing. Fourteen interviews, eleven from academia, all in five high-income countries, cannot establish that these five sources 'systematically shape' fairness risks in the broader domain. The absence of data providers, end users, and any low/middle-income perspective leaves open the possibility of missing sources—accessibility barriers, stigma, coercion. They acknowledge this in limitations, but the abstract and conclusion still use 'systematically,' which overreaches. Also, participants were asked to critique a framework adapted from the authors' own prior work, so the validation is mildly self-referential; that's not fatal, but worth noting.\n\nNo load-bearing flaw. The paper is honest about what it can and can't claim, and the taxonomy is a good starting point for designers and regulators. It deserves a serious referee; I'd suggest the authors soften the abstract wording and, ideally, supplement with broader stakeholder interviews before publication. I'd cite this in my own work.","headline":"Solid qualitative study with a useful taxonomy; the abstract's 'systematically' claims more than 14 researcher interviews can support.","tokens_in":20818,"tokens_out":1700,"would_cite":true,"duration_ms":18603,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues that fairness in passive wellbeing sensing is shaped mainly by five situated sources of inequity—comfort with monitoring, device access, digital literacy, cultural familiarity, and behavioral regularity—so post-hoc identity","keywords":["passive sensing","wellbeing","fairness","inequity","lifecycle","behavioral inference","qualitative interviews","governance"],"falsifier":"A quantitative audit of a deployed sensing system that found identity-based attributes such as race or gender explain more variance in data missingness, dropout, or prediction error than the five situated factors—or a direct survey of data providers showing they do not experience monitoring comfort, device access, literacy, cultural familiarity, or behavioral regularity as sources of unequal treatment—would undercut the paper's central claim.","tokens_in":19968,"feed_emoji":"⚖️","tokens_out":3829,"duration_ms":35590,"temperature":0.7,"pith_summary":"This paper tries to establish that fairness in passive wellbeing sensing—smartphones and wearables that infer stress, depression, or cognitive load—is decided long before a model is trained. Through interviews with 14 researchers and practitioners across five countries, it identifies five situated sources of inequity that cut across identity categories: comfort with monitoring, device and sensor access, digital literacy, cultural familiarity, and behavioral regularity. It maps 15 fairness risks and corresponding mitigations across the full sensing lifecycle and argues that post-hoc identity-based audits miss the main pathways by which unfairness enters. If this account is right, fair practice requires lifecycle-wide attention to representation, burden, and procedural agency, backed by funders, publication venues, ethics boards, and deploying institutions.","feed_headline":"Fairness gaps in wellbeing sensing start before the model","feed_subtitle":"Interviews map 15 fairness risks from study design to deployment, beyond identity-based audits.","key_machinery":"The carrying mechanism is a lifecycle model of the passive sensing pipeline, split into study planning, data preparation, modeling/system design, and deployment, with behavioral inference treated as a fairness-relevant intermediate stage. Within this model, the paper's central objects are the five situated sources of inequity (comfort with monitoring and vulnerability, device and sensor access, digital and data literacy, cultural or linguistic familiarity, and behavioral regularity) and the 15 risk–mitigation pairs, which together explain how representation, burden, and procedural agency get distributed unevenly across the system's lifetime.","core_discovery":"The central discovery is an empirically grounded account of how fairness breaks down in passive wellbeing sensing. The authors interviewed 14 researchers and practitioners across five countries and found that fairness risks cluster around five situated sources of inequity rather than only around demographic identity. They synthesize 15 fairness risks paired with mitigation strategies spanning study planning, data preparation, modeling and system design, and deployment, treating behavioral inference as a distinct stage where unverified interpretations of sensor signals can silently propagate into model errors. The paper also reports that researchers widely recognize these risks but face struc","pith_inferences":["Beyond the paper: the five situated sources could be operationalized as measurable covariates—for example, dropout-risk scores based on monitoring comfort or device class—and fed into quantitative fairness audits, giving the interview findings a direct testable form.","Beyond the paper: the lifecycle-situated framing likely transfers to other longitudinal or sensor-heavy ML domains, such as workplace productivity analytics or digital phenotyping for clinical trials, where similar upstream inference and burden dynamics appear.","Beyond the paper: if behavioral regularity is as powerful a driver of model error as the interviews suggest, existing sensing datasets could be re-analyzed to check whether irregular-routine participants show systematically higher prediction error; that is a low-cost, direct test of the paper's central claim."],"forward_implications":["If the account is right, fairness audits that compare error rates across demographic groups will keep missing the inequities that matter most, because the operative fault lines are device access, monitoring comfort, literacy, cultural familiarity, and behavioral regularity.","Mitigation has to start before data collection: pilot analyses for dropout risk, explicit inclusion targets for situational subgroups, and burden-calibrated study designs.","Behavioral inference—converting raw signals into constructs like location entropy—must be validated against participants' lived experience rather than treated as routine preprocessing or feature engineering.","Deploying institutions such as hospitals, universities, and employers acquire ongoing obligations to monitor post-deployment disparities, provide meaningful recourse, and disclose system limitations.","Funders and publication venues should require fairness documentation and pre-registered protocols, because individual researcher goodwill is systematically overridden by existing incentive structures."],"fun_headline_variants":["Fairness in sensing fails at five hidden sources","Interviews reveal 15 fairness risks in wellbeing sensing","Wellbeing sensing fairness starts long before the model","Five inequity sources plague passive sensing","Fairness risks span the whole sensing lifecycle"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The taxonomy rests on the assumption that 14 mostly academic researchers from five high-income countries can faithfully report how fairness risks arise—the paper itself notes that the analysis relies solely on researcher accounts rather than the lived experiences of data providers, end users, or domain experts.","fun_headline_variants_meta":{"raw":{"variants":["Fairness in sensing fails at five hidden sources","Interviews reveal 15 fairness risks in wellbeing sensing","Wellbeing sensing fairness starts long before the model","Five inequity sources plague passive sensing","Fairness risks span the whole sensing lifecycle"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000175,"raw_usage":{"total_tokens":1120,"prompt_tokens":736,"completion_tokens":384,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":480,"completion_tokens_details":{"reasoning_tokens":329}},"tokens_in":480,"tokens_out":384,"duration_ms":3931,"temperature":1.0,"reasoning_tokens":329,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T07:09:30.861744+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A quantitative audit of a deployed sensing system that found identity-based attributes such as race or gender explain more variance in data missingness, dropout, or prediction error than the five situated factors—or a direct survey of data providers showing they do not experience monitoring comfort, device access, literacy, cultural familiarity, or behavioral regularity as sources of unequal treatment—would undercut the paper's central claim.","supporting_citations":[],"review_version":1}