{"id":"438ef45e-0d96-4521-8dd7-b15a0fea89cd","arxiv_id":"2508.02680","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"The authors propose AnnoSense, a set of 15 expert-reviewed guidelines for everyday emotion data collection, derived from survey, interview, and focus group insights from 119 stakeholders.","lead":"This paper reports surveys, interviews, and focus groups with 119 stakeholders and 25 expert reviewers, and distills them into a 15-guideline framework called AnnoSense for collecting emotion data through wearables and phones in everyday settings. A smart generalist might care because the framework tackles a bottleneck in training emotion AI: how to get accurate, consent-aware, context-rich emotional labels outside the lab.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Expert Likert ratings of guideline clarity, usefulness, and adaptability are treated as evidence that AnnoSense improves real-world emotion data collection, but no deployment or baseline test exists; the framework's central value claim rests on this unvalidated linkage.","rationale":"The reader's weakest assumption identifies the same load-bearing concern: stakeholder preferences and expert Likert ratings are used as evidence of framework value without any deployment, baseline comparison, or inter-rater reliability for the thematic coding. My stress-test pass confirms this is the most critical gap. The paper is transparent about its methods and limitations, and the claims in Section 5 are stated narrowly as ratings of clarity, usefulness, and adaptability, not as measured improvements in data quality. However, the paper's framing in the abstract, introduction, and discussion repeatedly suggests that the framework will 'enhance' collection and analysis, which goes beyond what expert ratings establish. A two-arm field trial would directly test whether the guidelines produce measurable benefits in compliance, data completeness, annotation consistency, or downstream model performance. Until such a test exists, the appropriate verdict is conditional acceptance: the qualitative synthesis is valuable, but the framework's practical value remains unproven. The proposed test is expensive but would settle the concern; a cheaper interim check would be to re-analyze the existing expert ratings per guideline with statistical tests and inter-rater reliability, though that would only partially address the predictive validity question.","tokens_in":45205,"tokens_out":7376,"duration_ms":82177,"concrete_test":"Run a two-week field trial with two matched arms (n approximately 30 per arm). Arm A follows AnnoSense's pre-, during-, and post-data-collection guidelines (e.g., screening, calibration, training, adaptable structured and unstructured annotation, multi-source assessment where feasible). Arm B uses a standard fixed-schedule EMA protocol with the same wearable sensors. Collect identical physiological and self-report data. Compare: (1) annotation response rate; (2) proportion of self-reports with valid time-aligned physiological windows; (3) inter-annotator consistency (e.g., self-report versus expert-coded description); and (4) AUROC of a fixed baseline classifier (e.g., random forest on HR, EDA, accelerometer features) predicting self-reported valence. If Arm A does not significantly outperform Arm B on these metrics, the framework's central value claim is not supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The framework's central value claim is that following AnnoSense produces well-annotated physiological emotion data for AI. The evidence chain is: (1) guidelines derived from stakeholder surveys, interviews, and focus groups (Section 4), and (2) positive expert ratings of clarity, usefulness, and adaptability on 5-point Likert scales (Section 5). Neither step tests the causal linkage: stated preferences and expert opinion are assumed to predict actual data-collection outcomes such as response rates, annotation completeness, label reliability, or downstream model accuracy. Section 7 concedes the participant sample is well-educated, tech-savvy, and culturally homogeneous, limiting the input evidence itself. This is load-bearing because the guidelines include specific, potentially costly prescriptions (G1 screening, G5 psychosocial profiling, G9 adaptable annotation, G10 multi-source assessments) whose benefits have never been measured against a baseline. Without a field trial or a post-hoc evaluation on an existing deployment, the abstract's claim that AnnoSense can 'enhance the collection and analysis of emotion data in real-world contexts' is an assertion rather than an established result.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces AnnoSense, a framework of 15 action-oriented guidelines for collecting and annotating physiological emotion data in everyday settings for AI. The guidelines are derived from qualitative data gathered from 75 survey respondents, 32 semi-structured interviews with members of the public, and 3 focus groups with 12 mental health professionals. The framework spans pre-, during-, and post-data collection phases. It is subsequently evaluated by 25 emotion-AI experts who rated each guideline on clarity, usefulness, and adaptability via 5-point Likert scales. The paper also proposes prototype interface directions and discusses implications for future emotion-AI data collection.","tokens_in":45337,"tokens_out":10390,"duration_ms":121251,"significance":"If positioned appropriately, AnnoSense is a substantive methodological contribution to UbiComp, HCI, and affective computing. The manuscript is unusually transparent for a qualitative study: it reports piloting of instruments, consistency checks, a response-quality table, detailed demographic tables, and stepwise thematic-analysis procedures. The 119-stakeholder corpus is substantial for a guideline-generation study, and most guidelines are traceable to quoted participant or expert statements. The expert evaluation, although based on expert perception rather than outcome measurement, is a defensible heuristic-evaluation step. The main gap is that the framework's practical payoff—whether following AnnoSense improves annotation completeness, label reliability, participant compliance, or downstream model performance—is not measured, so the paper's claims of 'evaluated guidelines' and of 'enhancing' data collection need to be carefully scoped.","major_comments":[{"comment":"The evaluation consists solely of 25 experts' Likert ratings of guideline clarity, usefulness, and adaptability; there is no deployment, no baseline protocol, and no outcome measure such as annotation completeness, response rate, label reliability, or downstream AI model performance. Section 7, the Limitations section, discusses demographic and cultural sample limitations but does not state that the framework's effect on actual data-collection outcomes was never evaluated. Because the abstract and contributions say AnnoSense can 'enhance the collection and analysis of emotion data,' this missing causal link is load-bearing for the central claim of an 'evaluated' framework. My recommendation is to either (a) explicitly narrow the contribution to 'expert-evaluated' and add a limitations paragraph stating that the framework's causal impact on data quality has not been established, or (b) add a small field trial, a baseline comparison, or a post-hoc analysis of an existing deployment.","section":"Sections 5 and 7"},{"comment":"The paper reports only an aggregated figure for the expert ratings and asserts that the vast majority of ratings were 'Good' or 'Excellent.' No per-guideline descriptive statistics are provided, so the reader cannot verify uniformity of the positive reception, examine variability across the three phases, or identify weak guidelines. Please include a table with counts, percentages, or mean/SD for each of the 15 guidelines on clarity, usefulness, and adaptability, and note which guidelines, if any, received more than a trivial number of 'Average' or 'Fair' ratings.","section":"Section 5, Figure 3"}],"minor_comments":[{"comment":"The text states that the average completion rate for required open-ended questions was '85%,' but the values in Table 14 (Q4, Q5, Q7, Q8, Q9, Q10, Q11, Q12) average to 86.7%. Also, Q4 is listed as required yet has a completion rate of only 54.7%; please reconcile these numbers or explain why this required question had such a low completion rate.","section":"Section 3.1 and Table 14"},{"comment":"The terms 'objective methods' and 'subjective methods' are introduced with nonstandard definitions in Section 1, but the definitions are not repeated where the terms are used in the methodology. Please provide a short reminder or a cross-reference at first use in Sections 3.2 and 4 to avoid confusion with the conventional meaning of 'subjective' as self-report.","section":"Section 1 and Section 3.2"},{"comment":"When reporting percentages such as 27.2% and 19.6% for annotation-method preferences, please state the denominator (N=75) and clarify whether the underlying question was single-choice or multi-choice, as this materially affects how the reader interprets the distribution.","section":"Section 4.2.1, Figure 2"},{"comment":"Dataset names are used inconsistently (e.g., 'GReX' vs. 'G-REx', 'Dairyhelper' vs. 'DiaryHelper'). Please standardize names across text and tables, including the reference list.","section":"Tables 1 and 2 and Section 2.2"},{"comment":"The in-text citation 'Swain et al.' does not match the reference's lead author 'Das Swain' ([33]). Please align the in-text form with the reference entry.","section":"Section 2.2 and References"}],"recommendation":"major_revision","confidential_remarks":"The paper leans heavily on the authors' own prior work (references [137]–[139]) for motivational framing. The authors should explicitly articulate how AnnoSense goes beyond reference [138], which already examines participants' perspectives on physiological emotion data collection. I also suggest that the editorial team consider whether an outcome-oriented validation, even a small one, should be a condition for acceptance of a framework paper that claims to 'enhance' data collection. The current evidence supports 'expert-evaluated,' not 'outcome-validated.'"},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"AnnoSense does something real: it turns 119 stakeholder voices into 15 concrete, phase-organized guidelines for collecting physiological emotion data in the wild. The derivation is traceable — survey and interview themes lead to specific G1-G15 entries — and the methods write-up is unusually thorough for this kind of paper: piloting, consistency checks, open-ended response quality stats, thematic analysis steps. The authors also tell you where their sample is weak, and the limitation paragraph (Section 7) matches what a skeptic would say.\n\nThe soft spot is the evaluation. Twenty-five experts rate the guidelines on clarity, usefulness, and adaptability, and that rating is treated as evidence that AnnoSense improves real-world emotion data collection. It doesn't show that. Stated preferences and Likert scores do not establish that following G5 or G9 boosts annotation quality, response rates, or downstream model accuracy. There is no field trial, no baseline, and no inter-rater reliability on the coding. The stress-test is right that this is the load-bearing gap. But it is a gap of degree: this is a framework proposal, not an effectiveness claim. The genre expects a first-pass evaluation, and stakeholder grounding plus expert feedback is a reasonable first pass. The abstract's 'can enhance' should be softened to 'can inform' or 'may support', and the authors should acknowledge that the value proposition is still a hypothesis.\n\nThe overlap with the authors' earlier participant-perspective paper [138] deserves a sentence explaining what AnnoSense adds beyond it. That said, the citation pattern here is not abusive; self-citation in a research programme is normal.\n\nBottom line: if I needed to design an everyday emotion data collection study, these guidelines would be on the table. That is the real value. The paper is not a proof that the system works, but it is an honest, detailed synthesis of stakeholder needs plus a plausible expert-checked checklist. I'd send it to review; it belongs in the conversation at UbiComp or IMWUT. I'd ask for the causal language to be moderated and an explicit future-work commitment to a deployment comparison. No desk reject.\n\nWho for: affective computing, UbiComp, mobile mental health researchers. Reading group: maybe, if there's appetite for a long methods-focused piece. Cite: yes, as methodological reference. Serious thinker: yes.","headline":"A genuinely new, stakeholder-grounded 15-guideline framework whose main weakness is the expert-opinion evaluation standing in for evidence of real-world impact.","tokens_in":45910,"tokens_out":3363,"would_cite":true,"duration_ms":39874,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"AnnoSense introduces 15 expert-evaluated guidelines for collecting well-annotated emotion data from wearables and phones in everyday life, derived from 119 stakeholders and reviewed by 25 emotion AI experts.","keywords":["emotion annotation","emotion AI","physiological signals","wearable devices","everyday emotion data collection","qualitative stakeholder study","design guidelines","mental health"],"falsifier":"Run the comparison the paper omits: over the same period, two matched groups of participants log their emotions with wearable sensors, one group under a standard ecological-momentary-assessment protocol and one under a protocol built from the AnnoSense guidelines. If the AnnoSense group does not show higher annotation completion, lower dropout, and labels that are at least as reliable and that train downstream emotion-recognition models at least as accurately, the central claim that the framework improves everyday emotion data collection fails. No such trial appears in the paper.","tokens_in":44949,"feed_emoji":"🧠","tokens_out":7973,"duration_ms":82948,"temperature":0.7,"pith_summary":"The paper sets out to fix a bottleneck in emotion AI: algorithms that infer feelings from wearable and phone data are only as good as the labels attached to that data, and current labeling practice, whether lab studies with fixed scales or brief in-the-moment prompts in daily life, captures too little context and wears participants out. AnnoSense is presented as the answer: 15 actionable guidelines, grouped into pre-, during-, and post-data collection phases, that tell researchers how to recruit, train, profile, prompt, support, and validate participants so the resulting annotations are richer and more honest. The guidelines are grounded in 75 survey responses, 32 interviews with members of the public, and 3 focus groups with 12 mental health professionals, then rated by 25 emotion AI experts on clarity, usefulness, and adaptability. What the paper establishes is that such a stakeholder-grounded protocol can be built and receives positive expert review; it does not yet show that following the protocol yields datasets that train better AI models.","feed_headline":"15 rules for gathering better emotion data for AI","feed_subtitle":"Stakeholder interviews and expert ratings back a three-phase protocol for labeling emotions in daily life.","key_machinery":"The carrying object is the guideline set itself: 15 prescriptive items (G1-G15) organized into three phases, pre-data collection (G1-G6), during-data collection (G7-G11), and post-data collection (G12-G15). The argument that makes the framework credible has two moving parts: a qualitative pipeline that converts stakeholder testimony into guidelines, using inductive thematic analysis of 75 surveys, 32 interviews, and 3 focus group discussions; and an evaluation protocol adapted from heuristic evaluation, in which 25 emotion AI experts scored each guideline for clarity, usefulness, and adaptability on a 5-point Likert scale and gave open-ended suggestions used to revise the guidelines.","core_discovery":"On its own terms, the paper's discovery is the AnnoSense framework itself, which it claims is the first set of guidelines tailored specifically to real-world emotion data collection that has been evaluated. The framework's 15 guidelines (G1-G15) organize the entire data-collection pipeline: before collection, select and screen participants (including screening for alexithymia, the difficulty identifying and expressing emotions), obtain informed consent, calibrate devices, train participants in emotion labeling, and build detailed psycho-social and demographic profiles; during collection, preserve participant agency over prompt timing and annotation frequency, use those profiles to trigger participant-aware sampling, offer annotation methods that flex between quick scales and open-ended reflection based on emotional intensity, add multi-perspective assessments from trusted others and extra data streams, and keep participants engaged with learning and support; after collection, handle data securely with review and deletion rights, validate and normalize quality, ground labels holistically by combining qualitative and quantitative signals with psychosocial context, and share findings, limitations, and intended AI uses. Each guideline is tied to observations from the stakeholder data, and the expert evaluation rated the guidelines predominantly Good or Excellent on clarity, usefulness, and adaptability, with no Poor ratings.","pith_inferences":["A head-to-head field trial is the missing test: two matched groups collecting emotion data over the same weeks, one on a standard ecological-momentary-assessment protocol and one on AnnoSense, comparing completion rates, label reliability, dropout, and downstream model performance, would directly test whether the framework delivers its promised benefit.","The framework's context-rich labels imply a data format that today's emotion-recognition benchmarks are not built to consume, so widespread adoption would likely push the field toward new modeling conventions, such as learning from weighted multi-source labels rather than discrete categories.","Because all 119 stakeholders shared one country and a tech-literate profile, the guidelines' portability across cultures and literacy levels remains open; the paper's own limitation section concedes this, and testing G1-G6 in other populations would show which items are universal and which are local.","The framing in which each person's emotions are constructed from personal history and context sits in tension with training a single general model to map physiology to emotion; if the framework works as intended, it may push emotion AI toward personalized models and away from universal labels."],"forward_implications":["Research teams that adopt AnnoSense would screen participants for alexithymia and other conditions before collecting physiological emotion data, which the paper argues will reduce a known source of noisy labels.","Annotation interfaces built to G7-G9 would let participants switch between scale-based and open-ended reporting depending on emotional intensity, producing datasets that combine quick ratings with rich contextual descriptions.","Validation under G13-G14 would move away from single-source labels toward triangulated labels that weight self-reports, physiological signals, and peer or expert input by reliability, with confidence scores attached.","Published datasets following G15 would document intended AI applications and data limitations as standard practice, addressing the reproducibility problems the paper identifies in current lab and real-life datasets.","If AnnoSense becomes a default protocol, emotion data collection shifts from one-way, scale-driven prompts toward a participant-centered model in which agency, training, and support are part of the data pipeline."],"supporting_citations":[{"why":"Barrett's Theory of Constructed Emotion supplies the paper's working definition of emotions as individualized, context-dependent constructions, which motivates context-rich annotation.","marker":"[15]"},{"why":"Braun and Clarke's thematic analysis is the method used to code the interviews and focus groups and to derive the themes behind the guidelines.","marker":"[26]"},{"why":"Amershi et al.'s heuristic evaluation is adapted as the expert evaluation protocol that scores the guidelines for clarity, usefulness, and adaptability.","marker":"[9]"},{"why":"Jarrahi et al.'s principles of data-centric AI inform the screening, quality-validation, and documentation guidelines (G1, G13, G15).","marker":"[59]"},{"why":"WESAD is the canonical lab dataset whose stimulus-driven, self-report-free labeling the paper contrasts with its own participant-centered approach.","marker":"[129]"},{"why":"K-EmoPhone is a real-life wearable emotion dataset whose EMA-based labeling approach the guidelines are designed to improve upon.","marker":"[60]"},{"why":"Gao et al.'s critique of self-report practices in well-being computing supports the paper's claim that participant-level factors and context are missing from current annotation methods.","marker":"[41]"},{"why":"The authors' earlier study of participants' perspectives on physiological emotion data collection is the direct precursor whose findings feed into the derivation of the guidelines.","marker":"[138]"}],"fun_headline_variants":["AnnoSense framework offers 15 rules for real-world emotion data","Expert-evaluated 15 guidelines make emotion data collection smarter","New framework helps AI get real-world emotion data, says 25 experts","AnnoSense: 15 expert-rated rules for everyday emotion AI data"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The framework assumes that what stakeholders say they want, and what experts rate highly on a questionnaire, predicts how well a real data-collection protocol built on those guidelines will work when it is actually deployed in the field.","fun_headline_variants_meta":{"raw":{"variants":["AnnoSense framework offers 15 rules for real-world emotion data","Expert-evaluated 15 guidelines make emotion data collection smarter","New framework helps AI get real-world emotion data, says 25 experts","AnnoSense: 15 expert-rated rules for everyday emotion AI data"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000615,"raw_usage":{"total_tokens":2878,"prompt_tokens":984,"completion_tokens":1894,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":600,"completion_tokens_details":{"reasoning_tokens":1818}},"tokens_in":600,"tokens_out":1894,"duration_ms":15986,"temperature":1.0,"reasoning_tokens":1818,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T16:33:22.167923+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the comparison the paper omits: over the same period, two matched groups of participants log their emotions with wearable sensors, one group under a standard ecological-momentary-assessment protocol and one under a protocol built from the AnnoSense guidelines. If the AnnoSense group does not show higher annotation completion, lower dropout, and labels that are at least as reliable and that train downstream emotion-recognition models at least as accurately, the central claim that the framework improves everyday emotion data collection fails. No such trial appears in the paper.","supporting_citations":[{"cited_title":"Critiquing Self-report Practices for Human Mental and Wellbeing Computing at Ubicomp","cited_arxiv_id":"2311.15496","evidence_quote":"Gao et al.'s critique of self-report practices in well-being computing supports the paper's claim that participant-level factors and context are missing from current annotation methods."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The authors' earlier study of participants' perspectives on physiological emotion data collection is the direct precursor whose findings feed into the derivation of the guidelines."}],"review_version":1}