Pith. sign in

REVIEW 2 major objections 30 references

A content-controlled smartphone protocol can collect privacy-preserving prosodic features at scale by having people read fixed sentences while processing audio only on the device.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-13 23:21 UTC pith:OXKU4UVR

load-bearing objection Field-ready methods package that actually ships: content control + on-device OpenSMILE + delete-raw, evaluated at useful scale; feasibility and speaker signal hold, lexical fidelity is the softest claim. the 2 major comments →

arxiv 2603.17061 v2 pith:OXKU4UVR submitted 2026-03-17 cs.HC eess.AS

Collecting Prosody in the Wild: A Content-Controlled, Privacy-First Smartphone Protocol and Empirical Evaluation

classification cs.HC eess.AS
keywords prosodyspeech data collectionsmartphonesprivacyon-device processingecological momentary assessmentacoustic features
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Everyday speech research is hard because what people say confounds how they say it, raw audio is privacy-sensitive, and recording tasks often get skipped. This paper introduces a smartphone protocol that shows participants fixed, valence-balanced sentences to read aloud, extracts standard acoustic features on the phone, deletes the raw recording immediately, and transmits only the derived features. In a large field deployment (560 participants, 9,877 retained recordings) compliance was good and the resulting features passed basic acoustic quality checks. Diagnostic models recovered speaker sex with about 92% balanced accuracy under participant-blocked evaluation, while prediction of momentary valence and arousal was only modest. The result is a practical, field-ready way to gather controlled prosody without storing identifiable voice audio.

Core claim

The authors claim that a content-controlled, privacy-first smartphone protocol—scripted read-aloud sentences of controlled lexical valence, on-device extraction of standard acoustic feature sets, immediate deletion of raw audio, and transmission of features only—is feasible in everyday life at scale. Deployed with 560 participants and 9,877 retained recordings, it produced good compliance, analyzable prosodic summaries with substantial speaker-level stability, strong sex classification under blocked cross-validation, and weaker prediction of concurrent self-reported affect.

What carries the argument

The content-controlled, privacy-first smartphone protocol: at each evening prompt, participants read three valence-matched sentences (positive, neutral, negative); OpenSMILE extracts eGeMAPS and ComParE features on-device; the raw WAV is deleted at once; only feature vectors leave the phone. It standardizes lexical content (including valence) while capturing delivery variation and removes raw audio from the research pipeline.

Load-bearing premise

The protocol assumes people actually read the displayed sentences as written, which cannot be verified later because the raw audio is deleted immediately.

What would settle it

Re-run the protocol while temporarily retaining raw audio for a held-out subset; if a large share of clips that pass the voicing and HNR filters are paraphrased, silent, or otherwise non-compliant, or if participant-blocked sex classification falls far below the reported ~92% balanced accuracy on a new matched sample, the claim of reliable content-controlled, speaker-informative signal fails.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Scripted read-aloud modules can be added as semantic baselines inside larger in-the-wild speech studies.
  • Prosody panels become practical under strict privacy rules because raw audio never leaves the device.
  • Speaker-stable metrics from the protocol can calibrate analyses of unconstrained daily speech.
  • On-device features carry enough speaker signal for basic demographic recovery under blocked evaluation.
  • Affect prediction stays limited under fixed scripts, so the method is better used as a control baseline than as a primary emotion sensor.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Lightweight on-device keyword checks (without storing audio or text) could close the unverifiable lexical-fidelity gap the authors note.
  • The same design can serve as a private reference track for calibrating passive continuous-audio embeddings collected in parallel.
  • Pairing the scripted module with brief free-speech or acted probes in the same session would quantify how much affective signal content control removes.
  • Because engineered features can still be re-identifying, future deployments may need feature-level anonymization as inference models improve.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 0 minor

Summary. The paper introduces and evaluates a smartphone protocol for collecting prosodic speech data in everyday life that standardizes lexical content (including prompt valence) via scripted read-aloud sentences, extracts eGeMAPS and ComParE features on-device with OpenSMILE, immediately deletes raw audio, and transmits only feature vectors. Deployed in a quota-matched German panel (final N=560; 9,877 recordings after QC), it reports compliance (67.8% initiation; 96.8% completion once started), acoustic diagnostics and mixed-effects condition contrasts, and diagnostic prediction tasks under participant-blocked CV: ~92% balanced accuracy for speaker sex and modest performance for concurrent single-item valence/arousal. The authors position the protocol as a field-ready, privacy-preserving baseline for content-controlled prosody sampling.

Significance. If the feasibility and signal-characterization claims hold, the work supplies a practical, reproducible template that jointly addresses the prosody–semantics confound and raw-audio privacy barriers that currently limit large-scale in-the-wild prosody research. Strengths include the large quota sample, explicit QC rules, FDR-corrected mixed-effects contrasts, participant-blocked RF CV, and planned release of analysis scripts plus the Android pipeline. The sex-classification positive control and non-degenerate acoustic diagnostics after QC give concrete evidence that usable speaker-informative features can be obtained under realistic smartphone conditions. The residual privacy discussion and the framing of the protocol as a within-participant baseline for unconstrained speech are useful for the field.

major comments (2)
  1. §2.1 and §4.2: The central content-control claim (standardizing lexical content, including valence, so that prosody can be isolated from semantics) rests on participants reading the displayed sentences essentially verbatim. Because raw audio is deleted immediately, lexical fidelity cannot be verified post hoc; QC in §3.2 only flags low voicing probability, few/short voiced segments, or non-positive HNR. Occasional paraphrasing, skipping, or disfluency would reintroduce semantic variance. The paper acknowledges this limitation but supplies no quantitative bound or sensitivity analysis on the deviation rate. A concrete estimate (e.g., pilot ASR keyword-match rates, or a small retained-audio validation subset) would substantially strengthen the content-control claim without changing the privacy design.
  2. §3.4: Affect prediction is reported as modest (eGeMAPS R² ≈ 0.03–0.04 for arousal/valence; ComParE similar) with no condition differences. While the authors correctly treat these tasks as diagnostic rather than primary claims, the abstract and introduction still frame the protocol as enabling prosodic analysis of affective states. The manuscript would be clearer if it more sharply separated the strong feasibility/speaker-signal evidence from the weak affect-signal evidence, and if it quantified how much of the modest performance is attributable to single-item EMA reliability versus scripted-read limitations.

Circularity Check

0 steps flagged

No circularity: empirical protocol evaluation against external labels under blocked CV; no derivation reduces to its inputs by construction.

full rationale

This is a methods/protocol paper whose load-bearing claims are feasibility, compliance, acoustic quality after QC, and diagnostic predictive signal from on-device eGeMAPS/ComParE features. Sex classification (~92% balanced accuracy) and affect regression (modest R^{2}/MAE) use self-reported external targets under participant-blocked ten-fold CV; neither quantity is defined by the protocol or fitted parameters. Condition contrasts and ICCs are ordinary mixed-effects summaries of the collected features, not self-definitional. Self-citations ([15] panel protocol, [18] preregistration) document study context and analysis plan; they do not supply uniqueness theorems, forced ansatzes, or load-bearing premises that close a circular chain. Residual privacy discussion and the acknowledged lexical-fidelity limitation (raw audio deleted) are validity caveats, not circular reductions. The evaluation is self-contained against external benchmarks.

Axiom & Free-Parameter Ledger

4 free parameters · 4 axioms · 0 invented entities

The paper is a methods-and-evaluation contribution. Load-bearing elements are standard domain tools (OpenSMILE feature sets, RF defaults, EMA single items) plus operational choices (time windows, QC thresholds, three sentences per valence). No new physical entities are postulated; residual risk is that unverified lexical compliance reintroduces the semantic confound the design aims to remove.

free parameters (4)
  • RF hyperparameters (num.trees=500, mtry=sqrt, min.node.size=5) = 500 / sqrt(p) / 5
    Fixed defaults used for both sex classification and affect regression; performance numbers depend on them though not heavily tuned.
  • Voice-absence QC thresholds (mean voicing probability, voiced segments/s, mean voiced-segment length)
    Used to drop 232 clips; exact cutoffs determine the final N=9,877 analytic sample.
  • HNR > 0 dB robustness filter = > 0 dB
    Excluded 1,108 additional clips; directly shapes retained data quality and downstream prediction.
  • Recording duration bounds (min 4 s, max 12 s) = 4–12 s
    Chosen from extreme reading times of the three-sentence prompts; constrains available speech material.
axioms (4)
  • domain assumption eGeMAPS (88) and ComParE 2016 (6373) features extracted by on-device OpenSMILE are valid, comparable prosodic descriptors across heterogeneous Android devices and environments.
    Invoked throughout §2.2 and all prediction analyses; no device-level calibration study is provided.
  • domain assumption Participants largely read the displayed sentences as instructed, so lexical content (including valence) is effectively standardized.
    Core design premise of content control (§2.1); acknowledged as unverifiable in Limitations §4.2 because raw audio is deleted.
  • domain assumption Self-reported sex and single-item 6-point valence/arousal EMA items are sufficiently reliable targets for diagnostic prediction.
    Used as ground truth in §3.4; single-item affect reliability is known to be limited (cited).
  • domain assumption Immediate local deletion of WAV + CSV and SSL batch transfer of features adequately mitigates re-identification risk for the study’s purposes under GDPR.
    Privacy-first claim in §2.2 and Discussion; residual feature-inference risk is noted but not quantified.

pith-pipeline@v1.1.0-grok45 · 14720 in / 2777 out tokens · 40513 ms · 2026-07-13T23:21:43.509339+00:00 · methodology

0 comments
read the original abstract

Collecting everyday speech data for prosodic analysis is challenging due to the confounding of prosody and semantics, privacy constraints, and participant compliance. We introduce and empirically evaluate a content-controlled, privacy-first smartphone protocol that uses scripted read-aloud sentences to standardize lexical content (including prompt valence) while capturing naturalistic variation in prosodic delivery. The protocol performs on-device prosodic feature extraction, deletes raw audio immediately, and transmits only derived features for analysis. We deployed the protocol in a large study (N = 560; 9,877 recordings), evaluated compliance and data quality, and conducted diagnostic prediction tasks on the extracted features, predicting self-reported speaker sex and momentary affective states (valence, arousal). We discuss implications and directions for advancing and deploying the protocol.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

30 extracted references · 1 linked inside Pith

  1. [1]

    Although this shift enables large- scale and ecologically valid speech sampling, it also introduces two core challenges for prosodic data collection and analy- sis

    Introduction Speech research is increasingly moving from controlled labo- ratory settings to everyday life, for example, leveraging stan- dard smartphones [1, 2]. Although this shift enables large- scale and ecologically valid speech sampling, it also introduces two core challenges for prosodic data collection and analy- sis. First, real-world recordings ...

  2. [2]

    positive

    Protocol 2.1. V oice Recording The protocol was implemented as a module within thePhoneS- tudyapp’s smartphone-based ecological momentary assessment (EMA) procedure, which prompted participants multiple times per day. At the start of each prompt, participants were shown a brief introductory screen describing the voice-recording task. Here, they also had t...

  3. [3]

    How do you feel right now?

    Evaluation 3.1. Data Collection The protocol was implemented in a data collection that was part of a large panel study [15]. Data collection was approved by the ethics committee of the psychology department at LMU Mu- nich and all procedures adhered to the General Data Protection Regulation (GDPR). We recruited a quota-matched sample of N = 850 participan...

  4. [4]

    Discussion 4.1. Protocol Contributions The primary contribution of this work is a field-ready protocol for collecting prosodic speech data in everyday life that jointly addresses two challenges in naturalistic speech research: con- founding between semantic content and prosody and the privacy challenges associated with raw audio collection. The protocol s...

  5. [5]

    Conclusion We presented a content-controlled, privacy-first smartphone protocol for collecting prosodic speech data in everyday life and empirically evaluated it. By standardizing lexical content, performing on-device acoustic feature extraction, and deleting raw audio immediately, the protocol enables scalable, privacy- preserving prosodic data collectio...

  6. [6]

    This project was supported by the Swiss National Science Founda- tion (SNSF) under project number 215303 and a scholarship of the German Academic Scholarship foundation

    Acknowledgments We thank audEERING GmbH, Peter Ehrich, and Dominik Heinrich for their support with the technical implementation of the on-device voice feature extraction and the Leibniz In- stitute for Psychology (ZPID) for funding data collection. This project was supported by the Swiss National Science Founda- tion (SNSF) under project number 215303 and...

  7. [7]

    Fabla: A voice-based ecological as- sessment method for securely collecting spoken responses to re- searcher questions,

    D. M. Kaplan, S. J. A. Alvarez, R. Palitsky, H. Choi, G. D. Clif- ford, M. Crozier, B. W. Dunlop, G. H. Grant, M. N. Greenleaf, L. M. Johnson, J. Maples-Keller, H. F. Levin-Aspenson, J. S. Mas- caro, A. McDowall, N. S. Pozzo, C. L. Raison, A. J. Zarrabi, B. O. Rothbaum, and W. A. Lam, “Fabla: A voice-based ecological as- sessment method for securely colle...

  8. [8]

    The PRIORI Emotion Dataset: Linking Mood to Emotion Detected In-the-Wild,

    S. Khorram, M. Jaiswal, J. Gideon, M. McInnis, and E. Mower Provost, “The PRIORI Emotion Dataset: Linking Mood to Emotion Detected In-the-Wild,” inInterspeech 2018. ISCA, Sep. 2018, pp. 1903–1907

  9. [9]

    When emotional prosody and se- mantics dance cheek to cheek: ERP evidence,

    S. A. Kotz and S. Paulmann, “When emotional prosody and se- mantics dance cheek to cheek: ERP evidence,”Brain Research, vol. 1151, pp. 107–118, Jun. 2007

  10. [10]

    Emotional Speech Processing at the Intersection of Prosody and Semantics,

    R. Schwartz and M. D. Pell, “Emotional Speech Processing at the Intersection of Prosody and Semantics,”PLoS ONE, vol. 7, no. 10, p. e47279, Oct. 2012

  11. [11]

    IEEE recommended practice for speech quality mea- surements,

    IEEE, “IEEE recommended practice for speech quality mea- surements,”IEEE Transactions on Audio and Electroacoustics, vol. 17, no. 3, pp. 225–246, Sep. 1969

  12. [12]

    IEMOCAP: Interactive emotional dyadic motion capture database,

    C. Busso, M. Bulut, C.-C. Lee, A. Kazemzadeh, E. Mower, S. Kim, J. N. Chang, S. Lee, and S. S. Narayanan, “IEMOCAP: Interactive emotional dyadic motion capture database,”Language Resources and Evaluation, vol. 42, no. 4, pp. 335–359, Dec. 2008

  13. [13]

    (Not) hearing happiness: Predicting fluctuations in happy mood from acoustic cues using machine learning,

    A. C. Weidman, J. Sun, S. Vazire, J. Quoidbach, L. H. Ungar, and E. W. Dunn, “(Not) hearing happiness: Predicting fluctuations in happy mood from acoustic cues using machine learning,”Emotion (Washington, D.C.), vol. 20, no. 4, pp. 642–658, Jun. 2020

  14. [14]

    Towards privacy- preserving conversation analysis in everyday life: Exploring the privacy-utility trade-off,

    J. Pohlhausen, F. Nespoli, and J. Bitzer, “Towards privacy- preserving conversation analysis in everyday life: Exploring the privacy-utility trade-off,”Computer Speech & Language, vol. 95, p. 101823, Jan. 2026

  15. [15]

    On- Device Speech Filtering for Privacy-Preserving Acoustic Activ- ity Recognition,

    H. Zhou, S. Boovaraghavan, M. Goel, and Y . Agarwal, “On- Device Speech Filtering for Privacy-Preserving Acoustic Activ- ity Recognition,” inProceedings of the 30th Annual International Conference on Mobile Computing and Networking. Washington D.C. DC USA: ACM, Dec. 2024, pp. 1802–1804

  16. [16]

    FRILL: A Non-Semantic Speech Embedding for Mobile De- vices,

    J. Peplinski, J. Shor, S. Joglekar, J. Garrison, and S. Patel, “FRILL: A Non-Semantic Speech Embedding for Mobile De- vices,” inInterspeech 2021. ISCA, Aug. 2021, pp. 1204–1208

  17. [17]

    Emotional Speech Perception: A set of semantically validated German neutral and emotionally affective sentences,

    S. Defren, P. de Brito Castilho Wesseling, S. Allen, V . Shakuf, B. Ben-David, and T. Lachmann, “Emotional Speech Perception: A set of semantically validated German neutral and emotionally affective sentences,” in9th International Conference on Speech Prosody 2018. ISCA, Jun. 2018, pp. 714–718

  18. [18]

    openSMILE: The Mu- nich versatile and fast open-source audio feature extractor,

    F. Eyben, M. W ¨ollmer, and B. Schuller, “openSMILE: The Mu- nich versatile and fast open-source audio feature extractor,” in Proceedings of the 18th ACM International Conference on Multi- media (MM ’10). Firenze, Italy: ACM, 2010, pp. 1459–1462

  19. [19]

    The Geneva Minimalistic Acoustic Parameter Set (GeMAPS) for V oice Research and Affective Computing,

    F. Eyben, K. R. Scherer, B. Schuller, J. Sundberg, E. Andre, C. Busso, L. Y . Devillers, J. Epps, P. Laukka, S. S. Narayanan, and K. P. Truong, “The Geneva Minimalistic Acoustic Parameter Set (GeMAPS) for V oice Research and Affective Computing,”IEEE Transactions on Affective Computing, vol. 7, no. 2, pp. 190–202, Apr. 2016

  20. [20]

    The INTERSPEECH 2016 computational paralinguistics challenge: Deception, sincerity & native language,

    B. Schuller, S. Steidl, A. Batliner, J. Hirschberg, J. K. Burgoon, A. Baird, A. Elkins, Y . Zhang, E. Coutinho, and K. Evanini, “The INTERSPEECH 2016 computational paralinguistics challenge: Deception, sincerity & native language,” inProc. INTERSPEECH 2016, San Francisco, CA, USA, Sep. 2016, pp. 2001–2005

  21. [21]

    Basic Protocol: Smartphone Sensing Panel Study,

    R. Schoedel and M. Oldemeier, “Basic Protocol: Smartphone Sensing Panel Study,”PsychArchives, 2024

  22. [22]

    Remote Smartphone-Based Speech Col- lection: Acceptance and Barriers in Individuals with Major De- pressive Disorder,

    J. Dineley, G. Lavelle, D. Leightley, F. Matcham, S. Siddi, M. T. Pe˜narrubia-Mar´ıa, K. M. White, A. Ivan, C. Oetzmann, S. Sim- blett, E. Dawe-Lane, S. Bruce, D. Stahl, Y . Ranjan, Z. Rashid, P. Conde, A. A. Folarin, J. M. Haro, T. Wykes, R. J. Dobson, V . A. Narayan, M. Hotopf, B. W. Schuller, N. Cummins, and The Radar-Cns Consortium, “Remote Smartphone...

  23. [23]

    Accurate short-term analysis of the fundamental fre- quency and the harmonics-to-noise ratio of a sampled sound,

    P. Boersma, “Accurate short-term analysis of the fundamental fre- quency and the harmonics-to-noise ratio of a sampled sound,” in Proceedings of the Institute of Phonetic Sciences, vol. 17, 1993, pp. 97–110

  24. [24]

    Predicting Affective States from Acoustic V oice Cues Collected with Smartphones,

    T. Koch and R. Schoedel, “Predicting Affective States from Acoustic V oice Cues Collected with Smartphones,”Psy- chArchives, 2021

  25. [25]

    A Circumplex Model of Affect,

    J. Russell, “A Circumplex Model of Affect,”Journal of Personal- ity and Social Psychology, vol. 39, pp. 1161–1178, Dec. 1980

  26. [26]

    Acoustic profiles in vocal emo- tion expression,

    R. Banse and K. R. Scherer, “Acoustic profiles in vocal emo- tion expression,”Journal of Personality and Social Psychology, vol. 70, no. 3, pp. 614–636, 1996

  27. [27]

    V oice analytics in the wild: Validity and predictive accuracy of common audio- recording devices,

    F. Busquet, F. Efthymiou, and C. Hildebrand, “V oice analytics in the wild: Validity and predictive accuracy of common audio- recording devices,”Behavior Research Methods, vol. 56, no. 3, pp. 2114–2134, Mar. 2024

  28. [28]

    Assessing the reliability of single-item momentary affective measurements in experience sampling

    E. Dejonckheere, F. Demeyer, B. Geusens, M. Piot, F. Tuer- linckx, S. Verdonck, and M. Mestdagh, “Assessing the reliability of single-item momentary affective measurements in experience sampling.”Psychological Assessment, vol. 34, no. 12, pp. 1138– 1154, Dec. 2022

  29. [29]

    The V oicePrivacy 2020 Challenge: Results and findings,

    N. Tomashenko, X. Wang, E. Vincent, J. Patino, B. M. L. Srivas- tava, P.-G. No´e, A. Nautsch, N. Evans, J. Yamagishi, B. O’Brien, A. Chanclu, J.-F. Bonastre, M. Todisco, and M. Maouche, “The V oicePrivacy 2020 Challenge: Results and findings,”Computer Speech & Language, vol. 74, p. 101362, Jul. 2022

  30. [30]

    Convolutional neural networks for small-footprint keyword spotting,

    T. N. Sainath and C. Parada, “Convolutional neural networks for small-footprint keyword spotting,” inInterspeech 2015. ISCA, Sep. 2015, pp. 1478–1482