Pith. sign in

REVIEW 2 major objections 6 minor 49 references

Umm... With Transformers? Insights from Filled Pause Use across Four Slavic Parliaments

T0 review · 2 major / 6 minor · reviewed 2026-07-08 · glm-5.2

Pith's one-line read Women say 'um' more in South Slavic parliaments, upending prior findings

desk verdict Large-scale FP study with a real measurement-confounding gap on its headline gender-reversal finding read the letter →

arxiv 2607.05964 v1 pith:CCNVGLMS submitted 2026-07-07 cs.CL stat.AP

classification cs.CLstat.AP
keywords filledpausesparliamentaryspeechSlaviclanguagesdisfluencywithin-speakervariationMundlakcorrectionGEEsociolinguistics
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper analyzes roughly 4,000 hours of parliamentary speech across Croatian, Czech, Polish, and Serbian to test whether patterns of filled-pause (um, uh) production hold up at scale and across languages. Using a transformer-based detector for filled pauses and a regression technique that separates stable speaker traits from utterance-level variation, the authors argue that many previously reported patterns are not universal. The gender effect found in English-language conversational corpora — men producing more filled pauses than women — reverses direction in South Slavic parliaments (women produce more) and disappears entirely in West Slavic ones. Age shows a negative association with filled-pause rate, also contrary to some prior work. Speech rate replicates the expected inverse relationship, and crucially the within-speaker component dwarfs the between-speaker component, suggesting filled pauses reflect moment-to-moment planning demands rather than habitual speaking style. A novel finding is that positive sentiment consistently predicts higher filled-pause rates within individual speakers across all four parliaments. Opposition-party status is globally associated with fewer filled pauses than governing-coalition status, though this holds clearly in only two of four parliaments.

What carries the argument

Mundlak-corrected GEE (Generalised Estimating Equations) with Negative Binomial outcome, applied to transformer-detected (wav2vec2-bert) filled-pause counts across utterances nested within speakers.

What would settle it

If the filled-pause detector systematically over-detects pauses in female voices or in positive-sentiment speech, the gender reversal and sentiment–FP association would be measurement artifacts rather than linguistic phenomena.

Watch

Extended reading notes

Core claim

The paper's central claim is that filled-pause production is governed by domain-specific and sociolinguistic factors rather than universal cognitive patterns. The key methodological mechanism is the Mundlak correction, which decomposes each time-varying predictor into a speaker's average value (between-speaker, trait-level) and their deviation from that average in a given utterance (within-speaker, state-level). This decomposition reveals that the speech-rate effect is primarily a within-speaker phenomenon — the same speaker produces fewer filled pauses when speaking faster than their own baseline — supporting a planning-demand account over a habitual-style account. The gender reversal in Sl

Load-bearing premise

The automated filled-pause detector (F1≈0.92) and sentiment model (R²≈0.65) are assumed to introduce no systematic bias correlated with the predictor variables — for example, the detector is assumed not to produce more false positives for female voices or for positive-sentiment speech. If such biases exist, the reported associations could be artifacts of the measurement pipeline rather than genuine linguistic patterns.

Editorial extensions

If this is right

  • If filled-pause patterns are domain-specific rather than universal, meta-analyses pooling conversational and formal speech data may produce misleading generalizations.
  • The within-speaker dominance of the speech-rate effect suggests that filled pauses index real-time cognitive load, which could make them a useful proxy for planning difficulty in speech-production research.
  • The sentiment–filled-pause link, if robust, would suggest that positive affective states increase planning demands or reduce fluency monitoring, a direction not well explored in prior disfluency literature.
  • Opposition speakers producing fewer filled pauses could reflect preparation differences (opposition speeches may be more rehearsed) rather than cognitive differences, opening a testable question about speech preparation and disfluency.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 6 minor

Summary. This paper analyses filled-pause (FP) production across ~4,000 hours of parliamentary speech in four Slavic languages (Croatian, Czech, Polish, Serbian), using a wav2vec2-bert-based FP detector and an XLM-R-based sentiment model. FP counts are modelled with Negative Binomial GEE regression with Mundlak correction to separate within- from between-speaker effects. The authors replicate negative associations of age and speech rate with FP rate, report a novel gender reversal in South Slavic parliaments (women producing more FPs), find a consistent positive within-speaker sentiment–FP association, and report that opposition speakers tend to produce fewer FPs. The modelling framework is methodologically careful and the scale of the study is impressive.

Significance. The paper makes a genuine contribution by bringing large-scale, cross-linguistic, automatically annotated parliamentary data to bear on filled-pause research, which has historically relied on small single-language corpora. The use of Mundlak-corrected GEE to decompose within- and between-speaker effects is a methodological strength that goes beyond standard regression approaches in this area. The gender-reversal finding, if robust, would be a substantive and surprising result for the disfluency literature. The transparent treatment of exploratory versus confirmatory analyses and the honest reporting of parliament-specific inconsistencies are commendable.

major comments (2)
  1. §5.1, Gender finding: The paper's most novel claim — that women produce more FPs in South Slavic parliaments (HR: IRR≈0.52, RS: IRR≈0.40 for male speakers), reversing prior literature — rests on the assumption that the wav2vec2-bert FP detector has no gender-correlated error rates. The detector processes acoustic features (F0, formant spacing, spectral characteristics) that differ systematically between male and female voices, and it was trained on Slovenian data (a South Slavic language, like Croatian and Serbian). The paper reports only aggregate event-level F1≈0.92 across all four languages (§3), with no per-gender or per-language breakdown of precision and recall. If the detector has even a modest gender-correlated bias in error rates — and such bias could plausibly be language-specific given the Slovenian training data and the South Slavic test languages — the large sample (~1Mutter
  2. §3: The claim that both automated tools have performance 'comparable to the inter-annotator agreement upper bound' is stated without elaboration. For the FP detector, the cited work [9] reports event-level F1 of 0.87–0.94, but the paper does not specify which metric is being compared to inter-annotator agreement (F1? precision? recall?) or what the inter-annotator agreement upper bound actually is. Given that the detector reportedly 'surpasses human recall, albeit with a minor hit on precision' (§2.1), the asymmetry between precision and recall is itself relevant: if recall is high but precision is lower, false positives could inflate FP counts in ways that interact with speaker characteristics. The paper should report precision and recall separately, ideally disaggregated by language and gender, or at minimum discuss this asymmetry as a potential confound.
minor comments (6)
  1. §5.5: The within-speaker political orientation effect in HR (IRR=1.386, +38.6% per orientation point) is very large relative to other effects in the paper. The authors note this reflects 'relatively infrequent episodes in which speakers switch parties,' but no count of such episodes is given. Reporting the number of speakers/utterances contributing to this estimate would help readers gauge its reliability.
  2. §5.6: The Serbian power-status reversal (IRR=1.534, p=.061) is described as possibly reflecting 'particular political dynamics of the Serbian parliament.' This is speculative; a brief description of what those dynamics might be would strengthen the discussion.
  3. Figure 1 is referenced as showing IRR results across all models but is not legible in the reviewed version. Ensure that coefficient labels, confidence interval values, and the distinction between baseline and Mundlak-corrected specifications are clearly readable in the final version.
  4. §4.2, Eq. (1): The model specification lists 'C(Parliament)' as a covariate, but it is unclear whether this is included only in the global model or also in per-parliament models. Clarifying this would prevent confusion.
  5. §2.1: The phrase 'This performance significantly surpasses human recall, albeit with a minor hit on precision' could be more precisely stated with the actual human recall and precision figures for comparison.
  6. The paper would benefit from a brief discussion of the base FP rates (1.38–3.47 FPs/min, §5) in relation to rates reported in prior literature, to contextualise whether the parliamentary domain produces systematically different FP densities.

Simulated Author's Rebuttal

2 responses · 0 unresolved

We thank the referee for a careful and constructive reading of our manuscript. The two major comments both concern the reliability of our automatic FP detector, specifically whether gender-correlated or language-specific biases in detection error could confound our findings — particularly the gender-reversal result. These are well-taken points, and we address each below.

read point-by-point responses
  1. Referee: §5.1, Gender finding: The paper's most novel claim — that women produce more FPs in South Slavic parliaments — rests on the assumption that the wav2vec2-bert FP detector has no gender-correlated error rates. The detector processes acoustic features that differ systematically between male and female voices, and it was trained on Slovenian data (a South Slavic language, like Croatian and Serbian). The paper reports only aggregate event-level F1≈0.92, with no per-gender or per-language breakdown of precision and recall. If the detector has even a modest gender-correlated bias, the large sample could amplify it.

    Authors: The referee raises a legitimate and important concern. We agree that a gender-correlated bias in the FP detector is a plausible confound for our gender findings, and we should have addressed this explicitly in the manuscript. We will revise the manuscript to incorporate the following points and analyses. First, we note that the detector was trained on Slovenian data, which is South Slavic and closely related to Croatian and Serbian — the very languages where the gender reversal appears. This means the detector should, if anything, perform best on those languages, which somewhat complicates the bias story: if the detector were systematically biased against male voices due to training-data composition, we would expect this to manifest as a language-general artifact rather than one specific to South Slavic parliaments. However, we acknowledge that this argument is indirect and does not rule out a gender-correlated bias that interacts with language-specific acoustic properties. Second, we will conduct and report a per-gender evaluation of the detector on the available manually annotated test sets from [9], disaggregating precision and recall by gender and by language. If the annotated subsets contain sufficient gender balance, this will directly test whether the detector's error rates differ by gender. We will report these results in the revised manuscript. Third, regardless of the outcome of that analysis, we will add an explicit limitation paragraph noting that gender-correlated detector bias remains a possible confound and that the gender-reversal finding should be interpreted with appropriate caution until validated against manually annotated FP data with balanced gender representation. We agree that the current manuscript overstates the robustness of this finding bynot revision: yes

  2. Referee: §3: The claim that both automated tools have performance 'comparable to the inter-annotator agreement upper bound' is stated without elaboration. For the FP detector, the cited work [9] reports event-level F1 of 0.87–0.94, but the paper does not specify which metric is being compared to inter-annotator agreement, or what the IAA upper bound actually is. The paper should report precision and recall separately, ideally disaggregated by language and gender, or at minimum discuss the precision-recall asymmetry as a potential confound.

    Authors: The referee is correct that our statement in §3 is underspecified. We will revise the manuscript to (a) state explicitly what the inter-annotator agreement upper bound is and which metric it refers to, drawing on the detailed results in [9]; (b) report precision and recall separately for the FP detector, rather than only aggregate F1; and (c) disaggregate these by language where the annotated test sets permit. We will also add a brief discussion of the precision-recall asymmetry the referee highlights: the detector reportedly has higher recall than precision, meaning false positives are the more common error mode. If false positives are distributed unevenly across speaker groups (e.g., if certain voice types are more prone to false-positive FP detections), this could inflate or deflate FP counts in ways that interact with our predictors. We will acknowledge this as a potential confound and note that our per-gender evaluation (described in the response to the first comment) will partially address this concern for the gender variable specifically. For the sentiment model, we will likewise clarify what 'comparable to inter-annotator agreement' means — specifically, that the R²≈0.65 of XLM-R-ParlaSent is in the range of agreement between human annotators on the same parliamentary sentiment data, as reported in [37]. revision: yes

Circularity Check

0 steps flagged · score 1.0 of 10

No significant circularity; self-citations are to tools used as instruments, not to results being validated

full rationale

The paper's central findings (gender reversal in South Slavic parliaments, sentiment-FP association, power-status effect) are outputs of GEE Negative Binomial regressions fitted to detected FP counts and speaker metadata. None of these results reduce to their inputs by construction. The two automated tools cited — the wav2vec2-bert FP detector [9] and XLM-R-ParlaSent sentiment model [37] — are co-authored by the present authors, but they are used as measurement instruments with externally reported performance metrics (F1≈0.92, R²≈0.65), not as results being circularly validated in this paper. The ParlaSpeech dataset [36] is similarly a resource, not a claim. The GEE models with Mundlak decomposition are standard econometric techniques [38, 39] applied to the data; the IRRs are genuine regression estimates, not definitions or fits renamed as predictions. The self-citations are therefore of the normal, non-load-bearing kind: the authors use tools they built, report their known performance, and then ask new empirical questions with them. The skeptic's concern about potential gender-correlated detector bias is a validity threat (correctness risk), not a circularity problem — the regression outputs are not equivalent to the detector inputs by construction. Score 1 reflects the presence of self-cited tools that are integral to the pipeline but are treated as external instruments with stated assumptions, not as self-validating results.

Assumptions & free parameters 2 free parameters · 5 assumptions · 0 invented entities

The paper introduces no new entities, particles, forces, or postulated constructs. It uses existing statistical methods and automated detection tools applied to an existing dataset. The free parameters are standard regression coefficients estimated from data. The axioms are domain assumptions about tool accuracy and domain validity, plus standard statistical methodology.

free parameters (2)
  • GEE regression coefficients (IRR values) = Various (e.g., gender IRR=0.636, age IRR=0.862, speech rate IRR=0.645, sentiment IRR=1.060, power status IRR=0.792)
    These are the empirical estimates from the GEE models fitted to the data; they are the paper's primary output.
  • Negative Binomial overdispersion parameter = Not reported explicitly
    The Negative Binomial model requires an overdispersion parameter that is estimated from the data but not reported in the paper.
assumptions (5)
  • domain assumption The wav2vec2-bert FP detection model (event-level F1≈0.92) produces detections that are sufficiently accurate and unbiased for regression analysis.
    Stated in §3 and §2.1 (citing [9]). The paper assumes detection errors do not systematically confound the predictor-outcome relationships.
  • domain assumption The XLM-R-ParlaSent sentiment model (R²≈0.65) provides sentiment scores of sufficient quality for regression analysis.
    Stated in §3 (citing [37]). Moderate R² implies substantial noise in sentiment measurement.
  • domain assumption Parliamentary speech is a valid domain for studying general filled-pause production patterns, despite its formal register.
    Stated in §1 and acknowledged as a limitation in §7. The paper's claims are bounded to this domain.
  • standard math GEE with independent working correlation and robust sandwich standard errors yields valid population-average inference under the study's clustering structure (utterances nested within speeches within speakers).
    Standard statistical methodology (§4.1, citing [38]). The robustness of sandwich SEs to working correlation misspecification is a well-established result.
  • standard math The Mundlak correction (decomposing time-varying predictors into speaker means and deviations) adequately reduces omitted variable bias from fixed speaker traits.
    Standard econometric technique (§4.2, citing [39]). The adequacy depends on the assumption that remaining confounders are not correlated with the included predictors.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Umm... With Transformers? Insights from Filled Pause Use across Four Slavic Parliaments." pith.science (2026). https://pith.science/paper/CCNVGLMS

@misc{pith2026260705964,
  author       = {Pith},
  title        = {Pith review of: Umm... With Transformers? Insights from Filled Pause Use across Four Slavic Parliaments},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/CCNVGLMS}},
  note         = {Machine review of arXiv:2607.05964}
}
read the original abstract

Filled pauses (FPs) are a universal feature of spontaneous speech, yet most studies rely on small, single-language corpora, limiting the generalisability of their findings. We analyse ~4,000 hours of parliamentary speech across four related Slavic languages (Croatian, Czech, Polish, Serbian). FP occurrence is obtained via transformer-based automatic detection, while FP rate is modelled using Generalised Estimating Equations (GEE) with Mundlak correction to distinguish within- from between- speaker effects. We replicate a negative association of age and speech rate with FP rate, but find that gender effects are language-specific and directionally opposite to most prior literature. Novel analyses of sentiment, political orientation, and power status reveal a consistent positive association between sentiment and FP rate, alongside parliament-specific modulation by orientation and power status, with opposition speakers tending toward lower FP rates than governing coalition speakers.

Figures

Figures reproduced from arXiv: 2607.05964 by the authors.

Figure 1
Figure 1. Incidence Rate Ratios (IRR) from GEE Negative Binomial models with independent working correlation structure, comparing baseline (no Mundlak correction; blue) and Mundlak-corrected (orange) specifications across four parliaments and a pooled global model. The Mundlak decomposition partitions time-varying predictors into within-speaker deviation and between-speaker mean com￾ponents. Error bars show 95% confidence int… view at source ↗

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

49 extracted references · 49 canonical work pages

  1. [9]

    Generative AI Use Disclosure During the preparation of this manuscript, generative AI was used for improving the grammar and style of specific sections of the manuscript

  2. [1]

    Introduction Filled pauses — vocalisations such asuhandum— are a per- vasive feature of spontaneous speech, serving functions ranging from turn-holding to signalling lexical difficulty [1]. Despite sustained research interest, most empirical work relies on small, single-language corpora, which limits both statistical power and the generalisability of conc...

  3. [2]

    Umm... With Transformers? Insights from Filled Pause Use across Four Slavic Parliaments

    Related Work 2.1. Automatic FP Identification Early filled-pause detection relied on hand-crafted acoustic cues (pitch/F0, MFCCs, formants, spectral/vocal-tract stability) with modest precision and recall [2, 3, 4]. Prosodic-feature methods improved results (0.61 F1), and prosodic discontinuity features reached 0.83 F1 in spontaneous speech [5, 6]. Recent...

  4. [3]

    Speaker metadata (gender, age, political orientation, party power status) were taken from official parliamentary records

    Data We use the ParlaSpeech dataset 1 [36] of 6,000 hours of parliamentary speech from the Croatian (HR, 2015-2022), Czech (CZ, 2013-2023), Polish (PL, 2017-2022), and Ser- bian (RS, 2013-2022) parliaments, with aligned official tran- scripts. Speaker metadata (gender, age, political orientation, party power status) were taken from official parliamentary ...

  5. [4]

    Model Having detected FPs automatically, we model theirrate of occurrenceacross utterances to quantify how speaker- and utterance-level factors relate to FP production

    Method 4.1. Model Having detected FPs automatically, we model theirrate of occurrenceacross utterances to quantify how speaker- and utterance-level factors relate to FP production. We treat FP count per utterance as a Negative Binomial outcome, chosen for its robustness to overdispersion typical of count data, with 1https://clarinsi.github.io/parlaspeech/...

  6. [5]

    Results We first report general incidence of FPs in our data. HR and RS show substantially lower baseline FP rates (1.82 and 1.38 FPs/min, respectively) than CZ and PL (2.91 and 3.47 FPs/min), motivating the inclusion of parliament as a covariate in the global model and suggesting cross-linguistic differences in dis- fluency norms or parliamentary speech ...

  7. [6]

    For established predictors, we replicate a negative effect of age and confirm the well-documented inverse relationship be- tween speech rate and FP rate

    Conclusion We presented a large-scale cross-linguistic analysis of filled- pause production across almost 4,000 hours of parliamentary speech in four Slavic languages, combining transformer-based automatic FP detection with Mundlak-corrected GEE mod- elling to separate stable speaker traits from utterance-level vari- ation. For established predictors, we ...

  8. [7]

    Future Work The domain restriction to parliamentary speech limits general- isation; extending this framework to conversational or broad- cast data would test which effects are robust across registers. The gender reversal in South Slavic parliaments, and the consis- tent behaviour regarding power status with Serbia as a potential anomaly, both warrant targ...

Show all 49 references
  1. [8]

    Large Language Models for Digital Humanities

    Acknowledgments This work was supported in part by the project “Large Language Models for Digital Humanities” (GC-0002), the project “EPIC- SI - Early Parent-Child Communication in Slovenian: Corpus- based Insights” (J6-70222), the research programme “Language Resources and Te...

  2. [10]

    Using uh and um in spontaneous speaking,

    H. H. Clark and J. E. Fox Tree, “Using uh and um in spontaneous speaking,”Cognition, vol. 84, no. 1, pp. 73–111, 2002

  3. [11]

    A real-time filled pause de- tection system for spontaneous speech recognition,

    M. Goto, K. Itou, and S. Hayamizu, “A real-time filled pause de- tection system for spontaneous speech recognition,” inSixth Eu- ropean Conference on Speech Communication and Technology, 1999

  4. [12]

    A feature-based filled pause de- tection system for Dutch,

    F. Stouten and J.-P. Martens, “A feature-based filled pause de- tection system for Dutch,” in2003 IEEE Workshop on Auto- matic Speech Recognition and Understanding (IEEE Cat. No. 03EX721). IEEE, 2003, pp. 309–314

  5. [13]

    Formant-based technique for automatic filled-pause detection in spontaneous spoken English,

    K. Audhkhasi, K. Kandhway, O. D. Deshmukh, and A. Verma, “Formant-based technique for automatic filled-pause detection in spontaneous spoken English,” in2009 IEEE International Confer- ence on Acoustics, Speech and Signal Processing. IEEE, 2009, pp. 4857–4860

  6. [14]

    Experiments on automatic detection of filled pauses us- ing prosodic features,

    H. Medeiros, H. Moniz, F. Batista, I. Trancoso, H. Meinedo et al., “Experiments on automatic detection of filled pauses us- ing prosodic features,”Actas de Inforum, vol. 2013, pp. 335–345, 2013

  7. [15]

    Filled pause detection by prosodic discontinuity features,

    U. D. Reichel, B. Weiss, and T. Michael, “Filled pause detection by prosodic discontinuity features,”Studientexte zur Sprachkom- munikation: Elektronische Sprachsignalverarbeitung 2019, pp. 272–279, 2019

  8. [16]

    Automatic disflu- ency detection from untranscribed speech,

    A. Romana, K. Koishida, and E. M. Provost, “Automatic disflu- ency detection from untranscribed speech,”IEEE/ACM transac- tions on audio, speech, and language processing, vol. 32, pp. 4727–4740, 2024

  9. [17]

    SWITCH- BOARD: Telephone speech corpus for research and develop- ment,

    J. J. Godfrey, E. C. Holliman, and J. McDaniel, “SWITCH- BOARD: Telephone speech corpus for research and develop- ment,” inAcoustics, speech, and signal processing, ieee interna- tional conference on, vol. 1. IEEE Computer Society, 1992, pp. 517–520

  10. [18]

    Identifying filled pauses in speech across south and west slavic languages,

    N. Ljube ˇsi´c, I. Porupski, P. Rupnik, and T. Kuzman, “Identifying filled pauses in speech across south and west slavic languages,” inProceedings of the 10th Workshop on Slavic Natural Language Processing (Slavic NLP 2025), 2025, pp. 1–8

  11. [19]

    Sep-28k: A dataset for stuttering event detection from podcasts with people who stutter,

    C. Lea, V . Mitra, A. Joshi, S. Kajarekar, and J. P. Bigham, “Sep-28k: A dataset for stuttering event detection from podcasts with people who stutter,” inICASSP 2021-2021 IEEE Interna- tional Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2021, pp. 6798–6802

  12. [20]

    Ksof: The kassel state of fluency dataset–a therapy centered dataset of stuttering,

    S. Bayerl, A. W. von Gudenberg, F. H ¨onig, E. N¨oth, and K. Ried- hammer, “Ksof: The kassel state of fluency dataset–a therapy centered dataset of stuttering,” inProceedings of the Thirteenth Language Resources and Evaluation Conference, 2022, pp. 1780– 1787

  13. [21]

    Fluencybank timestamped: An updated data set for disfluency detection and automatic intended speech recognition,

    A. Romana, M. Niu, M. Perez, and E. M. Provost, “Fluencybank timestamped: An updated data set for disfluency detection and automatic intended speech recognition,”Journal of Speech, Lan- guage, and Hearing Research, vol. 67, no. 11, pp. 4203–4215, 2024

  14. [22]

    Speech dis- fluency detection with contextual representation and data distil- lation,

    P. Mohapatra, A. Pandey, B. Islam, and Q. Zhu, “Speech dis- fluency detection with contextual representation and data distil- lation,” inProceedings of the 1st ACM international workshop on intelligent acoustic systems and applications, 2022, pp. 19–24

  15. [23]

    Detecting dysfluencies in stuttering therapy using wav2vec 2.0,

    S. P. Bayerl, D. Wagner, E. Noth, and K. Riedhammer, “Detecting dysfluencies in stuttering therapy using wav2vec 2.0,” inInter- speech, 2022

  16. [24]

    Automatic speech dis- fluency detection using wav2vec2. 0 for different languages with variable lengths,

    J. Liu, A. Wumaier, D. Wei, and S. Guo, “Automatic speech dis- fluency detection using wav2vec2. 0 for different languages with variable lengths,”Applied Sciences, vol. 13, no. 13, p. 7579, 2023

  17. [25]

    Disfluency rates in conversation: effects of age, rela- tionship, topic, role, and gender,

    H. Bortfeld, S. D. Leon, J. E. Bloom, M. F. Schober, and S. E. Brennan, “Disfluency rates in conversation: effects of age, rela- tionship, topic, role, and gender,”Language and Speech, vol. 44, no. 2, pp. 123–147, 2001

  18. [26]

    Uh and um as sociolinguistic markers in British En- glish,

    G. Tottie, “Uh and um as sociolinguistic markers in British En- glish,”International Journal of Corpus Linguistics, vol. 16, no. 2, pp. 173–197, 2011

  19. [27]

    On the use of uh and um in American English,

    ——, “On the use of uh and um in American English,”Functions of Language, vol. 21, no. 1, pp. 6–29, 2014

  20. [28]

    Gen- der in everyday speech and language: a corpus-based study,

    D. Binnenpoorte, C. Van Bael, E. den Os, and L. Boves, “Gen- der in everyday speech and language: a corpus-based study,” in Interspeech 2005, 2005, pp. 2213–2216

  21. [29]

    Variation and change in the use of hesitation mark- ers in germanic languages,

    M. Wieling, J. Grieve, G. Bouma, J. Fruehwald, J. Coleman, and M. Liberman, “Variation and change in the use of hesitation mark- ers in germanic languages,”Language Dynamics and Change, vol. 6, no. 2, pp. 199–234, 2016

  22. [30]

    Crible,Discourse Markers and (Dis)fluency

    L. Crible,Discourse Markers and (Dis)fluency. Amster- dam/Philadelphia: John Benjamins Publishing Company, 2018

  23. [31]

    On the role of pauses – a qualitative and quanti- tative analysis of selected political speeches in the European Par- liament,

    M. Mikl ´ossiov´a, “On the role of pauses – a qualitative and quanti- tative analysis of selected political speeches in the European Par- liament,”Studia Anglica Resoviensia, vol. 20, no. 1, pp. 101–117, 2023

  24. [32]

    A corpus analysis of patterns of age-related change in conversational speech,

    W. S. Horton, D. H. Spieler, and E. Shriberg, “A corpus analysis of patterns of age-related change in conversational speech,”Psy- chology and Aging, vol. 25, no. 3, pp. 708–713, 2010

  25. [33]

    Um . . . who like says you know: Filler word use as a function of age, gender, and personality,

    C. M. Laserna, Y . T. Seih, and J. W. Pennebaker, “Um . . . who like says you know: Filler word use as a function of age, gender, and personality,”Journal of Language and Social Psychology, vol. 33, no. 3, pp. 328–338, 2014

  26. [34]

    Speaking style effects in the production of disfluencies,

    H. Moniz, F. Batista, A. I. Mata, and I. Trancoso, “Speaking style effects in the production of disfluencies,”Speech communication, vol. 65, pp. 20–35, 2014

  27. [35]

    Disfluency across the lifespan: an individual differences investigation,

    P. E. Engelhardt and I. Markostamou, “Disfluency across the lifespan: an individual differences investigation,”Aging, Neuropsychology, and Cognition, vol. 32, no. 1, pp. 93–117, 2025, published online 20 May 2024. [Online]. Available: https://doi.org/10.1080/13825585.2024.2354958

  28. [36]

    Hesitation phenomena in sponta- neous english speech,

    H. Maclay and C. E. Osgood, “Hesitation phenomena in sponta- neous english speech,”Word, vol. 15, no. 1, pp. 19–44, 1959

  29. [37]

    To ‘errrr’ is human: ecology and acoustics of speech disfluencies,

    E. Shriberg, “To ‘errrr’ is human: ecology and acoustics of speech disfluencies,”Journal of the International Phonetic Association, vol. 31, no. 1, pp. 153–169, 2001

  30. [38]

    Recognizing emotions in dia- logue with disfluences and non-verbal vocalisations,

    L. Tian, C. Lai, and J. D. Moore, “Recognizing emotions in dia- logue with disfluences and non-verbal vocalisations,” inProceed- ings of the 4th Interdisciplinary Workshop on Laughter and other Non-Verbal Vocalisations in Speech, 2015

  31. [39]

    Recognizing emotions in spoken dialogue with acoustic and lexical cues,

    L. Tian, “Recognizing emotions in spoken dialogue with acoustic and lexical cues,” Ph.D. dissertation, University of Edinburgh, 2018. [Online]. Available: https://era.ed.ac.uk/handle/ 1842/31284

  32. [40]

    Disfluencies in public and private speech,

    D. Verdonik, P. Rupnik, and N. Ljube ˇsi´c, “Disfluencies in public and private speech,”Language and Speech, 2025

  33. [41]

    The voice of the people: Populism and donald trump’s use of informal voice,

    J. Kjeldgaard-Christiansen, “The voice of the people: Populism and donald trump’s use of informal voice,”Society, vol. 61, pp. 289–302, 2024

  34. [42]

    An analysis of speech fillers used by Biden,

    R. F. Hammad Al-Faragy, S. H. Suleiman Al Khalifawi, and H. H. Hasan Alqaisi, “An analysis of speech fillers used by Biden,” Journal of Language Studies, vol. 9, no. 1, pp. 346–360, 2025

  35. [43]

    The Power of Prosody and Prosody of Power: An Acoustic Analysis of Finnish Parliamentary Speech,

    M. Vainio, A. Suni, J. ˇSimko, and S. Kakouros, “The Power of Prosody and Prosody of Power: An Acoustic Analysis of Finnish Parliamentary Speech,” inSpeech Prosody 2024, 2024, pp. 662– 666

  36. [44]

    Analyzing german parliamentary speeches: A machine learning approach for topic and sentiment classification,

    L. P ¨atz, M. Beyer, J. Sp¨ath, L. Bohlen, P. Zschech, M. Kraus, and J. Rosenberger, “Analyzing german parliamentary speeches: A machine learning approach for topic and sentiment classification,”

  37. [45]

    Available: https://arxiv.org/abs/2508.03181

    [Online]. Available: https://arxiv.org/abs/2508.03181

  38. [46]

    Par- laspeech 3.0: Richly Annotated Spoken Parliamentary Cor- pora of Croatian, Czech, Polish, and Serbian,

    N. Ljube ˇsi´c, P. Rupnik, I. Porupski, and T. K. Punger ˇsek, “Par- laspeech 3.0: Richly Annotated Spoken Parliamentary Cor- pora of Croatian, Czech, Polish, and Serbian,”arXiv preprint arXiv:2511.01619, 2025

  39. [47]

    The ParlaSent multilin- gual training dataset for sentiment identification in parliamentary proceedings,

    M. Mochtak, P. Rupnik, and N. Ljubeˇsi´c, “The ParlaSent multilin- gual training dataset for sentiment identification in parliamentary proceedings,” inProceedings of the 2024 Joint International Con- ference on Computational Linguistics, Language Resources and Evaluation (LREC...

  40. [48]

    Longitudinal data analysis using generalized linear models,

    K.-Y . Liang and S. L. Zeger, “Longitudinal data analysis using generalized linear models,”Biometrika, pp. 13–22, 1986

  41. [49]

    On the pooling of time series and cross section data,

    Y . Mundlak, “On the pooling of time series and cross section data,”Econometrica, pp. 69–85, 1978

Pith tools

Reviewed July 8, 2026 · model on record in the stance chip above.