Pith. sign in

REVIEW 3 major objections 4 minor 85 references

"Dyadosyncrasy", Idiosyncrasy and Demographic Factors in Turn-Taking

T0 review · 3 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read Your conversation partner, not your demographics, sets the timing of turn-taking.

desk verdict A solid large-corpus variance decomposition whose headline 'dyadosyncrasy' is a conversation-level intercept, not yet evidence about relationships. read the letter →

arxiv 2505.24736 v1 pith:FDYK4U32 submitted 2025-05-30 eess.AS cs.CL

classification eess.AScs.CL
keywords turn-takingtransitionflooroffsetdyadosyncrasyidiosyncrasymixed-effectsmodelstelephoneconversationsdemographiceffects
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Turn-taking in dialogue is fast and near-universal, but its exact timing varies from person to person and conversation to conversation. This paper measures that variation through the Transition Floor Offset (TFO), the gap or overlap between the end of one speaker's turn and the start of the next, across nearly 11,000 recorded telephone conversations in US English. It finds that sex and age have small but statistically significant effects (women and older speakers switch slightly faster; education has no effect), and that lighter topics come with shorter offsets. The larger result is that a random effect for the conversation pair (Dyad ID) explains far more TFO variance than a random effect for the individual speaker, with conditional $R^2$ changes of 0.415 versus 0.128. The paper concludes that the dyadic relationship and joint activity, rather than a speaker's stable style or demographics, are the strongest determinants of turn-transition timing.

What carries the argument

The load-bearing machinery is a linear mixed-effects model with two random intercepts: Speaker ID and Dyad ID. Each speaker in the corpus appears in up to four different dyads, and each dyad has exactly one recorded conversation, so the model can ask whether timing variance travels with the person or with the pair. TFO values are extracted from each channel using a voice activity detector and a published turn-shift procedure, then averaged per speaker per conversation before modelling. The 'dyadosyncrasy' effect is operationally defined as the change in conditional $R^2$ when the Dyad random intercept is dropped from the full model, compared with the analogous change for Speaker.

What would settle it

Record the same pair of speakers in multiple separate conversations with different topics and tasks; if the dyad-level variance remains large when conversation instance is included as a separate factor, dyadosyncrasy is a stable property of the pair, whereas if it collapses to speaker-level variance, the large conditional $R^2$ change found here is mostly conversation-specific.

Watch

Extended reading notes

Core claim

On the paper's own terms, the central discovery is that turn-timing is primarily a property of the dyad, not of the individual. Using linear mixed-effects models with Speaker ID and Dyad ID as random intercepts and Sex, Age, and Topic as fixed effects, the authors estimate that removing the Dyad random intercept costs 0.415 in conditional $R^2$, while removing the Speaker intercept costs 0.128; the fixed effects are an order of magnitude smaller (Sex 0.0089, Age 0.0081, Topic 0.0094). In substantive terms, two speakers in the same conversation match each other's gap-and-overlap behavior more closely than either speaker matches their own behavior in a different conversation. The authors interpret this 'dyadosyncrasy' as evidence that dialogue is a joint activity in which mutual adaptation creates a pair-specific interactional signature.

Load-bearing premise

Because each pair of speakers has only one recorded conversation, the model treats all conversation-level variance as belonging to the pair; if topic details, task demands, or technical conditions vary from one call to the next, those factors are absorbed into the dyad term and the 0.415 estimate would overstate a stable relationship effect.

Editorial extensions

If this is right

  • Predictive turn-taking models for conversational systems should condition on long-term dyad-specific or speaker-specific patterns rather than only on the immediately preceding audio context.
  • Studies of spoken interaction should treat the dyad, not the individual speaker, as the primary unit of analysis for timing behavior.
  • Demographic factors alone explain little of the timing variance, so demographic conditioning cannot substitute for speaker- or pair-level adaptation.
  • Conversation topic is a measurable influence on turn timing, with serious or emotionally charged topics slowing exchanges, so benchmark data and models should control or report topic.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A confound the paper leaves unquantified: because each pair of speakers has exactly one recorded conversation, Dyad ID is statistically indistinguishable from Conversation ID, so the 0.415 estimate bundles task, topic, and recording-session effects together with any genuinely relational effect.
  • If the dyad effect is relational rather than situational, gap and overlap timing should converge over the course of a single conversation; that within-conversation trajectory is a testable consequence not examined here.
  • For a spoken dialogue system, the same human user paired with different synthetic partners should show measurably different TFO distributions; if it does not, the dyad-level effect may require richer interaction than current systems provide.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. This paper analyses turn-taking in 10,950 telephone conversations from the Fisher corpus, using an averaged Transition Floor Offset (TFO) per speaker and conversation. Through ANOVAs and linear mixed-effects models with Speaker and Dyad random effects, the authors report small but statistically significant effects of sex and age, negligible education effects, modest topic effects, and much larger random-effect contributions: Dyad ID accounts for a change in conditional R2 of 0.415 versus 0.128 for Speaker ID. The paper interprets this as evidence for 'dyadosyncrasy'—a pair-specific, jointly constructed turn-taking signature—and argues that dyads, not individual speakers, are the primary unit of analysis.

Significance. The question of how much demographic, individual, and dyadic/conversation factors contribute to turn-taking timing is relevant both for theories of joint action and for personalized turn-taking models. The paper's strengths are its use of a large, established corpus; a clear variance-decomposition protocol based on Nakagawa and Schielzeth's R2; and a transparent presentation of the small fixed-effect sizes. If the dyad-vs-speaker decomposition were robust, the result would be a useful addition to the literature and would motivate further work on conversation-level conditioning in spoken dialogue systems. The reported magnitudes, however, need to be re-examined because the 'Dyad' random effect is structurally identical to a conversation-level intercept, and the interpretation as relationship-specific 'dyadosyncrasy' goes beyond what the current model can identify.

major comments (3)
  1. [Section 3.4 and Discussion (§4)] The paper states that each conversation has a unique Dyad ID and that the same pair never has more than one conversation. This makes the Dyad random intercept exactly a conversation-level intercept. The reported change in conditional R2 of 0.415 therefore captures all conversation-level sources of variation not already explained by the fixed effects—topic nuances, task demands, recording conditions, VAD/channel artifacts, and chance—not only the pair's relationship or joint activity. The abstract and Discussion attribute this variance to 'the dyadic relationship and joint activity'; that attribution is not supported by the model alone. I recommend reframing the claim as a conversation-level effect, adding conversation-level covariates or a replication study with the same pairs in multiple conversations, and at minimum acknowledging this confound in the Limitations paragraph, which currently mentions stranger pairings but not this issue.
  2. [Section 3.2, Table 1] The ANOVA is conducted on 21,900 speaker-conversation averages, but the two averages from each conversation are not independent because Dyad ID is a shared factor. The p-values in Table 1 are therefore anti-conservative, and the claim that sex and age have 'significant' effects is not supported at the stated level. I would like to see the demographic effects tested in a model that accounts for the conversation-level clustering, for example the mixed model from Section 3.4, or with cluster-robust standard errors. The small effect sizes are not in question, but the significance statements need to be based on a valid test.
  3. [Table 2 and Section 3.4] The central comparison (Dyad 0.415 versus Speaker 0.128) is reported without any uncertainty measure, and the design makes the two variance components difficult to separate for speakers who appear in only one conversation. The mean number of conversations per speaker is about 1.87, so many speakers are singletons; for a singleton speaker, the speaker and conversation random effects are not individually identified, and the estimation must rely on the subset of repeated speakers. As a result, stable individual traits may be partially absorbed into the conversation/dyad component, potentially inflating the reported gap. I would like to see bootstrap confidence intervals for the R2 changes and a sensitivity analysis restricted to speakers with at least two conversations.
minor comments (4)
  1. [General] There are a few typos and formatting artifacts, including 'ANOV A' and 'V oice' in Section 3.2 and Section 2.2, respectively, and 'dyadsyncratic' in Section 4; these should be corrected.
  2. [Figure 3 and Section 2.1] The age bins in Figure 3 (16-23, 24-31, 32-39, 40-47, 48-55, 56+) do not match the age categories reported in Section 2.1 (16-24, 25-34, 35-44, 45-54, 55+); please harmonize or explain the discrepancy.
  3. [Section 3.1] The paper should state explicitly at the start of Section 3 that all analyses use the average TFO per speaker and conversation (21,900 units), not per-turn TFO, because this choice affects the interpretation of every variance component and R2 value.
  4. [Reproducibility] No data or code availability statement is included; given that the central claims rest on the exact random-effect specification and R2 computation, a reproducibility statement or link to analysis code would be valuable.

Circularity Check

0 steps flagged · score 1.0 of 10

No significant circularity: the variance decomposition is empirical, and the 'dyadosyncrasy' label is an interpretive step, not a derivation.

full rationale

The paper is an empirical variance decomposition, not a derivation from first principles, and no step reduces a prediction to its own inputs. The central result (Dyad ID explaining a 0.415 change in conditional R2 versus 0.128 for Speaker ID) is estimated from the data via mixed-effects models. The claim that 'speakers in a dyad resemble each other more than they resemble themselves in different dyads' follows mathematically from the fitted variance components: in the model, the correlation between two speakers in the same dyad equals the dyad variance, while the correlation across a speaker's different dyads equals the speaker variance. The paper explicitly notes that each conversation has a unique Dyad ID and that the same pair never converses twice, so Dyad ID is exactly a conversation-level random intercept. This transparency means the larger conversation-level variance is an empirical fact, not an imposed result. The interpretation of this variance as 'dyadic relationship and joint activity' is a construct-validity choice, not a circular derivation, even though other conversation-level factors (topic nuances, recording conditions) could contribute. No load-bearing self-citations are used: the cited prior work by the authors (e.g., VAD, turn-taking prediction) provides tools and background, not the theoretical conclusion. The limitations paragraph appropriately acknowledges the stranger-pairing caveat. Therefore, no circular step meets the evidential bar of quoting a specific reduction or a fitted parameter renamed as a prediction.

Assumptions & free parameters 3 free parameters · 3 assumptions · 0 invented entities

The analysis is a statistical decomposition; its main commitments are the validity of TFO extraction, the accuracy of Fisher metadata, and the adequacy of the mixed-effects structure. The variance components are fitted parameters and the dyad factor is confounded with conversation because each pair only appears once.

free parameters (3)
  • Variance component for Dyad ID = conditional R2 change = 0.415
    Fitted by REML; the paper's central claim that dyadic interaction dominates TFO relies on this number.
  • Variance component for Speaker ID = conditional R2 change = 0.128
    Fitted by REML; used as the idiosyncrasy benchmark against which the dyad effect is compared.
  • Fixed-effect slopes for Sex, Age, and Topic = marginal R2 changes 0.0089, 0.0081, 0.0094 respectively
    Fitted population effects that support the claim that demographic and topic effects are small.
assumptions (3)
  • domain assumption VAD-based extraction of TFO, following Heldner and Edlund (2010), gives a valid measure of conversational turn transitions.
    The entire analysis depends on the accuracy of the voice activity detector and turn-shift identification, which are referenced to prior work [25,28] and not evaluated here.
  • domain assumption Fisher demographic labels (sex, age, education) are accurate and the sample supports population inference.
    The demographic claims rely on the LDC-provided CSV labels; the paper does not independently validate them.
  • domain assumption Cross-classified random intercepts for Speaker and Dyad adequately model the dependencies in the data; no random slopes or alternative structures are tested.
    The variance decomposition assumes that speaker and dyad random intercepts, plus fixed effects, capture the structure; misspecification could change the variance attribution.

how reviews work

0 comments
Cite this review

Pith. "Pith review of "Dyadosyncrasy", Idiosyncrasy and Demographic Factors in Turn-Taking." pith.science (2026). https://pith.science/paper/FDYK4U32

@misc{pith2026250524736,
  author       = {Pith},
  title        = {Pith review of: "Dyadosyncrasy", Idiosyncrasy and Demographic Factors in Turn-Taking},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/FDYK4U32}},
  note         = {Machine review of arXiv:2505.24736}
}
read the original abstract

Turn-taking in dialogue follows universal constraints but also varies significantly. This study examines how demographic (sex, age, education) and individual factors shape turn-taking using a large dataset of US English conversations (Fisher). We analyze Transition Floor Offset (TFO) and find notable interspeaker variation. Sex and age have small but significant effects female speakers and older individuals exhibit slightly shorter offsets - while education shows no effect. Lighter topics correlate with shorter TFOs. However, individual differences have a greater impact, driven by a strong idiosyncratic and an even stronger "dyadosyncratic" component - speakers in a dyad resemble each other more than they resemble themselves in different dyads. This suggests that the dyadic relationship and joint activity are the strongest determinants of TFO, outweighing demographic influences.

Figures

Figures reproduced from arXiv: 2505.24736 by the authors.

Figure 1
Figure 1. [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Histogram of average TFO per speaker and dialog. 3.2. Demographics: Effects of Sex, Age, and Education To investigate the effects of demographic factors on TFO, an ANOVA analysis was performed. The results are shown in Ta￾ble 1. All demographic factors show significant effects, but with small effect sizes: sex only explains about 3.2% of the variation and age 1.3%. The effect of education is negligible [PITH_FULL_I… view at source ↗
Figure 3
Figure 3. Bars: average TFO per speaker and dialog, split on age and sex, with SE bars. Lines: Number of instances. 3.3. TFO and Topic Another factor that is likely to influence the turn-taking dynam￾ics to a certain extent is the topic of the conversation. Since the speakers in the Fisher corpus data collection were asked to discuss specific topics, we can analyze this effect. The average TFO per topic is shown in [PITH_FUL… view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: Average TFO per speaker and dialog, split on topic, with SE bars [PITH_FULL_IMAGE:figures/full_fig_p003_4.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

85 extracted references · 75 canonical work pages

  1. [1]

    "Dyadosyncrasy", Idiosyncrasy and Demographic Factors in Turn-Taking

    Introduction Humans spend a substantial portion of their lives engaged in spoken interactions, seamlessly coordinating conversations through finely tuned turn-taking mechanisms. These mecha- nisms, central to human communication, display strikingly uni- versal characteristics across languages – rooted in a shared cog- nitive and linguistic infrastructure ...

  2. [2]

    Corpus The analyses were based on the Fisher corpus of spoken tele- phone conversations in US English [27], collected by the Lin- guistic Data Consortium (LDC)

    Materials and Method 2.1. Corpus The analyses were based on the Fisher corpus of spoken tele- phone conversations in US English [27], collected by the Lin- guistic Data Consortium (LDC). We used both part 1 and 2 of the dataset, containing 10,950 English telephone conversations with 11,684 different speakers. Each individual speaker was in- volved in up t...

  3. [3]

    Overall distribution The average TFO for each speaker and dialogue was computed and the distribution is shown in Figure 2

    Results 3.1. Overall distribution The average TFO for each speaker and dialogue was computed and the distribution is shown in Figure 2. The average TFO is 0.300 s, which is similar that what has been reported in the literature [2]. However, there is also a considerable variation between speakers and dialogs (SD= 0.228). 0.50 0.25 0.00 0.25 0.50 0.75 1.00 ...

  4. [4]

    If the Increase is Sufficient

    Minimum Wage. If the Increase is Sufficient

  5. [5]

    Acceptable Humor and Bad T aste

    Comedy. Acceptable Humor and Bad T aste

  6. [6]

    Hyp. Sit. Committing Perjury for a Loved One

  7. [7]

    One Million Dollars to leave the US

  8. [8]

    Hyp. Sit. Opening your own Business

Show all 85 references
  1. [9]

    Hyp. Sit. Time Travel. Changing the Past

  2. [10]

    One Million Dollars not to speak with Their Best friend

  3. [11]

    How to correct it

    US Public Schools. How to correct it

  4. [12]

    If it is a good policy

    Affirmative Action. If it is a good policy

  5. [13]

    Preferences and Last Movie Watched

    Movies. Preferences and Last Movie Watched

  6. [14]

    Favorite Game

    Computer games. Favorite Game

  7. [15]

    How they keep up with it

    Current Events. How they keep up with it

  8. [16]

    Favorite Hobbies and Time Invested

    Hobbies. Favorite Hobbies and Time Invested

  9. [17]

    Most Important Thing to Look For

    Life Partners. Most Important Thing to Look For

  10. [18]

    How people react to it

    T errorism. How people react to it

  11. [19]

    If They Should Exist

    T elevised Criminal Trials. If They Should Exist

  12. [20]

    Companies T esting Employees

    Drug testing. Companies T esting Employees

  13. [21]

    Thoughts on Divorce and Marriage

    Family Values. Thoughts on Divorce and Marriage

  14. [22]

    Right to Forbid Certain Books

    Censorship. Right to Forbid Certain Books

  15. [23]

    Starting and Maintaining Them

    Health and Fitness. Starting and Maintaining Them

  16. [24]

    If Their Lives Changed After

    September 11. If Their Lives Changed After

  17. [25]

    Strikes by Professional Athletes and Their High Salaries

  18. [26]

    Heightened Security and T errorism

    Airport Security. Heightened Security and T errorism

  19. [27]

    The US Involvement

    Issues in the Middle East. The US Involvement

  20. [28]

    Threats to safety

    Foreign Relations. Threats to safety

  21. [29]

    using thelme4package [30]. Model parameters were es- Figure 1:Illustration of how turn shifts were identified based on VAD data.S pos exemplifies a shift with positive TFO (a gap) attributed to speaker B.S neg exemplifies a shift with negative TFO (a between-speaker overlap) a...

  22. [30]

    What the Word Means to Them

    Family. What the Word Means to Them

  23. [31]

    Illegal business and Scandals

    Corporate Conduct. Illegal business and Scandals

  24. [32]

    If Computers are helpful or harmful

    Education. If Computers are helpful or harmful

  25. [33]

    Best Friends and Acquaintances

    Friends. Best Friends and Acquaintances

  26. [34]

    Outdoor Activities in Cold or Warm Weather

  27. [35]

    T aking care of one's Health

    Illness. T aking care of one's Health

  28. [36]

    Choosing the worst one from a list

    Personal Habits. Choosing the worst one from a list

  29. [37]

    Ban, Prevention, Ads, and Ideas

    Smoking. Ban, Prevention, Ads, and Ideas

  30. [38]

    What the US Should do

    Arms Inspections in Iraq. What the US Should do

  31. [39]

    Create a Holiday and Describe it

    Holidays. Create a Holiday and Describe it

  32. [40]

    What the US can do to Prevent it

    Bioterrorism. What the US can do to Prevent it

  33. [41]

    Owning and Caring for them

    Pets. Owning and Caring for them

  34. [42]

    Preferences and Perfect meal

    Food. Preferences and Perfect meal

  35. [43]

    Favorite Professional TV Sport and Time Invested

  36. [44]

    dyadosyncrasy

    Reality TV. Thoughts on Them and on Their Popularity. Figure 4:Average TFO per speaker and dialog, split on topic, with SE bars. Table 2:Quantifying each factor’s contribution to TFO vari- ance based on changes inR 2. Fixed effects Change inR 2 m Sex 0.0089 Age 0.0081 Topic 0....

  37. [45]

    dyadosyncratic

    Discussion In this study, we assessed the extent to which dyadic, individ- ual and demographic factors can influence turn-taking behav- ior. The findings suggest that while the cognitive mechanisms supporting turn-taking may be universal, there is considerable variation in tur...

  38. [46]

    Conclusion To our knowledge, this study is the first to systematically quan- tify how individual, demographic, and dyadic factors influ- ence turn-taking behavior using a large-scale spoken dialogue dataset. Our findings reveal that while sex and age modestly shape the transit...

  39. [47]

    Multi- dimensional analysis of language, discourse, and society,

    Acknowledgments This study was financed, in part, by the S ˜ao Paulo Research Foundation (FAPESP), Brazil. Grant #2024/06797-2. The research is connected to the thematic project titled “Multi- dimensional analysis of language, discourse, and society,” which is funded by FAPESP...

  40. [48]

    Turn-taking in human communication–origins and implications for language processing,

    S. C. Levinson, “Turn-taking in human communication–origins and implications for language processing,”Trends in cognitive sci- ences, vol. 20, no. 1, pp. 6–14, 2016

  41. [49]

    Universals and cultural variation in turn-taking in conversation,

    T. Stivers, N. J. Enfield, P. Brown, C. Englert, M. Hayashi, T. Heinemann, G. Hoymann, F. Rossano, J. P. De Ruiter, K.-E. Yoonet al., “Universals and cultural variation in turn-taking in conversation,”Proceedings of the National Academy of Sciences, vol. 106, no. 26, pp. 10 58...

  42. [50]

    Timing in turn-taking and its im- plications for processing models of language,

    S. C. Levinson and F. Torreira, “Timing in turn-taking and its im- plications for processing models of language,”Frontiers in psy- chology, vol. 6, p. 731, 2015

  43. [51]

    Children’s verbal turn-taking,

    S. Ervin-Tripp, “Children’s verbal turn-taking,”Developmental pragmatics, pp. 391–414, 1979

  44. [52]

    Turn-taking, timing, and planning in early language acquisition,

    M. Casillas, S. C. Bobb, and E. V . Clark, “Turn-taking, timing, and planning in early language acquisition,”Journal of child lan- guage, vol. 43, no. 6, pp. 1310–1337, 2016

  45. [53]

    A simplest systemat- ics for the organization of turn-taking for conversation,

    H. Sacks, E. A. Schegloff, and G. Jefferson, “A simplest systemat- ics for the organization of turn-taking for conversation,”language, vol. 50, no. 4, pp. 696–735, 1974

  46. [54]

    Predicting and regulating participation equality in human-robot conversations: Effects of age and gender,

    G. Skantze, “Predicting and regulating participation equality in human-robot conversations: Effects of age and gender,” in Proceedings of the 2017 acm/ieee international conference on human-robot interaction, 2017, pp. 196–204

  47. [55]

    Social correlates of turn-taking style,

    J. Grothendieck, A. L. Gorin, and N. Borges, “Social correlates of turn-taking style,”Computer Speech & Language, vol. 25, no. 4, pp. 789–801, 2011

  48. [56]

    A dimensional model of interaction style variation in spoken dialog,

    N. G. Ward and J. E. Avila, “A dimensional model of interaction style variation in spoken dialog,”Speech Communication, vol. 149, pp. 47–62, 2023

  49. [57]

    Switching pauses in cooperative and competitive conversations,

    C. Trimboli and M. B. Walker, “Switching pauses in cooperative and competitive conversations,”Journal of Experimental Social Psychology, vol. 20, no. 4, pp. 297–311, 1984

  50. [58]

    An analysis of the timing of turn- taking in a corpus of goal-oriented dialogue,

    M. Bull and M. P. Aylett, “An analysis of the timing of turn- taking in a corpus of goal-oriented dialogue,” in5th International Conference of Spoken Language Processing (ICSLP’98). ISCA, 1998, pp. 1175–1178

  51. [59]

    Very short utterances and timing in turn-taking

    M. Heldner, J. Edlund, A. Hjalmarsson, and K. Laskowski, “Very short utterances and timing in turn-taking.” inInterspeech, 2011, pp. 2837–2840

  52. [60]

    Durational as- pects of turn-taking in spontaneous face-to-face and telephone di- alogues,

    L. Ten Bosch, N. Oostdijk, and J. P. De Ruiter, “Durational as- pects of turn-taking in spontaneous face-to-face and telephone di- alogues,” inText, Speech and Dialogue: 7th International Con- ference, TSD 2004, Brno, Czech Republic, September 8-11, 2004. Proceedings 7. Spring...

  53. [61]

    Long gaps between turns are awkward for strangers but not for friends,

    E. M. Templeton, L. J. Chang, E. A. Reynolds, M. D. Cone LeBeaumont, and T. Wheatley, “Long gaps between turns are awkward for strangers but not for friends,”Philosophical Transactions of the Royal Society B, vol. 378, no. 1875, p. 20210471, 2023

  54. [62]

    Big five predictors of behavior and perceptions in initial dyadic interactions: Personality similarity helps extraverts and introverts, but hurts “disagreeables

    R. Cuperman and W. Ickes, “Big five predictors of behavior and perceptions in initial dyadic interactions: Personality similarity helps extraverts and introverts, but hurts “disagreeables”.”Journal of personality and social psychology, vol. 97, no. 4, p. 667, 2009

  55. [63]

    Identifying personality traits using overlap dynamics in multiparty dialogue,

    M. Yu, E. Gilmartin, and D. Litman, “Identifying personality traits using overlap dynamics in multiparty dialogue,”Proceedings of the Annual Conference of the International Speech Communica- tion Association, INTERSPEECH, 2019

  56. [64]

    Pragmatic aspects of temporal accommodation in turn-taking,

    ˇS. Be ˇnuˇs, A. Gravano, and J. Hirschberg, “Pragmatic aspects of temporal accommodation in turn-taking,”Journal of Pragmatics, vol. 43, no. 12, pp. 3001–3027, 2011

  57. [65]

    Measuring acoustic-prosodic entrainment with respect to multiple levels and dimensions,

    R. Levitan and J. B. Hirschberg, “Measuring acoustic-prosodic entrainment with respect to multiple levels and dimensions,” in Interspeech, 2011, pp. 3081–3084

  58. [66]

    Pause and gap length in face-to-face interaction,

    J. Edlund, J. B. Hirschberg, and M. Heldner, “Pause and gap length in face-to-face interaction,” inInterspeech, 2009, pp. 2779– 2782

  59. [67]

    Monitoring convergence of temporal features in spontaneous dialogue speech,

    S. Kousidis and D. Dorran, “Monitoring convergence of temporal features in spontaneous dialogue speech,” inDigital Media Centre Conference papers. Technological University Dublin, 2009

  60. [68]

    Synchronization among speakers reduces macro- scopic temporal variability,

    F. Cummins, “Synchronization among speakers reduces macro- scopic temporal variability,” inProceedings of the Annual Meet- ing of the Cognitive Science Society, vol. 26, no. 26, 2004

  61. [69]

    Measuring synchronization among speakers reading to- gether,

    ——, “Measuring synchronization among speakers reading to- gether,” inITRW on Experimental Linguistics, 2006

  62. [70]

    Turn-taking in conversational systems and human- robot interaction: a review,

    G. Skantze, “Turn-taking in conversational systems and human- robot interaction: a review,”Computer Speech & Language, vol. 67, p. 101178, 2021

  63. [71]

    Towards a general, continuous model of turn-taking in spo- ken dialogue using lstm recurrent neural networks,

    ——, “Towards a general, continuous model of turn-taking in spo- ken dialogue using lstm recurrent neural networks,” inProceed- ings of the 18th Annual SIGdial Meeting on Discourse and Dia- logue, 2017, pp. 220–230

  64. [72]

    V oice activity projection: Self- supervised learning of turn-taking events,

    E. Ekstedt and G. Skantze, “V oice activity projection: Self- supervised learning of turn-taking events,” inInterspeech, 2022, pp. 5190–5194

  65. [73]

    Multilingual turn-taking prediction using voice activity projec- tion,

    K. Inoue, B. Jiang, E. Ekstedt, T. Kawahara, and G. Skantze, “Multilingual turn-taking prediction using voice activity projec- tion,”arXiv preprint arXiv:2403.06487, 2024

  66. [74]

    The fisher corpus: A resource for the next generations of speech-to-text

    C. Cieri, D. Miller, and K. Walker, “The fisher corpus: A resource for the next generations of speech-to-text.” inLREC, vol. 4, 2004, pp. 69–71

  67. [75]

    Pauses, gaps and overlaps in conver- sations,

    M. Heldner and J. Edlund, “Pauses, gaps and overlaps in conver- sations,”Journal of Phonetics, vol. 38, no. 4, pp. 555–568, 2010

  68. [76]

    [Online]

    R Core Team,R: A Language and Environment for Statistical Computing, R Foundation for Statistical Computing, Vienna, Austria, 2021. [Online]. Available: https://www.R-project.org/

  69. [77]

    Fitting linear mixed-effects models using lme4,

    D. Bates, M. M ¨achler, B. Bolker, and S. Walker, “Fitting linear mixed-effects models using lme4,”Journal of Statistical Software, vol. 67, no. 1, pp. 1–48, 2015

  70. [78]

    A general and simple method for obtaining r2 from generalized linear mixed-effects models,

    S. Nakagawa and H. Schielzeth, “A general and simple method for obtaining r2 from generalized linear mixed-effects models,”Meth- ods in ecology and evolution, vol. 4, no. 2, pp. 133–142, 2013

  71. [79]

    performance: An R package for assessment, com- parison and testing of statistical models,

    D. L ¨udecke, M. S. Ben-Shachar, I. Patil, P. Waggoner, and D. Makowski, “performance: An R package for assessment, com- parison and testing of statistical models,”Journal of Open Source Software, vol. 6, no. 60, p. 3139, 2021

  72. [80]

    H. H. Clark,Using language. Cambridge university press, 1996

  73. [81]

    No gap, lots of overlap: Turn-taking patterns in the talk of women friends,

    J. Coates, “No gap, lots of overlap: Turn-taking patterns in the talk of women friends,”Researching language and literacy in social context, pp. 177–192, 1994

  74. [82]

    Backchan- nel behavior is idiosyncratic,

    P. Blomsma, J. Vaitonyt´e, G. Skantze, and M. Swerts, “Backchan- nel behavior is idiosyncratic,”Language and Cognition, pp. 1–24, 2024

  75. [83]

    Backchannels revisited from a multimodal perspective,

    R. Bertrand, G. Ferr ´e, P. Blache, R. Espesser, and S. Rauzy, “Backchannels revisited from a multimodal perspective,” in Auditory-visual Speech Processing, 2007, pp. 1–5

  76. [84]

    The can- dor corpus: Insights from a large multimodal dataset of naturalis- tic conversation,

    A. Reece, G. Cooney, P. Bull, C. Chung, B. Dawson, C. Fitz- patrick, T. Glazer, D. Knox, A. Liebscher, and S. Marin, “The can- dor corpus: Insights from a large multimodal dataset of naturalis- tic conversation,”Science Advances, vol. 9, no. 13, p. eadf3197, 2023

  77. [85]

    Multi- parametric analysis of speech timing in inter-talker identical twin pairs and cross-pair comparisons: Some forensic implications,

    J. C. Cavalcanti, A. Eriksson, and P. A. Barbosa, “Multi- parametric analysis of speech timing in inter-talker identical twin pairs and cross-pair comparisons: Some forensic implications,” Plos one, vol. 17, no. 1, p. e0262800, 2022

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.