Pith. sign in

REVIEW 5 major objections 5 minor 42 references

Articulatory modeling of the S-shaped F2 trajectories observed in \"Ohman's spectrographic analysis of VCV syllables

T0 review · 5 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read S-shaped formant transitions emerge from the interaction of separate vowel and consonant articulatory plans.

desk verdict A transparent, reproducible simulation that credibly produces S-shaped F2 transitions from two-tier articulatory planning, but the load-bearing geometry is hand-set and the comparison with Öhman is qualitative. read the letter →

arxiv 2505.22455 v1 pith:U3GIWILF submitted 2025-05-28 eess.AS cs.SD

classification eess.AScs.SD
keywords formanttransitionsS-shapedF2trajectoriescoarticulationlocusequationsMaedaarticulatorymodelsyllablesynthesisTauVCVsyllables
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tries to show that the S-shaped second-formant transitions seen in classic vowel-consonant-vowel measurements can be reproduced with a physiologically structured articulatory model, not just by tube acoustics. It plans vowel-to-vowel movement and consonant constriction separately, then combines them through a geometric graph in a polar plane and tau-type kinematics. The resulting synthetic data match the characteristic S-curves and locus equations of the original 75 VCV sequences. If the paper is right, the S-shape is not a direct sigmoidal gesture but an emergent product of the coordinated interaction of separate articulator groups.

What carries the argument

The machinery is a syllable graph in a polar plane: each vowel and consonant is a node with abstract polar coordinates standing for place of articulation and degree of opening, and each transition is an arc, with the vowel-to-vowel arc twice as slow (2T) as the vowel-to-consonant arcs (T). Movement along each arc is driven by a tau-type kinematic law, a cosine decay over the arc's duration, and the seven articulatory parameters of the Maeda model are split by selection vectors into a vocalic group and a consonantal group whose outputs are superimposed. The load-bearing detail is that the vocalic parameters themselves do not follow a sigmoid, so any S-shape in the resulting F2 track has to come from the interaction of the two control tiers.

What would settle it

Measure or simulate real front-to-back and back-to-front VCV syllables and compare the timing of the F2 inflection relative to consonant closure with the polar-distance rule: the paper predicts the transition falls after the constriction when V1 is closer to the consonant and before it when the consonant is closer to V2. If the inflection timing instead follows measured articulator speed or vocal-tract geometry, the proposed mechanism would be falsified; so would it be if a single sigmoidal vocalic gesture with no consonant tier already produced the same S-shapes.

Watch

Extended reading notes

Core claim

The central claim is that the S-shaped F2 trajectory in V1CV2 syllables is a composite phenomenon: it appears only when the slow vowel-to-vowel arc (duration 2T) and the faster vowel-to-consonant and consonant-to-vowel arcs (duration T) are planned separately and then superimposed through the articulatory model's parameters. Synthetic spectrograms produced this way reproduce the classic 75-utterance measurements, including the characteristic S-curves, and the locus equations derived from CV versions of the same planning graph are consistent with formant-to-cavity affiliations rather than with any single invariant consonant locus. The paper concludes that where the S-transition occurs in time, before or after the consonantal constriction, is governed by the converging directions of the vocalic and consonantal pathways in the polar plane, so the S-shape emerges from the coordinated synergy of all articulators rather than from a sigmoidal vocalic parameter change.

Load-bearing premise

The load-bearing premise is that the hand-assigned polar-plane coordinates for each vowel and consonant, together with the hand-chosen split of articulatory parameters into vocalic and consonantal groups, faithfully represent how the vocal tract is actually controlled.

Editorial extensions

If this is right

  • The S-shaped F2 transition in the classic VCV data can be reproduced with articulatory constraints, so the phenomenon does not require a sigmoidal gesture built into the vowel parameters.
  • The timing of the S-transition relative to the consonant closure is determined by which target the converging pathways reach first, giving a concrete, checkable prediction of the polar-plane graph.
  • Locus equations for F2 and F3 can be steep or flat depending on vowel context and lip parameters, undercutting the hypothesis of a single invariant consonant locus.
  • Because the same planning graph handles VCV and CV structures (the latter with an initial reduced vowel), the framework gives one unified account of coarticulation in both.
  • The complete 75-utterance set is reproduced with open-source software, so the claimed mechanism can be checked directly.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the polar-plane coordinates were learned from articulatory data rather than hand-set, the same graph formalism could turn measured tongue movements into inferred syllable plans, an extension the paper itself lists as open.
  • The distance rule invites a perceptual test: systematically vary the polar distance between vowels and consonant in synthetic syllables and ask whether listeners hear the F2 inflection move before or after the closure as the model predicts.
  • The locus-equation reinterpretation implies that acoustic features used for stop-consonant classification may be capturing articulator synergy rather than a stable place cue, with possible consequences for automatic speech recognition.
  • Because the S-shape is claimed to be insensitive to formant affiliation and driven by relative positions, the mechanism could generalize to other constriction types, although the paper only tests plosives.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 5 minor

Summary. This paper proposes an articulatory synthesis model for Öhman's VCV syllables using the Maeda articulatory model. Vowel-to-vowel and consonant-related Maeda parameters are planned separately as arcs in a polar 'syllable graph,' with Tau-model kinematics, and combined through mutually exclusive selection vectors. The authors simulate 75 VCVs, compare them visually with Öhman's spectrograms, and analyze F2/F3 locus equations. The central claim is that S-shaped F2 trajectories are not a direct product of a sigmoidal vocalic parameter change but emerge from the coordinated interaction of all articulators, and that this mechanism also explains and reinterprets locus equations. The paper provides open-source software and supplementary materials for reproducibility.

Significance. If the central claim holds, the paper offers a genuinely articulatory account of a classic coarticulatory phenomenon, extending earlier tube-model and area-function simulations to a parameterized articulatory model, and providing a possible mechanistic basis for locus equations. The use of the Maeda model with separate vowel and consonant control tiers is well motivated by prior work, and the provision of reproducible code and supplementary figures is a strength. However, the significance is currently tempered by the qualitative nature of the validation and the large number of hand-set parameters, so the result is better read as a plausible proof-of-concept than as a demonstrated explanation.

major comments (5)
  1. [Section 3, Figure 2] The central claim that synthetic data 'exhibit similar characteristics' to Öhman's spectrograms rests entirely on a visual, manually aligned superimposition. No quantitative comparison is reported: there are no error bars, no RMS differences between measured and synthetic F2 tracks, and no statistical test of the S-shape classification. Because the letter's main contribution is reproducing Öhman's observations, the manuscript should provide a quantitative validation, for example by digitizing the original F2 tracks and computing a fit metric, or by reporting blind inter-rater agreement on the presence of S-shapes.
  2. [Sections 2.1 and 3, Section 4] The polar coordinates of vowels and consonants, the articulator selection vectors, and several kinematic constants are free parameters chosen by hand, and Section 2.1 explicitly says the consonantal parameter choice 'was chosen experimentally, which may contradict Öhman's simulation results.' The case A/B classification and the claimed 'composite synergy' mechanism in Section 4 depend directly on whether the V1C and V1V2 arcs converge or diverge, which is determined by those hand-set positions. Without a sensitivity analysis (e.g., perturbing the polar coordinates and selection vectors within a plausible range) or an independent grounding of the coordinates in articulatory data, the emergent S-shape may be an artifact of the coordinate choice rather than a consequence of Maeda's articulatory structure.
  3. [Section 4, Figure 1d] The paper states that the S-shape 'is not strictly equivalent to the sigmoid transition of the vocalic parameters' (red dots in Figure 1b), but this is only a qualitative observation. A stronger test would compare the full model's F2 trajectory against a control simulation in which only the vocalic arc is active (or in which the consonantal parameters follow a single sigmoid), and show that the S-shape persists or appears only when the composite mechanism is present. As written, the distinction between 'S-shaped' and 'sigmoid' is asserted rather than demonstrated.
  4. [Section 3, Figure 3] The 3D comparison of F2v-F2c-F3c clusters with Lindblom et al.'s data is described as 'overall structure remains consistent,' but no quantitative measure of consistency is provided, such as cluster overlap metrics or per-consonant RMS errors. The noted larger F3 variation in the synthetic data could itself affect the cluster geometry, so the claim of consistency is not yet supported.
  5. [Conclusion] The concluding sentence acknowledges that the model 'is not data-driven and lacks observable control mechanisms.' This is a direct admission that the load-bearing parameter choices (coordinates, selection vectors) are not constrained by data. While the paper is honest about this limitation, the central explanation of S-shape emergence currently rests on these unconstrained choices, and the discussion should either strengthen the justification or explicitly reframe the contribution as a proof-of-concept subject to future data-driven inference.
minor comments (5)
  1. [Throughout] The author name and title use '¨Ohman' with a misplaced diaeresis; every occurrence should be 'Öhman'.
  2. [Section 3] Typo: 'the indentification of [g]' should be 'the identification of [g]'.
  3. [Section 4] Typo: 'Symetrically' should be 'Symmetrically'.
  4. [Figure 4 caption] Typo: 'Comparision' should be 'Comparison'.
  5. [Section 4] The sudden change to K=1000 for the A/B analysis is not justified; if K affects the geometry of the arcs, the effect of this change on the S-shape should be reported or at least discussed.

Circularity Check

1 steps flagged · score 6.0 of 10

S-shaped-F2 'emergence' is largely installed by construction: the hand-set polar coordinates and Maeda selection vectors, admitted to be 'chosen experimentally' (Sec. 2.1), fix the two-tier detour-plus-interpolation that deterministically produces the S-shape through Eqs.

  1. fitted input called prediction [Sec. 2.1 (planning, Eqs. 2 and 5), Sec. 3 (polar coordinate table), Abstract and Sec. 4 (emergence claim)]
    "All tongue parameters are involved in the production of [d] and [g], and this was chosen experimentally, which may contradict Öhman's simulation results [2]."

    Eqs. 2-5 map the hand-assigned polar coordinates and hand-set selection vectors Sv, Sc onto Maeda trajectories P(t); the consonantal assignment is admitted to have been 'chosen experimentally' (Sec. 2.1). Planning superimposes a slow V1-to-V2 interpolation (duration 2T) on an out-and-back consonant detour V1-to-C-to-V2 (durations T, T); the F2 S-shape, defined in Sec. 4 by five points around the consonant ('F2c1 at Δt=−40 ms', 'F2 at the onset center', 'F2c2 at Δt=40 ms'), is the deterministic image of that detour-plus-interpolation under the Maeda acoustic mapping.

full rationale

The derivation chain runs: hand-assigned polar coordinates (Sec. 3) and hand-set Maeda selection vectors Sv, Sc (Sec. 2.1) to Eq. 2 arc plans, then to Eq. 5 parameter trajectories P(t), then to formants via a transmission-line solver. The central claim, that S-shaped F2 trajectories emerge from a composite mechanism governed by the coordinated synergy of all articulators, is a statement about this pipeline: the S-shape is produced by superimposing a slow vowel-to-vowel interpolation on a faster out-and-back consonant detour, exactly the design chosen. Because the consonantal parameter choice is admitted to have been made experimentally and the polar positions are set by hand with no sensitivity analysis, the S-shape is in large part installed by construction rather than derived from Maeda's articulatory constraints; Section 4's case A/B rule, said to be driven by relative positions within the polar plane, likewise reads the pattern off the same hand-set geometry that generated the data. This matches the fitted-input-called-prediction pattern in its informal, non-statistical guise: inputs were tuned with the target phenomenon in view, and the resulting behavior is presented as an emergent finding. Mitigating factors that keep this at partial, not total, circularity: the Maeda model [1] and the Tau-model median kappa=0.4 [14] are external anchors, the code and supplementary materials are fully released [22, 31] and reproducible, and the paper explicitly discloses its status as 'not data-driven' with 'no observable control mechanisms', so nothing is concealed. The locus-equation analysis is in-sample (computed from synthetic data of the same hand-tuned model), which limits its evidentiary weight but is not itself circular. No load-bearing self-citation chain or imported uniqueness theorem is present, so the self-citation patterns contribute nothing to the score. The remaining independent content is the specific Maeda articulatory-to-acoustic mapping and the observable, reproducible synthetic statistics; the claimed 'emergence' itself is heavily constructed.

Assumptions & free parameters 8 free parameters · 5 assumptions · 0 invented entities

The model's output depends on a large number of hand-set coordinates, selection vectors, and kinematic constants, which are not independently measured. The paper does not introduce new physical entities; the polar plane is an abstract representational device. The main 'free' content is the polar geometry and the articulator assignment rules.

free parameters (8)
  • Polar coordinates of vowels = y:(0.7,11π/6), ø:(0.3,11π/6), a:(0.8,π), o:(0.8,π/2), u:(1,π/3)
    Set by hand to represent abstract place of articulation and degree of opening; not derived from data.
  • Polar coordinates of consonants = b:(1.2,π/3), d:(1.2,23π/16), gp:(1.1,23π/12), gv:(1.2,π/3)
    Chosen to place constrictions at appropriate places for each consonant; [gv] shares a position with [b].
  • Arc curvature parameter K = 30 for vowel arcs, 10 for consonant arcs, 1000 for straight-case analysis
    Large K values make arcs quasi-straight; the specific values are chosen to 'generate useful variability' (Section 2.1).
  • Tau exponent kappa = 0.5
    Chosen to approximate kappa=0.4, the median value from articulographic data in Elie et al. [14].
  • Consonantal articulator selection vectors = b:{1,2,6}, d:{1,2,3,4}, g:{1,2,3,4}
    Hand-assigned which Maeda parameters participate in each consonant; described as 'chosen experimentally' (Section 2.1).
  • Transition duration T = 16 steps of 10 ms (V1V2 arc duration 2T)
    Sets the time scale of the simulation.
  • Reduced vowel position = (0.5ρV, θV)
    Used for CV syllables as an anticipatory reduced vowel; position is half the radius of the target vowel.
  • Formant measurement offsets = Delta t = 0 ms and 40 ms from consonant onset center
    These time points are selected for locus-equation analysis (Section 4).
assumptions (5)
  • domain assumption The Maeda articulatory model provides a valid articulatory-to-acoustic mapping.
    The paper uses the Maeda model [1] as ground truth for formant generation without validating it on the specific VCV set beyond this work.
  • domain assumption The Tau model describes articulatory kinematics with the specified exponent.
    Equation 4 is adopted with kappa=0.5 as an approximation of the empirically identified kappa=0.4; the equivalence of Eq. 3 and Eq. 4 is asserted in Section 2.2.
  • domain assumption Vowel and consonant gestures can be planned independently and superimposed.
    This is the core Öhman/Perkell hypothesis, adopted at the outset; the paper does not test it, it uses it as the design principle.
  • domain assumption Maeda parameters are independently controllable via PCA and can be assigned to vocalic or consonantal arcs.
    Relies on the model properties described in [1], used to justify dropping the coarticulation function wc(x).
  • domain assumption Öhman's tabulated data (Tables II and IV) are accurate and representative.
    Used as the comparison target for the 3D locus-equation clusters; the comparison is qualitative.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Articulatory modeling of the S-shaped F2 trajectories observed in \"Ohman's spectrographic analysis of VCV syllables." pith.science (2026). https://pith.science/paper/U3GIWILF

@misc{pith2026250522455,
  author       = {Pith},
  title        = {Pith review of: Articulatory modeling of the S-shaped F2 trajectories observed in \"Ohman's spectrographic analysis of VCV syllables},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/U3GIWILF}},
  note         = {Machine review of arXiv:2505.22455}
}
read the original abstract

The synthesis of Ohman's VCV sequences with intervocalic plosive consonants was first achieved 30 years ago using the DRM model. However, this approach remains primarily acoustic and lacks articulatory constraints. In this study, the same 75 VCVs are analyzed, but generated with the Maeda model, using trajectory planning that differentiates vowel-to-vowel transitions from consonantal influences. Synthetic data exhibit similar characteristics to Ohman's sequences, including the presence of S-shaped F2 trajectories. Furthermore, locus equations (LEs) for F2 and F3 are computed from synthetic CV data to investigate their underlying determinism, leading to a reassessment of conventional interpretations. The findings indicate that, although articulatory planning is structured separately for vowel and consonant groups, S-shaped F2 trajectories emerge from a composite mechanism governed by the coordinated synergy of all articulators.

Figures

Figures reproduced from arXiv: 2505.22455 by the authors.

Figure 1
Figure 1. Syllable planning and synthesis of the utterance /ydu/. x-axis:    s(x;t) = v(x;t) + k(t) (c(x) − v(x;t)) wc(x) = (1 − k(t) wc(x)) v(x;t) + k(t) wc(x) c(x) v(x;t) = (1 − k ′ (t)) v1(x) + k ′ (t) v2(x) (1) The slow variation process is governed by its own monotonic kinematic term, k ′ (t), which transitions smoothly from 0 to 1, defining a trajectory between the two static vowels, v1(x) and v2(x). Note that the … view at source ↗
Figure 2
Figure 2. Successful superimposition of Ohman 1966 [15] and spectrograms of synthetic VCVs. ¨ A closely related harmonic form (additional details are provided in the supplementary materials [22] for verification of the fol￾lowing claims) describes the harmonic decay over a quarter pe￾riod, where t ∈ [1, D]: X(t) = X0 cos  πt 2D  1 κ (4) In the trajectory equations, this latter form is applied to ρ(t) with X0 = 1 in Eq. 2 bu… view at source ↗
Figure 3
Figure 3. 3D representation as in Lindblom et al. [24] [PITH_FULL_IMAGE:figures/full_fig_p003_3.png] view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: a-b, highlighting that LEs are also governed by vowel parameters: LipP and LipH for the front cavity, and Hy for the back cavity. Globally, at the consonant onset center ∆t = 0 ms in Figure 4a, the slopes of the F2v-F2c LEs are notably steep for [b] and [gv], challengi…

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

42 extracted references · 41 canonical work pages

  1. [1]

    To date, this transition has not been systematically modeled using an articulatory framework

    Introduction The S-shaped transition of the second formant (F2) in a V1CV2 sequence, where V 1 and V2 have opposite characteris- tics (front/back), is a key element for understanding the phe- nomenon of coarticulation. To date, this transition has not been systematically modeled using an articulatory framework. Addressing this issue—which extends beyond t...

  2. [2]

    Syllable synthesis with the Maeda model 2.1. Planning of trajectories The coarticulation between vowels and consonants in VCV se- quences was modeled by [2] using an equation that describes the variations of the vocal tract shape along the glottis-to-lips axis x. A reformulation of this equation reveals that it represents a blending process between a time...

  3. [3]

    Each vowel is assigned a position on the unit circle: [y]: (0.7, 11π 6 ), [ø]: (0.3, 11π 6 ), [a]: (0.8, π), [ o]: (0.8, π 2 ), and [ u]: (1, π 3 )

    Modeling the S-shaped transition The 75 syllables V 1CV2, composed of the vowels V 1 and V2, were selected from the set/y, ø, a, o, u/. Each vowel is assigned a position on the unit circle: [y]: (0.7, 11π 6 ), [ø]: (0.3, 11π 6 ), [a]: (0.8, π), [ o]: (0.8, π 2 ), and [ u]: (1, π 3 ). Similarly, the conso- nants are positioned independently of the vowels, ...

  4. [4]

    Where do S-shaped F2 transitions originate? To provide a physical interpretation of locus equations Bailly

  5. [5]

    All simulations and figures are fully reproducible using the open-source Python software provided with this work [31]

    Conclusion From simulations of syllable structure using an articulatory model, we derive plausible deterministic principles that ac- count for key historical observations of the coarticulation phe- nomenon. All simulations and figures are fully reproducible using the open-source Python software provided with this work [31]. This approach extends articulat...

  6. [6]

    His input was es- sential for connecting the acoustic patterns with the underlying articulatory mechanisms

    Acknowledgments I gratefully thank G´erard Bailly for his insightful explana- tions regarding the physical interpretation of the locus equations based on the resonances of the vocal tract. His input was es- sential for connecting the acoustic patterns with the underlying articulatory mechanisms. I also thank Louis-Jean Bo ¨e for per- sonally sharing the M...

  7. [7]

    V owel-consonant-vowel modeling by superposition of consonant closure on vowel-to-vowel ges- tures,

    R. Carr ´e and S. Chennoukh, “V owel-consonant-vowel modeling by superposition of consonant closure on vowel-to-vowel ges- tures,” Journal of Phonetics, vol. 23, no. 1, pp. 231–241, 1995

  8. [8]

    Compensatory articulation during speech: Evidence from the analysis and synthesis of vocal-tract shapes using an ar- ticulatory model,

    S. Maeda, “Compensatory articulation during speech: Evidence from the analysis and synthesis of vocal-tract shapes using an ar- ticulatory model,” in Speech Production and Speech Modelling , ser. NATO ASI Series, D. J. Hardcastle and A. Marchal, Eds. Springer Netherlands, Dordrecht, 1990, pp. 131–149

Show all 42 references
  1. [9]

    Numerical model of coarticulation,

    S. E. G. ¨Ohman, “Numerical model of coarticulation,” The Jour- nal of the Acoustical Society of America , vol. 41, no. 2, pp. 310– 320, 1967

  2. [10]

    J. S. Perkell, Physiology of Speech Production: Results and Impli- cations of a Quantitative Cineradiographic Study , ser. Research Monograph No. 53. Cambridge, Massachusetts, and London, England: The M.I.T. Press, 1969

  3. [11]

    A parametric model of the vocal tract area function for vowel and consonant simulation,

    B. H. Story, “A parametric model of the vocal tract area function for vowel and consonant simulation,” The Journal of the Acoustical Society of America, vol. 117, no. 5, pp. 3231–3254, 04

  4. [12]

    Syllable as a synchronization mechanism that makes human speech possible,

    Y . Xu, “Syllable as a synchronization mechanism that makes human speech possible,” Brain Sciences , vol. 15, no. 1, p. 33, 2025. [Online]. Available: https://doi.org/10.3390/ brainsci15010033

  5. [13]

    Modeling consonant-vowel coarticulation for artic- ulatory speech synthesis,

    P. Birkholz, “Modeling consonant-vowel coarticulation for artic- ulatory speech synthesis,” PLOS ONE , vol. 8, no. 4, pp. 1–17, 2013

  6. [14]

    Distinctive regions and modes: A new theory of speech production,

    M. Mrayati, R. Carr ´e, and B. Gu ´erin, “Distinctive regions and modes: A new theory of speech production,” Speech Commun., vol. 7, no. 3, p. 257–286, Oct. 1988

  7. [15]

    Coarticulation in vcv utterances: Spectro- graphic measurements,

    S. E. G. ¨Ohman, “Coarticulation in vcv utterances: Spectro- graphic measurements,” The Journal of the Acoustical Society of America, vol. 39, no. 1, pp. 151–168, 1966

  8. [16]

    Locus equations in the light of articulatory modeling,

    S. Chennoukh, R. Carr ´e, and B. Lindblom, “Locus equations in the light of articulatory modeling,” The Journal of the Acoustical Society of America , vol. 102, no. 4, pp. 2380–2389, 10 1997. [Online]. Available: https://doi.org/10.1121/1.419622

  9. [17]

    Carr ´e, P

    R. Carr ´e, P. Pierre Divenyi, and M. Mryati, Speech: A Dynamic Process. Berlin,Boston: De Gruyter, 2017

  10. [18]

    Articulatory phonology: An overview,

    C. P. Browman and L. M. Goldstein, “Articulatory phonology: An overview,”Phonetica, vol. 49, pp. 155–180, 1992

  11. [19]

    A dynamical approach to gestural patterning in speech production,

    E. L. Saltzman and K. G. Munhall, “A dynamical approach to gestural patterning in speech production,”Ecological Psychology, vol. 1, no. 4, pp. 333–382, 1989

  12. [20]

    Perception of articulatory dynamics from acoustic signatures,

    K. Iskarous, H. Nam, and D. H. Whalen, “Perception of articulatory dynamics from acoustic signatures,” The Journal of the Acoustical Society of America, vol. 127, no. 6, pp. 3717–3728, 06 2010. [Online]. Available: https://doi.org/10.1121/1.3409485

  13. [21]

    Lee’s 1976 paper,

    D. N. Lee, R. J. Bootsma, M. Land, D. Regan, and R. Gray, “Lee’s 1976 paper,”Perception, vol. 38, no. 6, pp. 837–858, 2009, pMID: 19806967. [Online]. Available: https://doi.org/10.1068/pmklee

  14. [22]

    Modeling trajectories of human speech articulators using general tau theory,

    B. Elie, D. N. Lee, and A. Turk, “Modeling trajectories of human speech articulators using general tau theory,” Speech Communi- cation, vol. 151, pp. 24–38, 2023. [Online]. Available: https:// www.sciencedirect.com/science/article/pii/S0167639323000614

  15. [23]

    Notes on vocal tract computations,

    P. Badin and G. Fant, “Notes on vocal tract computations,” Royal Institute of Technology, Stockholm, Sweden, STL- QPSR 2- 3/1984, 1984

  16. [24]

    An approach to explaining formants,

    B. H. Story, “An approach to explaining formants,” Perspectives of the ASHA Special Interest Groups , vol. 9, no. 2, pp. 461–471,

  17. [25]

    An inves- tigation of locus equations as a source of relational invariance for stop place categorization,

    H. M. Sussman, H. A. McCaffrey, and S. A. Matthews, “An inves- tigation of locus equations as a source of relational invariance for stop place categorization,” The Journal of the Acoustical Society of America, vol. 90, no. 3, pp. 1309–1325, 1991

  18. [26]

    When a constriction occurs in the vocal tract, the resonant cavities are decoupled, allowing formants to be affili- ated to specific cavities

    focused on tracking the resonances of the vocal tract, based on the premise that formants are associated with the front and back cavities. When a constriction occurs in the vocal tract, the resonant cavities are decoupled, allowing formants to be affili- ated to specific cavit...

  19. [27]

    Coarticulation and theories of extrinsic timing,

    C. Fowler, “Coarticulation and theories of extrinsic timing,” Jour- nal of Phonetics, vol. 8, p. 113–133, 1980

  20. [28]

    A mod- ular architecture for articulatory synthesis from gestural specifica- tion,

    R. Alexander, T. Sorensen, A. Toutios, and S. Narayanan, “A mod- ular architecture for articulatory synthesis from gestural specifica- tion,” The Journal of the Acoustical Society of America, vol. 146, no. 6, pp. 4458–4471, 2019

  21. [29]

    From emg to formant patterns of vowels: The implication of vowel spaces,

    S. Maeda and K. Honda, “From emg to formant patterns of vowels: The implication of vowel spaces,” Phonetica, vol. 51, no. 1-3, pp. 17–29, 1994. [Online]. Available: https://doi.org/10.1159/000261955

  22. [30]

    The trough effect: Implications for speech motor programming,

    B. Lindblom, H. Sussman, G. Modarresi Ghavami, and E. Burlingame, “The trough effect: Implications for speech motor programming,” Phonetica, vol. 59, pp. 245–62, 2002

  23. [31]

    Supplementary materials: Articulatory mod- eling of the s-shaped f2 trajectories observed in ¨Ohman’s spectrographic analysis of vcv syllables,

    F. Berthommier, “Supplementary materials: Articulatory mod- eling of the s-shaped f2 trajectories observed in ¨Ohman’s spectrographic analysis of vcv syllables,” 2025, interspeech 2025, Rotterdam, The Netherlands. [Online]. Available: https://doi.org/10.5281/zenodo.15475543

  24. [32]

    Sylber: Syllabic embedding representation of speech from raw audio,

    C. J. Cho, N. Lee, A. Gupta, D. Agarwal, E. Chen, A. W. Black, and G. K. Anumanchipalli, “Sylber: Syllabic embedding representation of speech from raw audio,” 2025. [Online]. Available: https://arxiv.org/abs/2410.07168

  25. [33]

    Coarticulation as incomplete interpolation,

    B. Lindblom, D. Krull, and H. M. Sussman, “Coarticulation as incomplete interpolation,” in Proceedings of the FONETIK 2010. Sweden: Department of Phonetics, Center for Language & Liter- ature, Lund University, 2010

  26. [35]

    Caracterisation of formant trajectories by tracking vocal tract resonances,

    G. Bailly, “Caracterisation of formant trajectories by tracking vocal tract resonances,” in Levels in Speech Communication: Relations and Interactions , C. Sorin, J. Mariani, H. Meloni, and J. Schoentgen, Eds. El- sevier, 1995, pp. 91–102. [Online]. Available: https: //www.res...

  27. [36]

    Review of text-to-speech conversion for english,

    D. H. Klatt, “Review of text-to-speech conversion for english,” The Journal of the Acoustical Society of America , vol. 82, no. 3, pp. 737–793, 09 1987. [Online]. Available: https: //doi.org/10.1121/1.395275

  28. [37]

    Acoustic loci and transitional cues for consonants,

    P. C. Delattre, A. M. Liberman, and F. S. Cooper, “Acoustic loci and transitional cues for consonants,” The Journal of the Acoustical Society of America , vol. 27, no. 4, pp. 769–773, 06

  29. [39]

    Dissecting coarticulation: How locus equations happen,

    B. Lindblom and H. M. Sussman, “Dissecting coarticulation: How locus equations happen,” Journal of Phonetics , vol. 40, no. 1, pp. 1–19, 2012. [Online]. Available: https://www. sciencedirect.com/science/article/pii/S0095447011000933

  30. [40]

    Locus equations are an acoustic expression of articulator synergy,

    K. Iskarous, C. A. Fowler, and D. H. Whalen, “Locus equations are an acoustic expression of articulator synergy,” The Journal of the Acoustical Society of America, vol. 128, no. 4, pp. 2021–2032, 2010

  31. [41]

    Software for syllable synthesis and sup- plementary for

    F. Berthommier, “Software for syllable synthesis and sup- plementary for ”articulatory modeling of the s-shaped f2 trajectories observed in ¨Ohman’s spectrographic analysis of vcv syllables”,” 2025, available on Zenodo. [Online]. Available: https://doi.org/10.5281/zenodo.15527267

  32. [1955]

    Available: 10.1121/1.1908024

    [Online]. Available: 10.1121/1.1908024

  33. [2005]

    Available: https://doi.org/10.1121/1.1869752

    [Online]. Available: https://doi.org/10.1121/1.1869752

  34. [2024]

    Available: https://pubs.asha.org/doi/abs/10.1044/ 2023 PERSP-23-00200

    [Online]. Available: https://pubs.asha.org/doi/abs/10.1044/ 2023 PERSP-23-00200

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.