Pith. sign in

REVIEW 3 major objections 6 minor 70 references

GIRAFE: Glottal Imaging Dataset for Advanced Segmentation, Analysis, and Facilitative Playbacks Evaluation

T0 review · 3 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read GIRAFE is a publicly released dataset of 65 high-speed color videoendoscopic vocal-fold recordings with expert manual glottal-gap segmentations, automatic segmentation baselines, and facilitative playbacks, designed to fill the gap left…

desk verdict A useful new dataset that overstates its own manual-annotation coverage in the abstract; fix that and it's a solid contribution. read the letter →

arxiv 2412.15054 v1 pith:HPHBPLNF submitted 2024-12-19 cs.CV cs.AIcs.SDeess.AS

classification cs.CVcs.AIcs.SDeess.AS
keywords GIRAFEhigh-speedvideoendoscopyvocalfoldsglottalgapsegmentationsemanticfacilitativeplaybacksmedicalimagingdatasetdeeplearning
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper introduces GIRAFE, a data repository meant to advance the automatic semantic segmentation of the glottal gap in high-speed videoendoscopic recordings of the vocal folds and to support the synthesis and evaluation of facilitative playbacks. It argues that the lack of public annotated datasets in this domain has limited reproducibility and the training of deep-learning models, and that GIRAFE provides a much-needed complement to the existing BAGLS dataset by adding color recordings, manual masks, and playback outputs. The dataset comprises 65 recordings from 50 patients, with manual segmentations covering one complete glottal cycle per annotated recording, and it already has been used in several published studies.

What carries the argument

The central object is the GIRAFE repository itself: a structured corpus of raw HSV videos (AVI), per-patient JSON metadata, manual and automatic segmentation masks stored as PKL files, facilitative playback outputs, and a training directory with 760 PNG frames and their corresponding labels. The load-bearing element is the expert manual annotation of the glottal gap, which serves as ground truth for all reported DICE, Jaccard, recall, and precision scores, and from which the facilitative playbacks are synthesized.

What would settle it

Concrete check: have two or more independent experts re-annotate a random subset of the 760 glottal-gap frames and compute inter-annotator DICE agreement; if agreement is low or if the released masks show clear systematic errors, the dataset's validity as a gold standard is undermined. Additionally, a simple audit of the released repository can verify whether manual masks exist for all 65 recordings or only for the 38 stated in the methods section.

Watch

Extended reading notes

Core claim

On its own terms, GIRAFE establishes a new public benchmark for glottal-gap segmentation: 65 high-speed color videoendoscopic recordings acquired at 4,000 fps from 50 patients, spanning healthy voices, diagnosed voice disorders, and unknown health states. The manual gold standard consists of expert delineations of the glottal gap, verified by an otolaryngologist, produced for 38 of the 65 videos and totaling 760 frames; the abstract of the paper describes all 65 recordings as manually annotated. Alongside the manual masks, the repository provides segmentations from two classical image-processing methods (InP and Loh), deep-learning baselines (UNet and SwinUnetV2), facilitative playbacks (GAW, GVG, PVG, and digital kymograms), and predefined training/validation/test splits with JSON metadata, all released openly.

Load-bearing premise

The load-bearing premise is that the single expert's manual delineations of the glottal gap are an accurate gold standard; since no inter-annotator agreement is reported and only one otolaryngologist verification pass is mentioned, any systematic bias in those masks would propagate into every reported segmentation score and into the dataset's value as a benchmark.

Editorial extensions

If this is right

  • Researchers can train and compare deep-learning glottal segmentation models on a standardized, openly available benchmark with predefined splits, enabling fair and reproducible evaluation.
  • The color recordings and the inclusion of playback outputs allow downstream evaluation of segmentation quality through facilitative playbacks, not only through pixel-wise metrics.
  • GIRAFE complements BAGLS by adding color data and playback-based analysis, potentially supporting cross-dataset generalization studies and foundation-model fine-tuning in laryngeal imaging.
  • The provided baseline results for InP, Loh, UNet, and SwinUnetV2 give immediate reference points for future segmentation methods on this dataset.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The discrepancy between the abstract's claim that all 65 recordings were manually annotated and the methods section's statement that manual segmentation was performed on 38 recordings should be resolved by users checking the released metadata; if only 38 have manual masks, the effective supervised training set is smaller than the headline number suggests.
  • Because the ground truth comes from a single expert with no reported inter-annotator agreement, benchmark scores on GIRAFE should be interpreted as reflecting one expert's delineation style; an independent multi-annotator study would materially strengthen the dataset's reliability.
  • The color versus grayscale advantage could be tested directly by comparing segmentation performance on GIRAFE against BAGLS under matched training conditions, a comparison the paper motivates but does not perform.
  • The inclusion of facilitative playbacks opens the possibility of treating playback fidelity as a task-specific evaluation metric, which may reward segmentations that preserve temporal and morphological features relevant to clinical assessment.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper introduces GIRAFE, a publicly available dataset of 65 high-speed videoendoscopic (HSV) recordings of the vocal folds from 50 patients, intended to support semantic segmentation of the glottal gap and evaluation of facilitative playbacks (GAW, GVG, PVG). The repository includes raw AVI videos, JSON metadata, manual segmentations (reported in Section 2.3 as 38 of the 65 videos, yielding 760 frames), automatic/semi-automatic segmentations from the authors' InP method and the Loh method, and the derived playbacks. A technical validation trains UNet and SwinUnetV2 on the manual masks and compares them against InP and Loh on four test patients (20 frames each) in Table 3. The paper positions GIRAFE as a complement to the existing BAGLS dataset, adding color recordings and playback evaluation, and provides open data and code.

Significance. If the stated coverage is corrected, GIRAFE is a useful and reproducible contribution to laryngeal imaging: it is, to the authors' knowledge, the first public color HSV dataset with glottal-gap annotations, and it ships open code, trained baseline models, and a documented directory structure. The dataset has already supported several prior studies, which is a positive indicator of practical utility. The main significance is as a benchmark resource rather than a methodological advance; the DL validation is explicitly illustrative. However, the abstract's claim that all 65 recordings were manually annotated is contradicted by the body of the paper, and this discrepancy directly affects what a user receives and how the dataset's value is advertised. With the correction, the resource is of moderate but real value to the community.

major comments (3)
  1. [Abstract; Section 2.3; Section 2.6] The abstract states that 'All of them were manually annotated by an expert, including the masks corresponding to the semantic segmentation of the glottal gap,' and Section 1 repeats the promise of 'annotations meticulously performed and rigorously reviewed.' This is contradicted by Section 2.3, which says 'Manual segmentation was conducted on 38 out of the 65 videos,' and by Section 2.6, which reports that the Manual_Segmentation folder contains exactly 38 subfolders. The discrepancy is load-bearing because it determines the dataset's actual content: 27 of the 65 recordings lack manual masks and can only be used with the InP (and possibly Loh) segmentations. The abstract and all downstream descriptions must be corrected to state the 38-video coverage, and the dataset documentation should report the aggregate count of the Man=Yes metadata flag so users can verify the coverage without opening each recording.
  2. [Section 2.8; Table 3] The technical validation uses only four test recordings, each with 20 frames (80 frames total per model), and reports no inter-annotator agreement, confidence intervals, or statistical significance for the DICE/Jaccard metrics in Table 3. The paper explicitly disclaims that the validation is 'not intended to be exhaustive,' but the claim that the dataset provides a 'gold standard' for supervised segmentation is weakened by the absence of any evidence about the reliability of the manual masks. I recommend adding at least a small inter-annotator study (even on a subset) and reporting per-frame variability, or alternatively reframing the manual masks as single-expert annotations and softening the 'gold standard' language throughout the paper.
  3. [Section 2.3] The sentence 'These segmentations were further validated using two established image processing techniques: InP [22] and Loh [20]' is misleading. Comparing automatic segmentations (InP, Loh) against the manual masks is an evaluation of those automatic methods, not a validation of the manual segmentations. If the intent is to demonstrate that the manual masks are reliable, the appropriate evidence is clinician re-annotation, inter-annotator agreement, or some other independent verification, none of which is reported. Please rephrase this sentence to describe what was actually done, and avoid implying that agreement between automatic methods and manual masks establishes the correctness of the masks.
minor comments (6)
  1. [Section 2.3] The phrase 'PKL Phyton ® format' contains a typo: 'Phyton' should be 'Python'.
  2. [Section 2.6] There is a duplicated article in 'analysis of the the vocal folds trajectory'; also, 'glottal axis’s' should be 'glottal axis' or 'glottal-axis' to avoid a possessive apostrophe error.
  3. [Section 2.8] The heading 'Experimental design, Material and Methods' appears mid-paper without a section number and is not listed in the initial section outline; please either format it consistently or integrate it into the preceding subsection.
  4. [Figure 5] The database tree structure is visually cluttered, especially the parent folder labels for multiple patients, making it difficult to infer the actual directory nesting; a simplified tree or a listing of folder names with the number of entries per folder would be clearer.
  5. [CRediT Author Statement] The CRediT statement credits 'D.P.S.' with data collection, but no such author is listed among the manuscript's authors; the contribution should be attributed in the author list or acknowledgments, with the consent of that person if appropriate.
  6. [Section 2.7] The text says the Training directory contains 760 images 'corresponding to the same number of consecutive frames extracted from the original sequences'; please clarify how the 760 frames are distributed across the 38 manually segmented videos (e.g., 20 consecutive frames per video), since the phrasing is ambiguous.

Circularity Check

0 steps flagged · score 0.0 of 10

No circular derivation found; the paper's 65-vs-38 manual annotation discrepancy is a data-description error, not a circular step.

full rationale

GIRAFE is a dataset release rather than a derivation of predictions from first principles, so the standard circularity patterns do not apply. The manual glottal-gap masks are presented as ground truth, and the quantitative validation in Table 3 evaluates UNet, SwinUnetV2, InP, and Loh on held-out frames against those masks; this is ordinary benchmark practice, not a case of a fitted parameter being renamed as a prediction. The authors' self-citations, including the InP baseline and the statement that 'Various subsets of the GIRAFE corpus have been extensively used in the past as an experimental basis for several research works [21, 22, 10, 1, 19]', are present but not load-bearing: the dataset's content and availability do not depend on those prior works, and the InP method is explicitly applied as one of several segmentation approaches rather than used to define the manual ground truth. The one substantive defect is an internal inconsistency: the abstract claims 'All of them were manually annotated by an expert', while Section 2.3 states 'Manual segmentation was conducted on 38 out of the 65 videos' and Section 2.6 confirms that the Manual_Segmentation folder contains 38 subfolders. This is a factual and data-description accuracy issue that materially affects what users receive, but it is not a circularity: no claimed result reduces to its own input by construction. The paper should correct the abstract and clearly report the count of Man=Yes recordings, but this does not raise the circularity score.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The dataset's usefulness rests on untested domain assumptions about annotation quality and representativeness; no free parameters or invented entities are introduced.

assumptions (3)
  • domain assumption Expert manual annotations of the glottal gap are a valid gold standard for training and evaluation.
    Section 2.3 treats the expert masks as ground truth for all baselines and DL models; no inter-annotator agreement or independent verification is reported, so the assumption is untested.
  • domain assumption One complete glottal cycle of about 20 frames per patient captures enough variability to train and evaluate segmentation models.
    Manual segmentation was performed on 20 frames per recording (760 frames from 38 videos), and these frames are used as the training set without evidence that a single cycle is representative of the full sequence variability.
  • domain assumption The recorded metadata (age, sex, disorder status) are accurate as provided by the clinical records.
    The dataset relies on hospital records for disorder labels and demographics; 24 of 65 recordings have unknown health status, which is acknowledged but limits diagnostic-label reliability.

how reviews work

0 comments
Cite this review

Pith. "Pith review of GIRAFE: Glottal Imaging Dataset for Advanced Segmentation, Analysis, and Facilitative Playbacks Evaluation." pith.science (2026). https://pith.science/paper/HPHBPLNF

@misc{pith2026241215054,
  author       = {Pith},
  title        = {Pith review of: GIRAFE: Glottal Imaging Dataset for Advanced Segmentation, Analysis, and Facilitative Playbacks Evaluation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/HPHBPLNF}},
  note         = {Machine review of arXiv:2412.15054}
}
read the original abstract

The advances in the development of Facilitative Playbacks extracted from High-Speed videoendoscopic sequences of the vocal folds are hindered by a notable lack of publicly available datasets annotated with the semantic segmentations corresponding to the area of the glottal gap. This fact also limits the reproducibility and further exploration of existing research in this field. To address this gap, GIRAFE is a data repository designed to facilitate the development of advanced techniques for the semantic segmentation, analysis, and fast evaluation of High-Speed videoendoscopic sequences of the vocal folds. The repository includes 65 high-speed videoendoscopic recordings from a cohort of 50 patients (30 female, 20 male). The dataset comprises 15 recordings from healthy controls, 26 from patients with diagnosed voice disorders, and 24 with an unknown health condition. All of them were manually annotated by an expert, including the masks corresponding to the semantic segmentation of the glottal gap. The repository is also complemented with the automatic segmentation of the glottal area using different state-of-the-art approaches. This data set has already supported several studies, which demonstrates its usefulness for the development of new glottal gap segmentation algorithms from High-Speed-Videoendoscopic sequences to improve or create new Facilitative Playbacks. Despite these advances and others in the field, the broader challenge of performing an accurate and completely automatic semantic segmentation method of the glottal area remains open.

Figures

Figures reproduced from arXiv: 2412.15054 by the authors.

Figure 1
Figure 1. Illustration of a complete glottal cycle extracted from laryngeal HSV, along with three FP synthesized from [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Workflow for generating the GIRAFE dataset. Participants are of varying ages, gender, and health conditions, [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Age and sex distribution of the recordings in the GIRAFE dataset. [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figures from the paper (5 more)
Figure 4
Figure 4. Figure 4: Illustration of four cases with voice disorders: (a) cervicotomy for cervical hernia with postoperative [PITH_FULL_IMAGE:figures/full_fig_p007_4.png]
Figure 5
Figure 5. Figure 5: Database tree structure: hierarchical representation of the data records and their organization within the [PITH_FULL_IMAGE:figures/full_fig_p008_5.png]
Figure 6
Figure 6. Figure 6: Visual illustration of the dictionary keys in [PITH_FULL_IMAGE:figures/full_fig_p010_6.png]
Figure 7
Figure 7. Figure 7: Boxplots illustrating the DICE metric for the four test patients. [PITH_FULL_IMAGE:figures/full_fig_p012_7.png]
Figure 8
Figure 8. Figure 8: Visual representation of segmentation results, with blue contours indicating ground truth ( [PITH_FULL_IMAGE:figures/full_fig_p014_8.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

70 extracted references · 61 canonical work pages

  1. [22]

    Andrade-Miranda, J

    G. Andrade-Miranda, J. I. Godino-Llorente, Glottal gap tracking by a continuous background modeling using inpainting, Medical & Biological Engineering & Computing 55 (12) (2017) 2123–2141. URL https://doi.org/10.1007/s11517-017-1652-8

  2. [20]

    Lohscheller, H

    J. Lohscheller, H. Toy, F. Rosanowski, U. Eysholdt, M. Dollinger, Clinically evaluated procedure for the recon- struction of vocal fold vibrations from endoscopic digital high-speed videos, Medical Image Analysis 11 (4) (2007) 400–413

  3. [1]

    Andrade-Miranda, Y

    G. Andrade-Miranda, Y . Stylianou, D. D. Deliyski, J. I. Godino-Llorente, N. Henrich Bernardoni, Laryngeal image processing of vocal folds motion, Applied Sciences 10 (5) (2020). URL https://www.mdpi.com/2076-3417/10/5/1556

  4. [2]

    J. I. Godino-Llorente, N. Sáenz-Lechón, V . Osma-Ruiz, S. Aguilera-Navarro, P. Gómez-Vilda, An integrated tool for the diagnosis of voice disorders, Medical Engineering & Physics 28 (3) (2006) 276–289

  5. [3]

    Leppävuori, G

    M. Leppävuori, G. Andrade-Miranda, N. H. Bernardoni, A.-M. Laukkanen, A. Geneid, Characterizing vocal- fold dynamics in singing vocal modes from complete vocal technique using high-speed laryngeal imaging and electroglottographic analysis, in: 13th Pan-European V oice Conference, 2019

  6. [4]

    C. P.-L. Marion Beaud, Benoît Amy de la Bretèque, N. H. Bernardoni, Clinical characteristics of singers attending a phoniatric outpatient clinic, Logopedics Phoniatrics V ocology 47 (3) (2022) 209–218, pMID: 34110262

  7. [5]

    Henrich Bernardoni, A

    N. Henrich Bernardoni, A. Di Corcia, F. Fussi, Laryngeal mechanisms in modern singing, in: COMET Virtual annual conference, Tel-Aviv, Israel, 2022

  8. [6]

    M. R. Mehl, S. Vazire, N. Ramírez-Esparza, R. B. Slatcher, J. W. Pennebaker, Are women really more talkative than men?, Science 317 (5834) (2007) 82

Show all 70 references
  1. [7]

    I. R. Titze, Principles of voice production, Prentice-Hall Inc, 1993

  2. [8]

    A. E. Aronson, D. Bless, Clinical voice disorders, Thieme Medical Publishers, New York, USA, 2009

  3. [9]

    Laver, The Phonetic Description of V oice Quality, V ol

    J. Laver, The Phonetic Description of V oice Quality, V ol. 31 of Cambridge Studies in Linguistics, Cambridge University Press, 2009

  4. [10]

    Andrade-Miranda, Analyzing of the vocal fold dynamics using laryngeal videos, Ph.D

    G. Andrade-Miranda, Analyzing of the vocal fold dynamics using laryngeal videos, Ph.D. thesis, Universidad Politécnica de Madrid (2017)

  5. [11]

    J. G. Švec, F. Šram, H. K. Schutte, Videokymography in V oice Disorders: What to Look For?, Annals of Otology, Rhinology and Laryngology 116 (3) (2007) 172–180

  6. [12]

    H. S. Shaw, D. D. Deliyski, Mucosal wave: A normophonic study across visualization techniques, Journal of V oice 22 (1) (2008) 23 – 33

  7. [13]

    V oigt, M

    D. V oigt, M. Döllinger, T. Braunschweig, A. Yang, U. Eysholdt, J. Lohscheller, Classification of functional voice disorders based on phonovibrograms, Artificial Intelligence in Medicine 49 (1) (2010) 51–59

  8. [14]

    A. P. Pinheiro, D. E. Stewart, C. D. Maciel, J. C. Pereira, S. Oliveira, Analysis of nonlinear dynamics of vocal folds using high-speed video observation and biomechanical modeling, Digital Signal Processing 22 (2) (2012) 304–313

  9. [15]

    Yamauchi, H

    A. Yamauchi, H. Yokonishi, H. Imagawa, K.-I. Sakakibara, T. Nito, N. Tayama, T. Yamasoba, Quantification of vocal fold vibration in various laryngeal disorders using high-speed digital imaging, Journal of V oice 30 (2) (2016) 205–214

  10. [16]

    Crocker, L

    C. Crocker, L. E. Toles, R. A. Morrison, A. C. Shembel, Relationships between vocal fold adduction patterns, vocal acoustic quality, and vocal effort in individuals with and without hyperfunctional voice disorders, Journal of V oice (2024)

  11. [17]

    Andrade-Miranda, N

    G. Andrade-Miranda, N. Henrich, J. I. Godino-Llorente, Optical-flow kymograms and glottovibrograms: A new way to present high-speed data for laryngeal assessment, in: 9th International workshop, Models and Analysis of V ocal Emissions for Biomedical Applications, MA VEBA, Fire...

  12. [18]

    Andrade-Miranda, N

    G. Andrade-Miranda, N. H. Bernardoni, J. I. Godino-Llorente, A new technique for assessing glottal dynamics in speech and singing by means of optical-flow computation, in: 16th Annual Conference of the International Speech Communication Association (INTERSPEECH), Dresden, Germ...

  13. [19]

    Andrade-Miranda, N

    G. Andrade-Miranda, N. Henrich Bernardoni, J. I. Godino-Llorente, Synthesizing the motion of the vocal folds using optical flow based techniques, Biomedical Signal Processing and Control 34 (2017) 25–35

  14. [21]

    Andrade-Miranda, J

    G. Andrade-Miranda, J. I. Godino-Llorente, L. Moro-Velázquez, J. A. Gómez-García, An automatic method to detect and track the glottal gap from high speed videoendoscopic images, BioMedical Engineering OnLine 14 (1) (2015) 100

  15. [23]

    S. Z. Karakozoglou, N. Henrich, C. D’Alessandro, Y . Stylianou, Automatic glottal segmentation using local- based active contours and application to glottovibrography, Speech Communication 54 (5) (2012) 641–654

  16. [24]

    Lohscheller, U

    J. Lohscheller, U. Eysholdt, Phonovibrogram visualization of entire vocal fold dynamics, The Laryngoscope 118 (4) (2008) 753–758

  17. [25]

    J. G. Švec, H. K. Schutte, Videokymography: High-speed line scanning of vocal fold vibration, Journal of V oice 10 (2) (1996) 201 – 205

  18. [26]

    J. G. Švec, F. Šram, Kymographic imaging of the vocal folds oscillations, in: 7th International Conference on Spoken Language Processing, V ol. 2, 2002, pp. 957–960

  19. [27]

    Wurzbacher, R

    T. Wurzbacher, R. Schwarz, M. Döllinger, U. Hoppe, U. Eysholdt, J. Lohscheller, Model-based classification of nonstationary vocal fold vibrations. model-based classification of nonstationary vocal fold vibrations., The Journal of the Acoustical Society of America 120 (2) (2006...

  20. [28]

    Lohscheller, U

    J. Lohscheller, U. Eysholdt, Phonovibrography: Mapping high-speed movies of vocal fold vibrations into 2- D diagrams for visualizing and analyzing the underlying laryngeal dynamics, IEEE Transactions on Medical Imaging 27 (3) (2008) 300–309

  21. [29]

    D. D. Mehta, D. D. Deliyski, T. F. Quatieri, R. E. Hillman, Automated measurement of vocal fold vibratory asymmetry from high-speed videoendoscopy recordings, Journal of Speech, Language, and Hearing Research 54 (1) (2013) 47–54

  22. [30]

    Lohscheller, J

    J. Lohscheller, J. G. Švec, M. Döllinger, V ocal fold vibration amplitude, open quotient, speed quotient and their variability along glottal length: kymographic data from normal subjects, Logopedics Phoniatrics V ocology 38 (4) (2013) 182–192

  23. [31]

    C. T. Herbst, J. Lohscheller, J. G. Švec, N. Henrich, G. Weissengruber, W. T. Fitch, Glottal opening and closing events investigated by electroglottography and super-high-speed video recordings., The Journal of Experimental Biology 217 (6) (2014) 955–963

  24. [32]

    Y . Yan, E. Damrose, D. Bless, Functional analysis of voice using simultaneous high-speed imaging and acoustic recordings., Journal of V oice 21 (5) (2007) 604–616

  25. [33]

    Ahmad, Y

    k. Ahmad, Y . Yan, D. Bless, V ocal fold vibratory characteristics in normal female speakers from high-speed digital imaging, Journal of V oice 26 (2) (2012) 239–253

  26. [34]

    R. R. Patel, K. Forrest, D. Hedges, Relationship between acoustic voice onset and offset and selected instances of oscillatory onset and offset in young healthy men and women, Journal of V oice 31 (3) (2017) 389.e9 – 389.e17

  27. [35]

    Schlegel, M

    P. Schlegel, M. Stingl, M. Kunduk, S. Kniesburges, C. Bohr, M. Döllinger, Dependencies and ill-designed pa- rameters within high-speed videoendoscopy and acoustic signal analysis, Journal of V oice (2018)

  28. [36]

    C. R. Krausert, A. E. Olszewski, L. N. Taylor, J. S. McMurray, S. H. Dailey, J. J. Jiang, Mucosal wave measure- ment and visualization techniques, Journal of V oice 25 (4) (2011) 395 – 405

  29. [37]

    Kaneko, O

    M. Kaneko, O. Shiromoto, M. Fujiu-Kurachi, Y . Kishimoto, I. Tateya, S. Hirano, Optimal duration for voice rest after vocal fold surgery: Randomized controlled clinical study, Journal of V oice 31 (1) (2017) 97 – 103

  30. [38]

    L. Li, Y . Zhang, A. L. Maytag, J. J. Jiang, Quantitative study for the surface dehydration of vocal folds based on high-speed imaging, Journal of V oice 29 (4) (2015) 403 – 409

  31. [39]

    C. T. Herbst, J. G. Švec, J. Lohscheller, R. Frey, M. Gumpenberger, A. S. Stoeger, W. T. Fitch, Complex vibratory patterns in an elephant larynx, Journal of Experimental Biology 216 (21) (2013) 4054–4064. doi:10.1242/ jeb.091009. 16 GIRAFE: Glottal Imaging Repository for Advan...

  32. [40]

    C. P. H. Elemans, J. H. Rasmussen, C. T. Herbst, D. N. Düring, S. A. Zollinger, H. Brumm, K. H. Srivastava, N. Svane, M. Ding, O. N. Larsen, S. J. Sober, J. G. Svec, Universal mechanisms of sound production and control in birds and mammals, Nature communications 6 (2015) 8978

  33. [41]

    C. T. Herbst, Biophysics of V ocal Production in Mammals, Springer International Publishing, Cham, 2016, pp. 159–189

  34. [42]

    V oigt, M

    D. V oigt, M. Döllinger, U. Eysholdt, A. Yang, E. Gürlen, J. Lohscheller, Objective detection and quantification of mucosal wave propagation, The Journal of the Acoustical Society of America 128 (5) (2010) EL347–EL353

  35. [43]

    Schlegel, S

    P. Schlegel, S. Kniesburges, S. Dürr, et al., Machine learning based identification of relevant parameters for functional voice disorders derived from endoscopic high-speed recordings, Scientific Reports 10 (2020) 10517. URL https://doi.org/10.1038/s41598-020-66405-y

  36. [44]

    Unger, J

    J. Unger, J. Lohscheller, M. Reiter, K. Eder, C. S. Betz, M. Schuster, A Noninvasive Procedure for Early-Stage Discrimination of Malignant and Precancerous V ocal Fold Lesions Based on Laryngeal Dynamics Analysis, Cancer Research 75 (1) (2015) 31–39

  37. [45]

    A. Kist, S. Dürr, A. Schützenberger, et al., Openhsv: an open platform for laryngeal high-speed videoendoscopy, Scientific Reports 11 (2021) 13760. URL https://doi.org/10.1038/s41598-021-93149-0

  38. [46]

    A. M. Kist, K. Breininger, M. Dörrich, et al., A single latent channel is sufficient for biomedical glottis segmen- tation, Scientific Reports 12 (2022) 14292. URL https://doi.org/10.1038/s41598-022-17764-1

  39. [47]

    Chen, P.-Y

    I.-M. Chen, P.-Y . Yeh, Y .-C. Hsieh, T.-C. Chang, S. Shih, W.-F. Shen, C.-L. Chin, 3d vosnet: Segmentation of endoscopic images of the larynx with subsequent generation of indicators, Heliyon 9 (3) (2023) e14242. URL https://www.sciencedirect.com/science/article/pii/S2405844023014494

  40. [48]

    Pedersen, C

    M. Pedersen, C. F. Larsen, B. Madsen, et al., Localization and quantification of glottal gaps on deep learning segmentation of vocal folds, Scientific Reports 13 (2023) 878. URL https://doi.org/10.1038/s41598-023-27980-y

  41. [49]

    Conze, G

    P.-H. Conze, G. Andrade-Miranda, V . K. Singh, V . Jaouen, D. Visvikis, Current and emerging trends in medical image segmentation with deep learning, IEEE Transactions on Radiation and Plasma Medical Sciences 7 (6) (2023) 545–569

  42. [50]

    Andrade-Miranda, V

    G. Andrade-Miranda, V . Jaouen, O. Tankyevych, C. Cheze Le Rest, D. Visvikis, P.-H. Conze, Multi-modal medical transformers: A meta-analysis for medical image segmentation in oncology, Computerized Medical Imaging and Graphics 110 (2023) 102308. URL https://www.sciencedirect.c...

  43. [51]

    B. Azad, R. Azad, S. Eskandari, A. Bozorgpour, A. Kazerouni, I. Rekik, D. Merhof, Foundational models in medical imaging: A comprehensive survey and future vision, arXiv preprint arXiv:2310.18689 (2023)

  44. [52]

    H. Wang, P. K. A. Vasu, F. Faghri, R. Vemulapalli, M. Farajtabar, S. Mehta, M. Rastegari, O. Tuzel, H. Pouransari, Sam-clip: Merging vision foundation models towards semantic and spatial understanding, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern R...

  45. [53]

    Alzubaidi, J

    L. Alzubaidi, J. Bai, A. Al-Sabaawi, et al., A survey on deep learning tools dealing with data scarcity: definitions, challenges, solutions, tips, and applications, Journal of Big Data 10 (2023) 46. URL https://doi.org/10.1186/s40537-023-00727-2

  46. [54]

    Andrade-Miranda, P

    G. Andrade-Miranda, P. Soto Vega, K. Taguelmimt, H.-P. Dang, D. Visvikis, J. Bert, Exploring transformer reliability in clinically significant prostate cancer segmentation: A comprehensive in-depth investigation, Com- puterized Medical Imaging and Graphics - (2024) –

  47. [55]

    A. E. Kavur, N. S. Gezer, M. Barı¸ s, S. Aslan, P.-H. Conze, V . Groza, D. D. Pham, S. Chatterjee, P. Ernst, S. Özkan, et al., Chaos challenge-combined (ct-mr) healthy abdominal organ segmentation, Medical Image Analysis 69 (2021) 101950

  48. [56]

    Y . Ji, H. Bai, J. Yang, C. Ge, Y . Zhu, R. Zhang, Z. Li, L. Zhang, W. Ma, X. Wan, et al., Amos: A large-scale abdominal multi-organ benchmark for versatile medical image segmentation, arXiv preprint arXiv:2206.08023 (2022)

  49. [57]

    J. Ma, Y . Zhang, S. Gu, C. Ge, S. Ma, A. Young, C. Zhu, K. Meng, X. Yang, Z. Huang, et al., Unleashing the strengths of unlabeled data in pan-cancer abdominal organ quantification: the flare22 challenge, arXiv preprint arXiv:2308.05862 (2023). 17 GIRAFE: Glottal Imaging Repos...

  50. [58]

    Koitka, G

    S. Koitka, G. Baldini, L. Kroll, et al., Saros: A dataset for whole-body region and organ segmentation in ct imaging, Scientific Data 11 (2024) 483. URL https://doi.org/10.1038/s41597-024-03337-6

  51. [59]

    J. Liu, Y . Zhang, J.-N. Chen, J. Xiao, Y . Lu, B. A. Landman, Y . Yuan, A. Yuille, Y . Tang, Z. Zhou, Clip-driven universal model for organ segmentation and tumor detection, arXiv preprint arXiv:2301.00785 (2023)

  52. [60]

    Gómez, A

    P. Gómez, A. M. Kist, P. Schlegel, et al., Bagls, a multihospital benchmark for automatic glottis segmentation, Scientific Data 7 (2020) 186. URL https://doi.org/10.1038/s41597-020-0526-3

  53. [61]

    M. A. Mazurowski, H. Dong, H. Gu, J. Yang, N. Konz, Y . Zhang, Segment anything model for medical image analysis: An experimental study, Medical Image Analysis 89 (2023) 102918

  54. [62]

    Siblini, G

    L. Siblini, G. Andrade-Miranda, K. Taguelmimt, D. Visvkis, J. Bert, Optimal prompting in sam for few-shot and weakly supervised medical image segmentation, in: MICCAI 2024 2nd International Workshop on Foundation Models for General Medical AI. Accepted on July 15, 2024

  55. [63]

    Ronneberger, P.Fischer, T

    O. Ronneberger, P.Fischer, T. Brox, U-net: Convolutional networks for biomedical image segmentation, in: Medical Image Computing and Computer-Assisted Intervention (MICCAI), V ol. 9351 of LNCS, Springer, 2015, pp. 234–241, (available on arXiv:1505.04597 [cs.CV]). URL http://lm...

  56. [64]

    Z. Liu, H. Hu, Y . Lin, Z. Yao, Z. Xie, Y . Wei, J. Ning, Y . Cao, Z. Zhang, L. Dong, et al., Swin Transformer v2: Scaling up capacity and resolution, in: IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 12009–12019

  57. [65]

    Deliyski, P

    D. Deliyski, P. P. Petrushev, H. Bonilha, T. T. Gerlach, B. Martin-Harris, R. E. Hillman, Clinical implementation of laryngeal high-speed videoendoscopy: Challenges and evolution, Folia Phoniatrica et Logopaedica 60 (1) (2008) 33–44

  58. [66]

    Kendall, R

    K. Kendall, R. Leonard, Laryngeal Evaluation: Indirect Laryngoscopy to High-speed Digital Imaging, Thieme, 2010, Ch. 14. Normal Glottic Configuration, pp. 245–270

  59. [67]

    Braunschweig, J

    T. Braunschweig, J. Flaschka, P. Schelhorn-Neise, M. Döllinger, High-speed video analysis of the phonation onset, with an application to the diagnosis of functional dysphonias, Medical Engineering and Physics 30 (1) (2008) 59–66

  60. [68]

    P. Birkholz, Glottalimageexplorer - an open source tool for glottis segmentation in endoscopic high-speed videos of the vocal folds, in: Studientexte zur Sprachkommunikation: Elektronische Sprachsignalverarbeitung 2016, Dresden, Germany, 2016

  61. [69]

    M. J. Cardoso, W. Li, R. Brown, N. Ma, E. Kerfoot, Y . Wang, B. Murrey, A. Myronenko, C. Zhao, D. Yang, V . Nath, Y . He, Z. Xu, A. Hatamizadeh, A. Myronenko, W. Zhu, Y . Liu, M. Zheng, Y . Tang, I. Yang, M. Zephyr, B. Hashemian, S. Alle, M. Z. Darestani, C. Budd, M. Modat, T....

  62. [70]

    Andrade-Miranda, D

    G. Andrade-Miranda, D. Poletti-Serafini, K. Chatzipapas, J. D. A. Londoño, , J. I. Godino-Llorente, GIRAFE: Glottal Imaging Repository for Advanced Segmentation, Analysis, and Facilitative Playbacks Evaluation (1.0.0) (Sep. 2024). doi:https://doi.org/10.5281/zenodo.13773163. 18

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.