REVIEW 3 major objections 6 minor 70 references
GIRAFE: Glottal Imaging Dataset for Advanced Segmentation, Analysis, and Facilitative Playbacks Evaluation
T0 review · 3 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read GIRAFE is a publicly released dataset of 65 high-speed color videoendoscopic vocal-fold recordings with expert manual glottal-gap segmentations, automatic segmentation baselines, and facilitative playbacks, designed to fill the gap left…
desk verdict A useful new dataset that overstates its own manual-annotation coverage in the abstract; fix that and it's a solid contribution. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the GIRAFE repository itself: a structured corpus of raw HSV videos (AVI), per-patient JSON metadata, manual and automatic segmentation masks stored as PKL files, facilitative playback outputs, and a training directory with 760 PNG frames and their corresponding labels. The load-bearing element is the expert manual annotation of the glottal gap, which serves as ground truth for all reported DICE, Jaccard, recall, and precision scores, and from which the facilitative playbacks are synthesized.
What would settle it
Concrete check: have two or more independent experts re-annotate a random subset of the 760 glottal-gap frames and compute inter-annotator DICE agreement; if agreement is low or if the released masks show clear systematic errors, the dataset's validity as a gold standard is undermined. Additionally, a simple audit of the released repository can verify whether manual masks exist for all 65 recordings or only for the 38 stated in the methods section.
Extended reading notes
Core claim
On its own terms, GIRAFE establishes a new public benchmark for glottal-gap segmentation: 65 high-speed color videoendoscopic recordings acquired at 4,000 fps from 50 patients, spanning healthy voices, diagnosed voice disorders, and unknown health states. The manual gold standard consists of expert delineations of the glottal gap, verified by an otolaryngologist, produced for 38 of the 65 videos and totaling 760 frames; the abstract of the paper describes all 65 recordings as manually annotated. Alongside the manual masks, the repository provides segmentations from two classical image-processing methods (InP and Loh), deep-learning baselines (UNet and SwinUnetV2), facilitative playbacks (GAW, GVG, PVG, and digital kymograms), and predefined training/validation/test splits with JSON metadata, all released openly.
Load-bearing premise
The load-bearing premise is that the single expert's manual delineations of the glottal gap are an accurate gold standard; since no inter-annotator agreement is reported and only one otolaryngologist verification pass is mentioned, any systematic bias in those masks would propagate into every reported segmentation score and into the dataset's value as a benchmark.
Editorial extensions
If this is right
- Researchers can train and compare deep-learning glottal segmentation models on a standardized, openly available benchmark with predefined splits, enabling fair and reproducible evaluation.
- The color recordings and the inclusion of playback outputs allow downstream evaluation of segmentation quality through facilitative playbacks, not only through pixel-wise metrics.
- GIRAFE complements BAGLS by adding color data and playback-based analysis, potentially supporting cross-dataset generalization studies and foundation-model fine-tuning in laryngeal imaging.
- The provided baseline results for InP, Loh, UNet, and SwinUnetV2 give immediate reference points for future segmentation methods on this dataset.
Reading between the lines
- The discrepancy between the abstract's claim that all 65 recordings were manually annotated and the methods section's statement that manual segmentation was performed on 38 recordings should be resolved by users checking the released metadata; if only 38 have manual masks, the effective supervised training set is smaller than the headline number suggests.
- Because the ground truth comes from a single expert with no reported inter-annotator agreement, benchmark scores on GIRAFE should be interpreted as reflecting one expert's delineation style; an independent multi-annotator study would materially strengthen the dataset's reliability.
- The color versus grayscale advantage could be tested directly by comparing segmentation performance on GIRAFE against BAGLS under matched training conditions, a comparison the paper motivates but does not perform.
- The inclusion of facilitative playbacks opens the possibility of treating playback fidelity as a task-specific evaluation metric, which may reward segmentations that preserve temporal and morphological features relevant to clinical assessment.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces GIRAFE, a publicly available dataset of 65 high-speed videoendoscopic (HSV) recordings of the vocal folds from 50 patients, intended to support semantic segmentation of the glottal gap and evaluation of facilitative playbacks (GAW, GVG, PVG). The repository includes raw AVI videos, JSON metadata, manual segmentations (reported in Section 2.3 as 38 of the 65 videos, yielding 760 frames), automatic/semi-automatic segmentations from the authors' InP method and the Loh method, and the derived playbacks. A technical validation trains UNet and SwinUnetV2 on the manual masks and compares them against InP and Loh on four test patients (20 frames each) in Table 3. The paper positions GIRAFE as a complement to the existing BAGLS dataset, adding color recordings and playback evaluation, and provides open data and code.
Significance. If the stated coverage is corrected, GIRAFE is a useful and reproducible contribution to laryngeal imaging: it is, to the authors' knowledge, the first public color HSV dataset with glottal-gap annotations, and it ships open code, trained baseline models, and a documented directory structure. The dataset has already supported several prior studies, which is a positive indicator of practical utility. The main significance is as a benchmark resource rather than a methodological advance; the DL validation is explicitly illustrative. However, the abstract's claim that all 65 recordings were manually annotated is contradicted by the body of the paper, and this discrepancy directly affects what a user receives and how the dataset's value is advertised. With the correction, the resource is of moderate but real value to the community.
major comments (3)
- [Abstract; Section 2.3; Section 2.6] The abstract states that 'All of them were manually annotated by an expert, including the masks corresponding to the semantic segmentation of the glottal gap,' and Section 1 repeats the promise of 'annotations meticulously performed and rigorously reviewed.' This is contradicted by Section 2.3, which says 'Manual segmentation was conducted on 38 out of the 65 videos,' and by Section 2.6, which reports that the Manual_Segmentation folder contains exactly 38 subfolders. The discrepancy is load-bearing because it determines the dataset's actual content: 27 of the 65 recordings lack manual masks and can only be used with the InP (and possibly Loh) segmentations. The abstract and all downstream descriptions must be corrected to state the 38-video coverage, and the dataset documentation should report the aggregate count of the Man=Yes metadata flag so users can verify the coverage without opening each recording.
- [Section 2.8; Table 3] The technical validation uses only four test recordings, each with 20 frames (80 frames total per model), and reports no inter-annotator agreement, confidence intervals, or statistical significance for the DICE/Jaccard metrics in Table 3. The paper explicitly disclaims that the validation is 'not intended to be exhaustive,' but the claim that the dataset provides a 'gold standard' for supervised segmentation is weakened by the absence of any evidence about the reliability of the manual masks. I recommend adding at least a small inter-annotator study (even on a subset) and reporting per-frame variability, or alternatively reframing the manual masks as single-expert annotations and softening the 'gold standard' language throughout the paper.
- [Section 2.3] The sentence 'These segmentations were further validated using two established image processing techniques: InP [22] and Loh [20]' is misleading. Comparing automatic segmentations (InP, Loh) against the manual masks is an evaluation of those automatic methods, not a validation of the manual segmentations. If the intent is to demonstrate that the manual masks are reliable, the appropriate evidence is clinician re-annotation, inter-annotator agreement, or some other independent verification, none of which is reported. Please rephrase this sentence to describe what was actually done, and avoid implying that agreement between automatic methods and manual masks establishes the correctness of the masks.
minor comments (6)
- [Section 2.3] The phrase 'PKL Phyton ® format' contains a typo: 'Phyton' should be 'Python'.
- [Section 2.6] There is a duplicated article in 'analysis of the the vocal folds trajectory'; also, 'glottal axis’s' should be 'glottal axis' or 'glottal-axis' to avoid a possessive apostrophe error.
- [Section 2.8] The heading 'Experimental design, Material and Methods' appears mid-paper without a section number and is not listed in the initial section outline; please either format it consistently or integrate it into the preceding subsection.
- [Figure 5] The database tree structure is visually cluttered, especially the parent folder labels for multiple patients, making it difficult to infer the actual directory nesting; a simplified tree or a listing of folder names with the number of entries per folder would be clearer.
- [CRediT Author Statement] The CRediT statement credits 'D.P.S.' with data collection, but no such author is listed among the manuscript's authors; the contribution should be attributed in the author list or acknowledgments, with the consent of that person if appropriate.
- [Section 2.7] The text says the Training directory contains 760 images 'corresponding to the same number of consecutive frames extracted from the original sequences'; please clarify how the 760 frames are distributed across the 38 manually segmented videos (e.g., 20 consecutive frames per video), since the phrasing is ambiguous.
Circularity Check
No circular derivation found; the paper's 65-vs-38 manual annotation discrepancy is a data-description error, not a circular step.
full rationale
GIRAFE is a dataset release rather than a derivation of predictions from first principles, so the standard circularity patterns do not apply. The manual glottal-gap masks are presented as ground truth, and the quantitative validation in Table 3 evaluates UNet, SwinUnetV2, InP, and Loh on held-out frames against those masks; this is ordinary benchmark practice, not a case of a fitted parameter being renamed as a prediction. The authors' self-citations, including the InP baseline and the statement that 'Various subsets of the GIRAFE corpus have been extensively used in the past as an experimental basis for several research works [21, 22, 10, 1, 19]', are present but not load-bearing: the dataset's content and availability do not depend on those prior works, and the InP method is explicitly applied as one of several segmentation approaches rather than used to define the manual ground truth. The one substantive defect is an internal inconsistency: the abstract claims 'All of them were manually annotated by an expert', while Section 2.3 states 'Manual segmentation was conducted on 38 out of the 65 videos' and Section 2.6 confirms that the Manual_Segmentation folder contains 38 subfolders. This is a factual and data-description accuracy issue that materially affects what users receive, but it is not a circularity: no claimed result reduces to its own input by construction. The paper should correct the abstract and clearly report the count of Man=Yes recordings, but this does not raise the circularity score.
Assumptions & free parameters
assumptions (3)
- domain assumption Expert manual annotations of the glottal gap are a valid gold standard for training and evaluation.
- domain assumption One complete glottal cycle of about 20 frames per patient captures enough variability to train and evaluate segmentation models.
- domain assumption The recorded metadata (age, sex, disorder status) are accurate as provided by the clinical records.
Cite this review
Pith. "Pith review of GIRAFE: Glottal Imaging Dataset for Advanced Segmentation, Analysis, and Facilitative Playbacks Evaluation." pith.science (2026). https://pith.science/paper/HPHBPLNF
@misc{pith2026241215054,
author = {Pith},
title = {Pith review of: GIRAFE: Glottal Imaging Dataset for Advanced Segmentation, Analysis, and Facilitative Playbacks Evaluation},
year = {2026},
howpublished = {\url{https://pith.science/paper/HPHBPLNF}},
note = {Machine review of arXiv:2412.15054}
}
read the original abstract
The advances in the development of Facilitative Playbacks extracted from High-Speed videoendoscopic sequences of the vocal folds are hindered by a notable lack of publicly available datasets annotated with the semantic segmentations corresponding to the area of the glottal gap. This fact also limits the reproducibility and further exploration of existing research in this field. To address this gap, GIRAFE is a data repository designed to facilitate the development of advanced techniques for the semantic segmentation, analysis, and fast evaluation of High-Speed videoendoscopic sequences of the vocal folds. The repository includes 65 high-speed videoendoscopic recordings from a cohort of 50 patients (30 female, 20 male). The dataset comprises 15 recordings from healthy controls, 26 from patients with diagnosed voice disorders, and 24 with an unknown health condition. All of them were manually annotated by an expert, including the masks corresponding to the semantic segmentation of the glottal gap. The repository is also complemented with the automatic segmentation of the glottal area using different state-of-the-art approaches. This data set has already supported several studies, which demonstrates its usefulness for the development of new glottal gap segmentation algorithms from High-Speed-Videoendoscopic sequences to improve or create new Facilitative Playbacks. Despite these advances and others in the field, the broader challenge of performing an accurate and completely automatic semantic segmentation method of the glottal area remains open.
Figures
Figures from the paper (5 more)
Reference graph
Works this paper leans on
-
[22]
G. Andrade-Miranda, J. I. Godino-Llorente, Glottal gap tracking by a continuous background modeling using inpainting, Medical & Biological Engineering & Computing 55 (12) (2017) 2123–2141. URL https://doi.org/10.1007/s11517-017-1652-8
-
[20]
J. Lohscheller, H. Toy, F. Rosanowski, U. Eysholdt, M. Dollinger, Clinically evaluated procedure for the recon- struction of vocal fold vibrations from endoscopic digital high-speed videos, Medical Image Analysis 11 (4) (2007) 400–413
work page 2007
-
[1]
G. Andrade-Miranda, Y . Stylianou, D. D. Deliyski, J. I. Godino-Llorente, N. Henrich Bernardoni, Laryngeal image processing of vocal folds motion, Applied Sciences 10 (5) (2020). URL https://www.mdpi.com/2076-3417/10/5/1556
work page 2020
-
[2]
J. I. Godino-Llorente, N. Sáenz-Lechón, V . Osma-Ruiz, S. Aguilera-Navarro, P. Gómez-Vilda, An integrated tool for the diagnosis of voice disorders, Medical Engineering & Physics 28 (3) (2006) 276–289
work page 2006
-
[3]
M. Leppävuori, G. Andrade-Miranda, N. H. Bernardoni, A.-M. Laukkanen, A. Geneid, Characterizing vocal- fold dynamics in singing vocal modes from complete vocal technique using high-speed laryngeal imaging and electroglottographic analysis, in: 13th Pan-European V oice Conference, 2019
work page 2019
-
[4]
C. P.-L. Marion Beaud, Benoît Amy de la Bretèque, N. H. Bernardoni, Clinical characteristics of singers attending a phoniatric outpatient clinic, Logopedics Phoniatrics V ocology 47 (3) (2022) 209–218, pMID: 34110262
work page 2022
-
[5]
N. Henrich Bernardoni, A. Di Corcia, F. Fussi, Laryngeal mechanisms in modern singing, in: COMET Virtual annual conference, Tel-Aviv, Israel, 2022
work page 2022
-
[6]
M. R. Mehl, S. Vazire, N. Ramírez-Esparza, R. B. Slatcher, J. W. Pennebaker, Are women really more talkative than men?, Science 317 (5834) (2007) 82
work page 2007
Show all 70 references
-
[7]
I. R. Titze, Principles of voice production, Prentice-Hall Inc, 1993
1993
-
[8]
A. E. Aronson, D. Bless, Clinical voice disorders, Thieme Medical Publishers, New York, USA, 2009
2009
-
[9]
Laver, The Phonetic Description of V oice Quality, V ol
J. Laver, The Phonetic Description of V oice Quality, V ol. 31 of Cambridge Studies in Linguistics, Cambridge University Press, 2009
2009
-
[10]
Andrade-Miranda, Analyzing of the vocal fold dynamics using laryngeal videos, Ph.D
G. Andrade-Miranda, Analyzing of the vocal fold dynamics using laryngeal videos, Ph.D. thesis, Universidad Politécnica de Madrid (2017)
2017
-
[11]
J. G. Švec, F. Šram, H. K. Schutte, Videokymography in V oice Disorders: What to Look For?, Annals of Otology, Rhinology and Laryngology 116 (3) (2007) 172–180
2007
-
[12]
H. S. Shaw, D. D. Deliyski, Mucosal wave: A normophonic study across visualization techniques, Journal of V oice 22 (1) (2008) 23 – 33
2008
-
[13]
V oigt, M
D. V oigt, M. Döllinger, T. Braunschweig, A. Yang, U. Eysholdt, J. Lohscheller, Classification of functional voice disorders based on phonovibrograms, Artificial Intelligence in Medicine 49 (1) (2010) 51–59
2010
-
[14]
A. P. Pinheiro, D. E. Stewart, C. D. Maciel, J. C. Pereira, S. Oliveira, Analysis of nonlinear dynamics of vocal folds using high-speed video observation and biomechanical modeling, Digital Signal Processing 22 (2) (2012) 304–313
2012
-
[15]
Yamauchi, H
A. Yamauchi, H. Yokonishi, H. Imagawa, K.-I. Sakakibara, T. Nito, N. Tayama, T. Yamasoba, Quantification of vocal fold vibration in various laryngeal disorders using high-speed digital imaging, Journal of V oice 30 (2) (2016) 205–214
2016
-
[16]
Crocker, L
C. Crocker, L. E. Toles, R. A. Morrison, A. C. Shembel, Relationships between vocal fold adduction patterns, vocal acoustic quality, and vocal effort in individuals with and without hyperfunctional voice disorders, Journal of V oice (2024)
2024
-
[17]
Andrade-Miranda, N
G. Andrade-Miranda, N. Henrich, J. I. Godino-Llorente, Optical-flow kymograms and glottovibrograms: A new way to present high-speed data for laryngeal assessment, in: 9th International workshop, Models and Analysis of V ocal Emissions for Biomedical Applications, MA VEBA, Fire...
2015
-
[18]
Andrade-Miranda, N
G. Andrade-Miranda, N. H. Bernardoni, J. I. Godino-Llorente, A new technique for assessing glottal dynamics in speech and singing by means of optical-flow computation, in: 16th Annual Conference of the International Speech Communication Association (INTERSPEECH), Dresden, Germ...
2015
-
[19]
Andrade-Miranda, N
G. Andrade-Miranda, N. Henrich Bernardoni, J. I. Godino-Llorente, Synthesizing the motion of the vocal folds using optical flow based techniques, Biomedical Signal Processing and Control 34 (2017) 25–35
2017
-
[21]
Andrade-Miranda, J
G. Andrade-Miranda, J. I. Godino-Llorente, L. Moro-Velázquez, J. A. Gómez-García, An automatic method to detect and track the glottal gap from high speed videoendoscopic images, BioMedical Engineering OnLine 14 (1) (2015) 100
2015
-
[23]
S. Z. Karakozoglou, N. Henrich, C. D’Alessandro, Y . Stylianou, Automatic glottal segmentation using local- based active contours and application to glottovibrography, Speech Communication 54 (5) (2012) 641–654
2012
-
[24]
Lohscheller, U
J. Lohscheller, U. Eysholdt, Phonovibrogram visualization of entire vocal fold dynamics, The Laryngoscope 118 (4) (2008) 753–758
2008
-
[25]
J. G. Švec, H. K. Schutte, Videokymography: High-speed line scanning of vocal fold vibration, Journal of V oice 10 (2) (1996) 201 – 205
1996
-
[26]
J. G. Švec, F. Šram, Kymographic imaging of the vocal folds oscillations, in: 7th International Conference on Spoken Language Processing, V ol. 2, 2002, pp. 957–960
2002
-
[27]
Wurzbacher, R
T. Wurzbacher, R. Schwarz, M. Döllinger, U. Hoppe, U. Eysholdt, J. Lohscheller, Model-based classification of nonstationary vocal fold vibrations. model-based classification of nonstationary vocal fold vibrations., The Journal of the Acoustical Society of America 120 (2) (2006...
2006
-
[28]
Lohscheller, U
J. Lohscheller, U. Eysholdt, Phonovibrography: Mapping high-speed movies of vocal fold vibrations into 2- D diagrams for visualizing and analyzing the underlying laryngeal dynamics, IEEE Transactions on Medical Imaging 27 (3) (2008) 300–309
2008
-
[29]
D. D. Mehta, D. D. Deliyski, T. F. Quatieri, R. E. Hillman, Automated measurement of vocal fold vibratory asymmetry from high-speed videoendoscopy recordings, Journal of Speech, Language, and Hearing Research 54 (1) (2013) 47–54
2013
-
[30]
Lohscheller, J
J. Lohscheller, J. G. Švec, M. Döllinger, V ocal fold vibration amplitude, open quotient, speed quotient and their variability along glottal length: kymographic data from normal subjects, Logopedics Phoniatrics V ocology 38 (4) (2013) 182–192
2013
-
[31]
C. T. Herbst, J. Lohscheller, J. G. Švec, N. Henrich, G. Weissengruber, W. T. Fitch, Glottal opening and closing events investigated by electroglottography and super-high-speed video recordings., The Journal of Experimental Biology 217 (6) (2014) 955–963
2014
-
[32]
Y . Yan, E. Damrose, D. Bless, Functional analysis of voice using simultaneous high-speed imaging and acoustic recordings., Journal of V oice 21 (5) (2007) 604–616
2007
-
[33]
Ahmad, Y
k. Ahmad, Y . Yan, D. Bless, V ocal fold vibratory characteristics in normal female speakers from high-speed digital imaging, Journal of V oice 26 (2) (2012) 239–253
2012
-
[34]
R. R. Patel, K. Forrest, D. Hedges, Relationship between acoustic voice onset and offset and selected instances of oscillatory onset and offset in young healthy men and women, Journal of V oice 31 (3) (2017) 389.e9 – 389.e17
2017
-
[35]
Schlegel, M
P. Schlegel, M. Stingl, M. Kunduk, S. Kniesburges, C. Bohr, M. Döllinger, Dependencies and ill-designed pa- rameters within high-speed videoendoscopy and acoustic signal analysis, Journal of V oice (2018)
2018
-
[36]
C. R. Krausert, A. E. Olszewski, L. N. Taylor, J. S. McMurray, S. H. Dailey, J. J. Jiang, Mucosal wave measure- ment and visualization techniques, Journal of V oice 25 (4) (2011) 395 – 405
2011
-
[37]
Kaneko, O
M. Kaneko, O. Shiromoto, M. Fujiu-Kurachi, Y . Kishimoto, I. Tateya, S. Hirano, Optimal duration for voice rest after vocal fold surgery: Randomized controlled clinical study, Journal of V oice 31 (1) (2017) 97 – 103
2017
-
[38]
L. Li, Y . Zhang, A. L. Maytag, J. J. Jiang, Quantitative study for the surface dehydration of vocal folds based on high-speed imaging, Journal of V oice 29 (4) (2015) 403 – 409
2015
-
[39]
C. T. Herbst, J. G. Švec, J. Lohscheller, R. Frey, M. Gumpenberger, A. S. Stoeger, W. T. Fitch, Complex vibratory patterns in an elephant larynx, Journal of Experimental Biology 216 (21) (2013) 4054–4064. doi:10.1242/ jeb.091009. 16 GIRAFE: Glottal Imaging Repository for Advan...
2013
-
[40]
C. P. H. Elemans, J. H. Rasmussen, C. T. Herbst, D. N. Düring, S. A. Zollinger, H. Brumm, K. H. Srivastava, N. Svane, M. Ding, O. N. Larsen, S. J. Sober, J. G. Svec, Universal mechanisms of sound production and control in birds and mammals, Nature communications 6 (2015) 8978
2015
-
[41]
C. T. Herbst, Biophysics of V ocal Production in Mammals, Springer International Publishing, Cham, 2016, pp. 159–189
2016
-
[42]
V oigt, M
D. V oigt, M. Döllinger, U. Eysholdt, A. Yang, E. Gürlen, J. Lohscheller, Objective detection and quantification of mucosal wave propagation, The Journal of the Acoustical Society of America 128 (5) (2010) EL347–EL353
2010
-
[43]
Schlegel, S
P. Schlegel, S. Kniesburges, S. Dürr, et al., Machine learning based identification of relevant parameters for functional voice disorders derived from endoscopic high-speed recordings, Scientific Reports 10 (2020) 10517. URL https://doi.org/10.1038/s41598-020-66405-y
2020 doi
-
[44]
Unger, J
J. Unger, J. Lohscheller, M. Reiter, K. Eder, C. S. Betz, M. Schuster, A Noninvasive Procedure for Early-Stage Discrimination of Malignant and Precancerous V ocal Fold Lesions Based on Laryngeal Dynamics Analysis, Cancer Research 75 (1) (2015) 31–39
2015
-
[45]
A. Kist, S. Dürr, A. Schützenberger, et al., Openhsv: an open platform for laryngeal high-speed videoendoscopy, Scientific Reports 11 (2021) 13760. URL https://doi.org/10.1038/s41598-021-93149-0
2021 doi
-
[46]
A. M. Kist, K. Breininger, M. Dörrich, et al., A single latent channel is sufficient for biomedical glottis segmen- tation, Scientific Reports 12 (2022) 14292. URL https://doi.org/10.1038/s41598-022-17764-1
2022 doi
-
[47]
Chen, P.-Y
I.-M. Chen, P.-Y . Yeh, Y .-C. Hsieh, T.-C. Chang, S. Shih, W.-F. Shen, C.-L. Chin, 3d vosnet: Segmentation of endoscopic images of the larynx with subsequent generation of indicators, Heliyon 9 (3) (2023) e14242. URL https://www.sciencedirect.com/science/article/pii/S2405844023014494
2023
-
[48]
Pedersen, C
M. Pedersen, C. F. Larsen, B. Madsen, et al., Localization and quantification of glottal gaps on deep learning segmentation of vocal folds, Scientific Reports 13 (2023) 878. URL https://doi.org/10.1038/s41598-023-27980-y
2023 doi
-
[49]
Conze, G
P.-H. Conze, G. Andrade-Miranda, V . K. Singh, V . Jaouen, D. Visvikis, Current and emerging trends in medical image segmentation with deep learning, IEEE Transactions on Radiation and Plasma Medical Sciences 7 (6) (2023) 545–569
2023
-
[50]
Andrade-Miranda, V
G. Andrade-Miranda, V . Jaouen, O. Tankyevych, C. Cheze Le Rest, D. Visvikis, P.-H. Conze, Multi-modal medical transformers: A meta-analysis for medical image segmentation in oncology, Computerized Medical Imaging and Graphics 110 (2023) 102308. URL https://www.sciencedirect.c...
2023
-
[51]
B. Azad, R. Azad, S. Eskandari, A. Bozorgpour, A. Kazerouni, I. Rekik, D. Merhof, Foundational models in medical imaging: A comprehensive survey and future vision, arXiv preprint arXiv:2310.18689 (2023)
2023 arXiv
-
[52]
H. Wang, P. K. A. Vasu, F. Faghri, R. Vemulapalli, M. Farajtabar, S. Mehta, M. Rastegari, O. Tuzel, H. Pouransari, Sam-clip: Merging vision foundation models towards semantic and spatial understanding, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern R...
2024
-
[53]
Alzubaidi, J
L. Alzubaidi, J. Bai, A. Al-Sabaawi, et al., A survey on deep learning tools dealing with data scarcity: definitions, challenges, solutions, tips, and applications, Journal of Big Data 10 (2023) 46. URL https://doi.org/10.1186/s40537-023-00727-2
2023 doi
-
[54]
Andrade-Miranda, P
G. Andrade-Miranda, P. Soto Vega, K. Taguelmimt, H.-P. Dang, D. Visvikis, J. Bert, Exploring transformer reliability in clinically significant prostate cancer segmentation: A comprehensive in-depth investigation, Com- puterized Medical Imaging and Graphics - (2024) –
2024
-
[55]
A. E. Kavur, N. S. Gezer, M. Barı¸ s, S. Aslan, P.-H. Conze, V . Groza, D. D. Pham, S. Chatterjee, P. Ernst, S. Özkan, et al., Chaos challenge-combined (ct-mr) healthy abdominal organ segmentation, Medical Image Analysis 69 (2021) 101950
2021
-
[56]
Y . Ji, H. Bai, J. Yang, C. Ge, Y . Zhu, R. Zhang, Z. Li, L. Zhang, W. Ma, X. Wan, et al., Amos: A large-scale abdominal multi-organ benchmark for versatile medical image segmentation, arXiv preprint arXiv:2206.08023 (2022)
2022 arXiv
-
[57]
J. Ma, Y . Zhang, S. Gu, C. Ge, S. Ma, A. Young, C. Zhu, K. Meng, X. Yang, Z. Huang, et al., Unleashing the strengths of unlabeled data in pan-cancer abdominal organ quantification: the flare22 challenge, arXiv preprint arXiv:2308.05862 (2023). 17 GIRAFE: Glottal Imaging Repos...
2023 arXiv
-
[58]
Koitka, G
S. Koitka, G. Baldini, L. Kroll, et al., Saros: A dataset for whole-body region and organ segmentation in ct imaging, Scientific Data 11 (2024) 483. URL https://doi.org/10.1038/s41597-024-03337-6
2024 doi
-
[59]
J. Liu, Y . Zhang, J.-N. Chen, J. Xiao, Y . Lu, B. A. Landman, Y . Yuan, A. Yuille, Y . Tang, Z. Zhou, Clip-driven universal model for organ segmentation and tumor detection, arXiv preprint arXiv:2301.00785 (2023)
2023 arXiv
-
[60]
Gómez, A
P. Gómez, A. M. Kist, P. Schlegel, et al., Bagls, a multihospital benchmark for automatic glottis segmentation, Scientific Data 7 (2020) 186. URL https://doi.org/10.1038/s41597-020-0526-3
2020 doi
-
[61]
M. A. Mazurowski, H. Dong, H. Gu, J. Yang, N. Konz, Y . Zhang, Segment anything model for medical image analysis: An experimental study, Medical Image Analysis 89 (2023) 102918
2023
-
[62]
Siblini, G
L. Siblini, G. Andrade-Miranda, K. Taguelmimt, D. Visvkis, J. Bert, Optimal prompting in sam for few-shot and weakly supervised medical image segmentation, in: MICCAI 2024 2nd International Workshop on Foundation Models for General Medical AI. Accepted on July 15, 2024
2024
-
[63]
Ronneberger, P.Fischer, T
O. Ronneberger, P.Fischer, T. Brox, U-net: Convolutional networks for biomedical image segmentation, in: Medical Image Computing and Computer-Assisted Intervention (MICCAI), V ol. 9351 of LNCS, Springer, 2015, pp. 234–241, (available on arXiv:1505.04597 [cs.CV]). URL http://lm...
2015 arXiv
-
[64]
Z. Liu, H. Hu, Y . Lin, Z. Yao, Z. Xie, Y . Wei, J. Ning, Y . Cao, Z. Zhang, L. Dong, et al., Swin Transformer v2: Scaling up capacity and resolution, in: IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 12009–12019
2022
-
[65]
Deliyski, P
D. Deliyski, P. P. Petrushev, H. Bonilha, T. T. Gerlach, B. Martin-Harris, R. E. Hillman, Clinical implementation of laryngeal high-speed videoendoscopy: Challenges and evolution, Folia Phoniatrica et Logopaedica 60 (1) (2008) 33–44
2008
-
[66]
Kendall, R
K. Kendall, R. Leonard, Laryngeal Evaluation: Indirect Laryngoscopy to High-speed Digital Imaging, Thieme, 2010, Ch. 14. Normal Glottic Configuration, pp. 245–270
2010
-
[67]
Braunschweig, J
T. Braunschweig, J. Flaschka, P. Schelhorn-Neise, M. Döllinger, High-speed video analysis of the phonation onset, with an application to the diagnosis of functional dysphonias, Medical Engineering and Physics 30 (1) (2008) 59–66
2008
-
[68]
P. Birkholz, Glottalimageexplorer - an open source tool for glottis segmentation in endoscopic high-speed videos of the vocal folds, in: Studientexte zur Sprachkommunikation: Elektronische Sprachsignalverarbeitung 2016, Dresden, Germany, 2016
2016
-
[69]
M. J. Cardoso, W. Li, R. Brown, N. Ma, E. Kerfoot, Y . Wang, B. Murrey, A. Myronenko, C. Zhao, D. Yang, V . Nath, Y . He, Z. Xu, A. Hatamizadeh, A. Myronenko, W. Zhu, Y . Liu, M. Zheng, Y . Tang, I. Yang, M. Zephyr, B. Hashemian, S. Alle, M. Z. Darestani, C. Budd, M. Modat, T....
2022 arXiv
-
[70]
Andrade-Miranda, D
G. Andrade-Miranda, D. Poletti-Serafini, K. Chatzipapas, J. D. A. Londoño, , J. I. Godino-Llorente, GIRAFE: Glottal Imaging Repository for Advanced Segmentation, Analysis, and Facilitative Playbacks Evaluation (1.0.0) (Sep. 2024). doi:https://doi.org/10.5281/zenodo.13773163. 18
2024 doi
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.