REVIEW 3 major objections 5 minor 84 references
A Survey on Spoken Italian Datasets and Corpora
T0 review · 3 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read This survey catalogs 66 spoken Italian datasets, organizes them by speech type, source, and demographic and linguistic features, and publishes the complete inventory openly, while identifying persistent gaps in dialect, child, and…
desk verdict Useful inventory for Italian speech researchers, but the 'comprehensive' claim can't be audited without the full list and search protocol in the paper itself. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central artifact is the curated inventory of 66 datasets, organized along three axes: speech type, source and context, and demographic and linguistic features. The inventory, publicly available on GitHub and archived on Zenodo, is what carries the survey's claims, because the gap analysis, the application tables, and the recommendations all derive from the way these datasets are sorted and compared. The paper also uses a set of named tool references, such as Praat, ELAN, and WebMAUS, to illustrate annotation practice, but those serve as examples of methodology rather than as carriers of the central claim.
What would settle it
Check the public Zenodo archive of the inventory and look for an actively maintained spoken Italian dataset, documented as of November 2024, that is absent from the 66; finding a significant missing dataset would falsify the comprehensiveness claim. A second check is to verify a random sample of the 66 entries against their cited source pages, and if many sizes, speech types, or availability statuses are wrong, the categorization would collapse.
Extended reading notes
Core claim
The paper's central claim is that it provides a comprehensive examination of 66 spoken Italian datasets, characterized by speech type, source and context, and demographic and linguistic features, with the full inventory publicly available via GitHub and archived on Zenodo. It finds that the existing resources concentrate on read and conversational standard Italian, leaving dialects, minority languages, children's speech, field recordings, and specialized domains comparatively underrepresented. It also identifies technical and ethical challenges, such as inconsistent annotation, large-file overhead, privacy concerns, and access restrictions, and proposes standardized annotation, open-access models, and collaborative collection as the main remedies.
Load-bearing premise
The survey assumes that the 66 datasets it assembled, with sizes and features taken from creators' own documentation, accurately represent the full landscape of usable spoken Italian resources as of November 2024, and that this documentation is reliable.
Editorial extensions
If this is right
- Researchers can use the public inventory to locate an Italian speech dataset by speech type, source, or demographic need, instead of rediscovering resources through scattered catalogs.
- The gap analysis gives concrete targets for new data collection: dialects and minority languages, child speech, outdoor field recordings, and specialized domains such as medical or non-native speech.
- Adopting the paper's recommendations for standardized annotation formats, open-access licensing, and richer metadata would make Italian speech datasets easier to combine and compare.
- The survey's reliance on creator documentation means its characterizations are a starting point for selection, not an independent quality certification of the datasets themselves.
Reading between the lines
- The three-axis categorization is simple enough to serve as a lightweight template for comparable surveys of other under-resourced languages, not just Italian.
- Because the inventory is frozen as of November 2024 and hosted on a public repository, it is positioned to become a living resource; community updates would be a natural next step that the paper only gestures at.
- The paper highlights a few dozen datasets in its tables and discussion while the full inventory contains 66 entries, so a reader who wants the complete map must consult the GitHub or Zenodo archive rather than rely only on the body text.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents a survey of 66 spoken Italian datasets and corpora, organizing them by speech type, source/context, and demographic/linguistic features, and discussing collection and annotation methodologies, applications, challenges, and future directions. The authors state that the complete inventory is available on GitHub and archived on Zenodo with DOI 10.5281/zenodo.14246196, while the article itself presents selected examples in three tables. The survey's stated aim is to provide a comprehensive overview of usable spoken Italian resources as of November 2024, along with a gap analysis and recommendations.
Significance. If the inventory is accurate, this survey fills a genuine gap: there is no comparable recent reference cataloging spoken Italian resources across academic, commercial, and crowdsourced sources. The authors' decision to publish the full inventory as an open, versioned artifact on GitHub/Zenodo is a concrete contribution that can be updated and reused, and the categorization by speech type, source, and demographic/linguistic features is useful for navigating the resource landscape. The paper is also transparent in stating that dataset quality and availability were not independently verified. However, the central 'comprehensive' claim and the resulting gap analysis in Sections V and VI depend on the completeness and correct attribution of the 66-item inventory, and this dependence is currently the paper's weakest point because the selection protocol is unspecified and the full list is not in the manuscript.
major comments (3)
- [I.B, V, VII] The claim of comprehensiveness over 66 datasets is not auditable from the manuscript. Section I.B does not state the search queries, databases queried, inclusion/exclusion criteria, or a working definition of 'actively maintained,' and Section VII directs readers to an external GitHub/Zenodo archive for the full list. Within the paper, Table 1 shows 23 rows, Table 2 shows 23 datasets, and Table 3 shows 9 datasets, so the reader cannot check membership, duplicates, or overlap handling from the article itself. Because Section V's scarcity/accessibility analysis and Section VI's recommendations are derived from what is absent from this list, an incomplete or misattributed inventory would change the conclusions, not just the numbers. The authors should provide the full 66-item inventory as a supplementary appendix or, at minimum, an appendix table with the same fields as Table 1, and they should document the search and inclusion protocol.
- [Table 1, II.A.2] Several table entries appear to report multilingual totals rather than Italian-specific figures, which is inconsistent with the survey's stated scope. The VoxPopuli row lists '400,000 hours' and labels the source as 'Standard European languages,' but the paper's subject is spoken Italian datasets; no Italian-specific hour count is given in the table or in Section II.A.2. Similarly, the C-ORAL-ROM row reports '1,200,000 words' for 'Spontaneous speech in Romance languages,' which is a multilingual total. If Italian-specific quantities are unavailable or not reported, the table should state 'not reported' rather than a multilingual total that can be misread as an Italian dataset size.
- [V.A, I.B] The gap analysis in Section V.A draws conclusions about restricted access and commercial availability (e.g., 'commercial datasets such as Aurora Project [28] and Defined.ai [3], [7], [25] require significant financial investment') from data that the authors explicitly did not independently verify in Section I.B. While the authors' caveat is honest, it does not resolve the inconsistency between the caveat and the strength of the claims in Section V.A. The manuscript should either verify accessibility at the time of writing or mark each entry's availability as reported rather than confirmed, and it should define 'actively maintained' as used in Section I.B.
minor comments (5)
- [I.C] Section I.C refers to 'Section 2,' 'Section 3,' etc., while the actual section headings use Roman numerals (II, III, etc.); the numbering should be made consistent.
- [Table 3] The MuST-C row in Table 3 lists 'Creative Commons (temporarily suspended)' as the availability, which is an odd status for a table entitled 'Summary of Publicly Available Italian Speech Datasets'; the table should clarify what 'temporarily suspended' means and whether the dataset is currently downloadable.
- [II.C.1] In Section II.C.1, C-ORAL-ROM is described as 'though multilingual, includes a substantial component of Italian speech,' but neither the text nor Table 1 quantifies the Italian component; adding the Italian-specific size would strengthen the categorization.
- [IV.A.1] Reference [56], cited for EMOVO, is incomplete and lacks a publisher, venue, or URL; it should be completed for readers wishing to locate the resource.
- [III.C] The sentence 'according to our current knowledge, none of the documented datasets made use of this tool in validation processes' reads as an aside and would be clearer as a footnote or a statement of the limitation of the survey's information, since the authors already disclaim independent verification.
Circularity Check
No significant circularity: the survey derives no fitted quantities, makes no predictions from its own inputs, and rests on external dataset documentation rather than on a self-citation chain.
full rationale
This paper is a cataloging survey, not a derivation. It contains no equations, no fitted parameters, and no prediction that is later compared with the data that defined it. The central claim, that the paper provides a comprehensive overview of 66 spoken Italian datasets, is supported by citations to external dataset pages, repositories, and published corpus descriptions, not by the authors' own prior results. The only self-referential element is the statement in Section I.B and the Conclusion that the complete inventory is hosted on GitHub and archived on Zenodo (DOI: 10.5281/zenodo.14246196). That is an availability claim, not a circular step: the paper does not derive its categorization or gap analysis from the contents of that archive, nor does it treat the archive as evidence for its completeness. The manuscript explicitly discloses in Section I.B that 'dataset quality and availability were not independently verified' and that the survey 'synthesizes information provided by creators or documentation,' which correctly identifies that the inventory is compiled from external sources. Concerns about whether the 66-item inventory is auditable, whether inclusion criteria are stated, and whether some table entries attribute multilingual totals to Italian resources are legitimate reproducibility and verifiability risks, but they are not circularity: the conclusions do not reduce by construction to a fitted input or to an unverified self-citation. No load-bearing self-citations, uniqueness theorems, or ansatz-smuggling citations appear in the text. Accordingly, no circular step can be quoted or exhibited, and the honest finding is that the paper exhibits no significant circularity.
Assumptions & free parameters
assumptions (2)
- domain assumption Dataset documentation supplied by creators is accurate
- domain assumption The set of 66 datasets identified by the authors is comprehensive
Cite this review
Pith. "Pith review of A Survey on Spoken Italian Datasets and Corpora." pith.science (2026). https://pith.science/paper/2OHSQ2YD
@misc{pith2026250106557,
author = {Pith},
title = {Pith review of: A Survey on Spoken Italian Datasets and Corpora},
year = {2026},
howpublished = {\url{https://pith.science/paper/2OHSQ2YD}},
note = {Machine review of arXiv:2501.06557}
}
read the original abstract
Spoken language datasets are vital for advancing linguistic research, Natural Language Processing, and speech technology. However, resources dedicated to Italian, a linguistically rich and diverse Romance language, remain underexplored compared to major languages like English or Mandarin. This survey provides a comprehensive analysis of 66 spoken Italian datasets, highlighting their characteristics, methodologies, and applications. The datasets are categorized by speech type, source and context, and demographic and linguistic features, with a focus on their utility in fields such as Automatic Speech Recognition, emotion detection, and education. Challenges related to dataset scarcity, representativeness, and accessibility are discussed alongside recommendations for enhancing dataset creation and utilization. The full dataset inventory is publicly accessible via GitHub and archived on Zenodo, serving as a valuable resource for researchers and developers. By addressing current gaps and proposing future directions, this work aims to support the advancement of Italian speech technologies and linguistic research.
Reference graph
Works this paper leans on
-
[28]
“AURORA Project database - Subset of SpeechDat-Car - Italian database - Evaluation Package – ELRA Catalogue.” [Online]. Available: https://catalogue.elra.info/en-us/repository/browse/ ELRA-AURORA-CD0003_05/
-
[3]
Italian Spontaneous Dialogue Dataset | Defined.ai
“Italian Spontaneous Dialogue Dataset | Defined.ai.” [Online]. Available: https://www.defined.ai/datasets/italian-spontaneous-dialogue
-
[7]
Italian Scripted Monologue Dataset | Defined.ai
“Italian Scripted Monologue Dataset | Defined.ai.” [Online]. Available: https://www.defined.ai/datasets/italian-scripted-monologue
-
[25]
Italian Lifestyle Podcast Dataset | Defined.ai
“Italian Lifestyle Podcast Dataset | Defined.ai.” [Online]. Available: https://www.defined.ai/datasets/italian-lifestyle-podcast
-
[1]
Corpus KIParla – L’italiano parlato e chi parla italiano
“Corpus KIParla – L’italiano parlato e chi parla italiano.” [Online]. Available: https://kiparla.it/
-
[2]
KIParla Corpus: A New Resource for Spoken Italian,
C. Mauri, S. Ballarè, E. Goria, M. Cerruti, and F. Suriano, “KIParla Corpus: A New Resource for Spoken Italian,” Proceedings of the Sixth Italian Conference on Computational Linguistics, vol. 2481
-
[4]
A. Vietti and D. Mereu, “DIA-Dialogic ItAlian corpus,” 2021, publisher: The Language Archive - Max Planck Institute for Psycholinguistics. [Online]. Available: https://bia.unibz.it/esploro/outputs/ dataset/DIA-Dialogic-ItAlian-corpus/991006608798401241
-
[5]
Robust Speech Recognition via Large-Scale Weak Supervision,
A. Radford, J. W. Kim, T. Xu, G. Brockman, C. Mcleavey, and I. Sutskever, “Robust Speech Recognition via Large-Scale Weak Supervision,” in Proceedings of the 40th International Conference on Machine Learning. 14 VOLUME XX, 202X PMLR, Jul. 2023, pp. 28 492–28 518, iSSN: 2640-3498. [Online]. Available: https://proceedings.mlr.press/v202/radford23a.html
2023
Show all 84 references
-
[6]
WavLM: Large-Scale Self-Supervised Pre- Training for Full Stack Speech Processing,
S. Chen, C. Wang, Z. Chen, Y . Wu, S. Liu, Z. Chen, J. Li, N. Kanda, T. Yoshioka, X. Xiao, J. Wu, L. Zhou, S. Ren, Y . Qian, Y . Qian, J. Wu, M. Zeng, X. Yu, and F. Wei, “WavLM: Large-Scale Self-Supervised Pre- Training for Full Stack Speech Processing,” Jun. 2022, arXiv:2110....
2022
-
[8]
voxpopuli,
“voxpopuli,” Nov. 2024, original-date: 2021-01-08T18:47:36Z. [Online]. Available: https://github.com/facebookresearch/voxpopuli
2024
-
[9]
V oxPopuli: A Large- Scale Multilingual Speech Corpus for Representation Learning, Semi- Supervised Learning and Interpretation,
C. Wang, M. Rivière, A. Lee, A. Wu, C. Talnikar, D. Haziza, M. Williamson, J. Pino, and E. Dupoux, “V oxPopuli: A Large- Scale Multilingual Speech Corpus for Representation Learning, Semi- Supervised Learning and Interpretation,” Jul. 2021, arXiv:2101.00390. [Online]. Availabl...
2021 arXiv
-
[10]
Spontaneous Speech Characterization and Detection in Large Audio Database,
R. Dufour, V . Jousse, Y . Estève, F. Béchet, and G. Linarès, “Spontaneous Speech Characterization and Detection in Large Audio Database,” in 13-th International Conference on Speech and Computer (SPECOM 2009), St Petersburg, Russia, 2009. [Online]. Available: https://hal.scie...
2009
-
[11]
A comparison of ASR and human errors for transcription of non- native spontaneous speech,
M. Mulholland, M. Lopez, K. Evanini, A. Loukina, and Y . Qian, “A comparison of ASR and human errors for transcription of non- native spontaneous speech,” in 2016 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Mar. 2016, pp. 5855–5859, iSSN:...
2016
-
[12]
The LABLITA Speech Resources,
E. CRESTI, L. GREGORI, M. MONEGLIA, C. NICOLÁS, and A. PANUNZI, “The LABLITA Speech Resources,” Corpora e Studi Linguistici, no. 6, pp. 85–108, 2022. [Online]. Available: https: //doi.org/10.17469/O2106SLI000005
2022 doi
-
[13]
Lablita
“Lablita.” [Online]. Available: http://corpus.lablita.it/?locale=en
-
[14]
Multi-media edition; tools of analysis; standard linguistic mea- surements for validation in HLT – ELRA Catalogue.” [Online]
“C-ORAL-ROM - Integrated reference corpora for spoken romance languages. Multi-media edition; tools of analysis; standard linguistic mea- surements for validation in HLT – ELRA Catalogue.” [Online]. Available: https://catalogue.elra.info/en-us/repository/browse/ELRA-S0172/
-
[15]
Hey ASR System! Why Aren’t You More Inclusive? Automatic Speech Recognition Systems’ Bias and Proposed Bias Mitigation Techniques. A Literature Review,
M. K. Ngueajio and G. Washington, “Hey ASR System! Why Aren’t You More Inclusive? Automatic Speech Recognition Systems’ Bias and Proposed Bias Mitigation Techniques. A Literature Review,” Nov. 2022, arXiv:2211.09511. [Online]. Available: http://arxiv.org/abs/2211.09511
2022 arXiv
-
[16]
Speak, Read and Prompt: High-Fidelity Text-to-Speech with Minimal Supervision,
E. Kharitonov, D. Vincent, Z. Borsos, R. Marinier, S. Girgin, O. Pietquin, M. Sharifi, M. Tagliasacchi, and N. Zeghidour, “Speak, Read and Prompt: High-Fidelity Text-to-Speech with Minimal Supervision,” Feb. 2023, arXiv:2302.03540. [Online]. Available: http://arxiv.org/abs/2302.03540
2023 arXiv
-
[17]
MLS: A Large-Scale Multilingual Dataset for Speech Research,
V . Pratap, Q. Xu, A. Sriram, G. Synnaeve, and R. Collobert, “MLS: A Large-Scale Multilingual Dataset for Speech Research,” in Interspeech 2020, Oct. 2020, pp. 2757–2761, arXiv:2012.03411 [cs, eess]. [Online]. Available: http://arxiv.org/abs/2012.03411
2020 arXiv
-
[18]
MLS - Multilingual Librispeech
“MLS - Multilingual Librispeech.” [Online]. Available: https://www. openslr.org/94/
-
[19]
Common V oice Mozilla
“Common V oice Mozilla.” [Online]. Available: https://commonvoice. mozilla.org/
-
[20]
Combining Deep Learning with Domain Adaptation and Filtering Techniques for Speech Recognition in Noisy Environments,
E. De J. Velásquez-Martínez, A. Becerra-Sánchez, J. I. De La Rosa- Vargas, E. González-Ramírez, A. Rodarte-Rodríguez, G. Zepeda-Valles, N. I. Escalante-García, and J. E. Olvera-González, “Combining Deep Learning with Domain Adaptation and Filtering Techniques for Speech Recogn...
2023
-
[21]
Analysis of Deep Clustering as Preprocessing for Automatic Speech Recognition of Sparsely Overlapping Speech,
T. Menne, I. Sklyar, R. Schlüter, and H. Ney, “Analysis of Deep Clustering as Preprocessing for Automatic Speech Recognition of Sparsely Overlapping Speech,” Sep. 2019, arXiv:1905.03500. [Online]. Available: http://arxiv.org/abs/1905.03500
2019 arXiv
-
[22]
Identifying Fake News on Social Networks Based on Natural Language Processing: Trends and Challenges,
N. R. De Oliveira, P. S. Pisa, M. A. Lopez, D. S. V . De Medeiros, and D. M. F. Mattos, “Identifying Fake News on Social Networks Based on Natural Language Processing: Trends and Challenges,” Information, vol. 12, no. 1, p. 38, Jan. 2021. [Online]. Available: https://www.mdpi....
2021
-
[23]
The MGB challenge: Evaluating multi-genre broadcast media recognition,
P. Bell, M. J. F. Gales, T. Hain, J. Kilgour, P. Lanchantin, X. Liu, A. McParland, S. Renals, O. Saz, M. Wester, and P. C. Woodland, “The MGB challenge: Evaluating multi-genre broadcast media recognition,” in 2015 IEEE Workshop on Automatic Speech Recognition and Understanding...
2015
-
[24]
IBNC - An Italian Broadcast News Corpus – ELRA Catalogue
“IBNC - An Italian Broadcast News Corpus – ELRA Catalogue.” [Online]. Available: https://catalogue.elra.info/en-us/repository/browse/ ELRA-S0093/
-
[26]
Exploring AI-Driven Customer Service: Evolution, Archi- tectures, Opportunities, Challenges and Future Directions,
S. M. I. , “Exploring AI-Driven Customer Service: Evolution, Archi- tectures, Opportunities, Challenges and Future Directions,” International Journal For Multidisciplinary Research, vol. 6, no. 3, p. 22283, Jun. 2024. [Online]. Available: https://www.ijfmr.com/research-paper.p...
2024
-
[27]
Italian(Italy) Spontaneous Dialogue Telephony speech dataset - Nexdata
“Italian(Italy) Spontaneous Dialogue Telephony speech dataset - Nexdata.” [Online]. Available: https://m.nexdata.ai/datasets/speechrecog/ 1232?source=Huggingface
-
[29]
MuST-C: a Multilingual Speech Translation Corpus,
M. A. Di Gangi, R. Cattoni, L. Bentivogli, M. Negri, and M. Turchi, “MuST-C: a Multilingual Speech Translation Corpus,” in Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, V olume 1 ...
2019
-
[30]
EmoFilm - A multilingual emotional speech corpus,
E. Parada-Cabaleiro, G. Costantini, A. Batliner, A. Baird, and B. Schuller, “EmoFilm - A multilingual emotional speech corpus,” Sep. 2018. [Online]. Available: https://zenodo.org/records/7665999
2018
-
[31]
Vivaldi
“Vivaldi.” [Online]. Available: https://www2.hu-berlin.de/vivaldi/index. php
-
[32]
DEMoS: an Italian emotional speech corpus,
E. Parada-Cabaleiro, G. Costantini, A. Batliner, M. Schmitt, and B. W. Schuller, “DEMoS: an Italian emotional speech corpus,” Language Resources and Evaluation, vol. 54, no. 2, pp. 341–383, Jun. 2020. [Online]. Available: https://doi.org/10.1007/s10579-019-09450-y
2020 doi
-
[33]
DEMoS: an Italian emotional speech corpus. Elicitation methods, machine learning, and perception,
E. Parada-Cabaleiro, G. Costantini, A. Batliner, M. Schmitt, and B. Schuller, “DEMoS: an Italian emotional speech corpus. Elicitation methods, machine learning, and perception,” Feb. 2019. [Online]. Available: https://zenodo.org/records/2544829
2019
-
[34]
ITALIC: An Italian Intent Classification Dataset,
A. Koudounas, M. L. Quatra, L. Vaiani, L. Colomba, G. Attanasio, E. Pastor, L. Cagliero, and E. Baralis, “ITALIC: An Italian Intent Classification Dataset,” Jun. 2023, arXiv:2306.08502. [Online]. Available: http://arxiv.org/abs/2306.08502
2023 arXiv
-
[35]
ITALIC: An Italian Intent Classification Dataset,
A. Koudounas, M. La Quatra, L. Vaiani, L. Colomba, G. Attanasio, E. Pastor, L. Cagliero, and E. Baralis, “ITALIC: An Italian Intent Classification Dataset,” Jun. 2023. [Online]. Available: https://zenodo.org/ records/8040649
2023
-
[36]
Europarl ST corpus
“Europarl ST corpus.” [Online]. Available: https://mllp.upv.es/europarl-st/
-
[37]
Europarl-ST: A Multilingual Corpus for Speech Translation of Parliamentary Debates,
J. Iranzo-Sánchez, J. A. Silvestre-Cerdà, J. Jorge, N. Roselló, A. Giménez, A. Sanchis, J. Civera, and A. Juan, “Europarl-ST: A Multilingual Corpus for Speech Translation of Parliamentary Debates,” in ICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Si...
2020
-
[38]
PortMedia French and Italian corpus – ELRA Catalogue
“PortMedia French and Italian corpus – ELRA Catalogue.” [Online]. Available: https://catalogue.elra.info/en-us/repository/browse/ ELRA-S0371/
-
[39]
m-ailabs-dataset,
i. celeste [witchzard, “m-ailabs-dataset,” Nov. 2024, original-date: 2019- 03-21T07:46:35Z. [Online]. Available: https://github.com/imdatceleste/ m-ailabs-dataset
2024
-
[40]
MULTEXT Prosodic database – ELRA Catalogue
“MULTEXT Prosodic database – ELRA Catalogue.” [Online]. Available: https://catalogue.elra.info/en-us/repository/browse/ELRA-S0060/
-
[41]
APASCI – ELRA Catalogue
“APASCI – ELRA Catalogue.” [Online]. Available: https://catalogue.elra. info/en-us/repository/browse/ELRA-S0039/
-
[42]
Italian Kids Speech Recognition Corpus (Desktop) – ELRA Catalogue
“Italian Kids Speech Recognition Corpus (Desktop) – ELRA Catalogue.” [Online]. Available: https://catalogue.elra.info/en-us/repository/browse/ ELRA-S0228_98/
-
[43]
VOLIP: a corpus of spoken Italian and a virtuous example of reuse of linguistic resources
I. Alfano, F. Cutugno, A. D. Rosa, C. Iacobini, and R. Savy, “VOLIP: a corpus of spoken Italian and a virtuous example of reuse of linguistic resources.”
-
[44]
Methods and Tools for Prosodic Analysis of a Spoken Italian Corpus
M. Savino, M. Refice, and D. Daleno, “Methods and Tools for Prosodic Analysis of a Spoken Italian Corpus.” VOLUME XX, 202X 15
-
[45]
More Harm than Good: Why Dictionaries Using Orthographic Transcription Instead of the IPA Should Be Handled with Care,
A. Bryła-Cruz, “More Harm than Good: Why Dictionaries Using Orthographic Transcription Instead of the IPA Should Be Handled with Care,” Research in Language, vol. 20, no. 2, pp. 133–152, Dec. 2022, number: 2. [Online]. Available: https://czasopisma.uni.lodz.pl/research/ articl...
2022
-
[46]
RECOApy: Data Recording, Pre-Processing and Phonetic Transcription for End-to-End Speech-Based Applications,
A. Stan, “RECOApy: Data Recording, Pre-Processing and Phonetic Transcription for End-to-End Speech-Based Applications,” Interspeech 2020, pp. 586–590, Oct. 2020, conference Name: Interspeech 2020 Publisher: ISCA. [Online]. Available: https://www.isca-archive.org/ interspeech_2...
2020
-
[47]
Praat: doing Phonetics by Computer
“Praat: doing Phonetics by Computer.” [Online]. Available: https: //www.fon.hum.uva.nl/praat/
-
[48]
ELAN Annotator | The Language Archive
“ELAN Annotator | The Language Archive.” [Online]. Available: https://archive.mpi.nl/tla/elan
-
[49]
WebMAUS | Bavarian Archive for Speech Signals
“WebMAUS | Bavarian Archive for Speech Signals.” [Online]. Available: https://clarin.phonetik.uni-muenchen.de/BASWebServices/ interface/WebMAUSGeneral
-
[50]
EMU-SDMS: Advanced speech database management and analysis in R,
R. Winkelmann, J. Harrington, and K. Jänsch, “EMU-SDMS: Advanced speech database management and analysis in R,” Computer Speech & Language, vol. 45, pp. 392–410, Sep. 2017. [Online]. Available: https://linkinghub.elsevier.com/retrieve/pii/S0885230816302601
2017
-
[51]
SPPAS - SPeech Phonetization Alignment and Syllabification
“SPPAS - SPeech Phonetization Alignment and Syllabification.” [Online]. Available: https://sppas.org/
-
[52]
SPPAS: a tool for the phonetic segmentations of Speech,
B. Bigi, “SPPAS: a tool for the phonetic segmentations of Speech,” in The eighth international conference on Language Resources and Evaluation, Istanbul, Turkey, May 2012, pp. 1748–1755. [Online]. Available: https://hal.science/hal-00983701
2012
-
[53]
Annotald
“Annotald.” [Online]. Available: https://annotald.github.io/
-
[54]
Inter-annotator Agreement,
R. Artstein, “Inter-annotator Agreement,” in Handbook of Linguistic Annotation, N. Ide and J. Pustejovsky, Eds. Dordrecht: Springer Netherlands, 2017, pp. 297–313. [Online]. Available: https://doi.org/10. 1007/978-94-024-0881-2_11
2017
-
[55]
Audacity ® | Free Audio editor, recorder, music making and more!
“Audacity ® | Free Audio editor, recorder, music making and more!” [Online]. Available: https://www.audacityteam.org/
-
[56]
EMOVO Corpus: an Italian Emotional Speech Database
G. Costantini, I. Iadarola, A. Paoloni, and M. Todisco, “EMOVO Corpus: an Italian Emotional Speech Database.”
-
[57]
Speech emotion recognition with artificial intelligence for contact tracing in the COVID-19 pandemic,
F. Pucci, P. Fedele, and G. M. Dimitri, “Speech emotion recognition with artificial intelligence for contact tracing in the COVID-19 pandemic,” Cognitive Computation and Systems, vol. 5, no. 1, pp. 71–85, 2023, _eprint: https://onlinelibrary.wiley.com/doi/pdf/10.1049/ccs2.1207...
2023 doi
-
[58]
Evaluating and Optimizing Prosodic Alignment for Automatic Dubbing,
M. Federico, Y . Virkar, R. Enyedi, and R. Barra-Chicote, “Evaluating and Optimizing Prosodic Alignment for Automatic Dubbing,” in Interspeech
-
[59]
For a mapping of the languages/dialects of Italy and regional varieties of Italian,
P. Boula de Mareüil, E. Bilinski, F. Vernier, V . de Iacovo, and A. Romano, “For a mapping of the languages/dialects of Italy and regional varieties of Italian,” in New Ways of Analyzing Dialectal Variation, A. Thibault, M. Avanzi, N. L. Vecchio, and AliceMillour, Eds. Édition...
2021
-
[60]
The "untamed
B. Huszthy, “The "untamed" /s/ of Italian dialects: An overview of the singular behaviour of Italo-Romance sibilants,” Verbum – Analecta Neolatina, vol. 18, no. 1-2, pp. 189–214, Dec. 2017, number: 1-2. [Online]. Available: https://ojs.ppke.hu/verbum/article/view/345
2017
-
[61]
Speech Analysis of Language Varieties in Italy,
M. La Quatra, A. Koudounas, E. Baralis, and S. M. Siniscalchi, “Speech Analysis of Language Varieties in Italy,” in Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024), N. Calzolari, M.-Y . K...
2024
-
[62]
Custom Data Augmentation for low resource ASR using Bark and Retrieval-Based V oice Conversion,
A. Kamble, A. Tathe, S. Kumbharkar, A. Bhandare, and A. C. Mitra, “Custom Data Augmentation for low resource ASR using Bark and Retrieval-Based V oice Conversion,” 2023, publisher: arXiv Version Number: 3. [Online]. Available: https://arxiv.org/abs/2311.14836
2023 arXiv
-
[63]
ITAcotron 2: Trans- fering English Speech Synthesis Architectures and Speech Features to Italian
A. Favaro, L. Sbattella, R. Tedesco, and V . Scotti, “ITAcotron 2: Trans- fering English Speech Synthesis Architectures and Speech Features to Italian.”
-
[64]
Italian SpeechDat-Car database – ELRA Catalogue
“Italian SpeechDat-Car database – ELRA Catalogue.” [Online]. Available: https://catalogue.elra.info/en-us/repository/browse/ELRA-S0144/
-
[65]
Wake Words & V oice Commands Speech Data: Italian (Italy)
“Wake Words & V oice Commands Speech Data: Italian (Italy).” [Online]. Available: https://www.futurebeeai.com/dataset/ monologue-speech-dataset/wake-words-and-commands-italian-italy
-
[66]
Travel Scripted Monologue Speech Data: Italian (Italy)
“Travel Scripted Monologue Speech Data: Italian (Italy).” [Online]. Avail- able: https://www.futurebeeai.com/dataset/monologue-speech-dataset/ travel-scripted-speech-monologues-italian-italy
-
[67]
Telecom Scripted Monologue Speech Data: Italian (Italy)
“Telecom Scripted Monologue Speech Data: Italian (Italy).” [Online]. Available: https: //www.futurebeeai.com/dataset/monologue-speech-dataset/ telecom-scripted-speech-monologues-italian-italy
-
[68]
Retail & E-commerce Scripted Monologue Speech Data: Italian (Italy)
“Retail & E-commerce Scripted Monologue Speech Data: Italian (Italy).” [Online]. Available: https://www.futurebeeai.com/dataset/monologue-speech-dataset/ retail-scripted-speech-monologues-italian-italy
-
[69]
Healthcare Scripted Monologue Speech Data: Italian (Italy)
“Healthcare Scripted Monologue Speech Data: Italian (Italy).” [Online]. Available: https: //www.futurebeeai.com/dataset/monologue-speech-dataset/ healthcare-scripted-speech-monologues-italian-italy
-
[70]
Delivery & Logistics Scripted Monologue Speech Data: Italian (Italy)
“Delivery & Logistics Scripted Monologue Speech Data: Italian (Italy).” [Online]. Available: https: //www.futurebeeai.com/dataset/monologue-speech-dataset/ delivery-scripted-speech-monologues-italian-italy
-
[71]
BFSI Scripted Monologue Speech Data: Italian (Italy)
“BFSI Scripted Monologue Speech Data: Italian (Italy).” [Online]. Avail- able: https://www.futurebeeai.com/dataset/monologue-speech-dataset/ bfsi-scripted-speech-monologues-italian-italy
-
[72]
Travel Call Center Speech Data: Italian (Italy)
“Travel Call Center Speech Data: Italian (Italy).” [On- line]. Available: https://www.futurebeeai.com/dataset/speech-dataset/ travel-call-center-conversation-italian-italy
-
[73]
Telecom Call Center Speech Data: Italian (Italy)
“Telecom Call Center Speech Data: Italian (Italy).” [On- line]. Available: https://www.futurebeeai.com/dataset/speech-dataset/ telecom-call-center-conversation-italian-italy
-
[74]
Real Estate Call Center Speech Data: Italian (Italy)
“Real Estate Call Center Speech Data: Italian (Italy).” [Online]. Available: https://www.futurebeeai.com/dataset/speech-dataset/ realestate-call-center-conversation-italian-italy
-
[75]
Healthcare Call Center Speech Data: Italian (Italy)
“Healthcare Call Center Speech Data: Italian (Italy).” [On- line]. Available: https://www.futurebeeai.com/dataset/speech-dataset/ healthcare-call-center-conversation-italian-italy
-
[76]
BFSI Call Center Speech Data: Italian (Italy)
“BFSI Call Center Speech Data: Italian (Italy).” [On- line]. Available: https://www.futurebeeai.com/dataset/speech-dataset/ bfsi-call-center-conversation-italian-italy
-
[77]
Retail & E-commerce Call Center Speech Data: Italian (Italy)
“Retail & E-commerce Call Center Speech Data: Italian (Italy).” [Online]. Available: https://www.futurebeeai.com/dataset/speech-dataset/ retail-call-center-conversation-italian-italy
-
[78]
Delivery & Logistics Call Center Speech Data: Italian (Italy)
“Delivery & Logistics Call Center Speech Data: Italian (Italy).” [Online]. Available: https://www.futurebeeai.com/dataset/speech-dataset/ delivery-call-center-conversation-italian-italy
-
[79]
Italian (Italy) General Conversation Speech Dataset
“Italian (Italy) General Conversation Speech Dataset.” [On- line]. Available: https://www.futurebeeai.com/dataset/speech-dataset/ general-conversation-italian-italy
-
[80]
voxforge.org
“voxforge.org.” [Online]. Available: https://www.voxforge.org/home
-
[81]
ASR-ItaCSC: An Italian Conversational Speech Corpus - MagicHub
“ASR-ItaCSC: An Italian Conversational Speech Corpus - MagicHub.” [Online]. Available: https://magichub.com/datasets/ italian-conversational-speech-corpus/
-
[82]
Consonant gemination in Italian: the nasal and liquid case,
M.-G. D. Benedetto and L. D. Nardis, “Consonant gemination in Italian: the nasal and liquid case,” Sep. 2020, arXiv:2005.06960. [Online]. Available: http://arxiv.org/abs/2005.06960
2020 arXiv
-
[83]
GDPR - Regulation - 2016/679 - EN - EUR-Lex,
“GDPR - Regulation - 2016/679 - EN - EUR-Lex,” doc ID: 32016R0679 Doc Sector: 3 Doc Title: Regulation (EU) 2016/679 of the European Parliament and of the Council of 27 April 2016 on the protection of natural persons with regard to the processing of personal data and on the fre...
2016
-
[2020]
2020, pp
ISCA, Oct. 2020, pp. 1481–1485. [Online]. Available: https: //www.isca-archive.org/interspeech_2020/federico20_interspeech.html
2020
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.