REVIEW 2 major objections 5 minor 1 cited by
TaigiSpeech supplies a real elderly Taiwanese Hokkien intent corpus and shows that models trained only on mined drama speech lose substantial accuracy on it.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-13 20:49 UTC pith:4OAJCZXP
load-bearing objection Solid elderly Taigi intent corpus plus a clean domain-mismatch result; mining is useful scaffolding, not the main claim. the 2 major comments →
TaigiSpeech: A Low-Resource Real-World Speech Intent Dataset and Preliminary Results with Scalable Data Mining In-the-Wild
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Models trained solely on intent segments mined from Taiwanese drama videos suffer large accuracy drops when evaluated on real-world elderly TaigiSpeech recordings, demonstrating a domain mismatch that makes a purpose-built elderly corpus necessary for low-resource spoken intent recognition.
What carries the argument
TaigiSpeech (the 8-intent elderly corpus) together with the two mining pipelines—keyword-match mining plus LLM pseudo-labeling via Mandarin subtitles, and PE-AV audio-visual retrieval—used to create the pre-training data against which the domain gap is measured.
Load-bearing premise
The claim rests on treating LLM pseudo-labels generated from Mandarin drama subtitle windows as accurate enough ground truth for both the mined training pool and the drama test set.
What would settle it
Independent human re-annotation of a random sample of the LLM-pseudo-labeled drama segments would show low agreement with the automatic labels, or models trained on those labels would perform near chance even on held-out drama speech.
If this is right
- Realistic elderly-speech benchmarks become a prerequisite before home-assistant or emergency systems can be trusted for low-resource languages.
- Keyword-match mining with an intermediate high-resource language yields stronger subsequent adaptation than pure audio-visual mining under current multimodal encoders.
- End-to-end SSL speech classifiers outperform cascaded foundation ASR-plus-LLM pipelines on this task while using far fewer parameters.
- A few hundred in-domain elderly utterances are enough to recover most of the accuracy lost to domain mismatch.
Where Pith is reading between the lines
- Comparable domain gaps are likely for other primarily oral or unwritten languages when training data are mined from broadcast drama or television.
- The persistent gap between lightweight MatchboxNet models and SSL models points to a need for better distillation if elderly intent systems must run on edge devices.
- Mining from more naturalistic elderly speech sources, rather than drama alone, could shrink the mismatch without requiring additional manual labels.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces TaigiSpeech, a spoken-intent corpus of 3,079 utterances from 21 elderly Taiwanese Hokkien speakers (ages 54–78) covering eight intents (four emergency, four functional home-assistant commands). It further explores two in-the-wild mining pipelines—keyword matching on Mandarin drama subtitles followed by LLM pseudo-labeling, and weakly supervised audio-visual retrieval with PE-AV—and evaluates lightweight MatchboxNet and SSL models (HuBERT, WavLM). Table 7 shows a clear domain gap: models trained only on mined drama data drop substantially on the held-out TaigiSpeech test set (e.g., WavLM-large 92.36 % → 70.00 % under the 5-class keyword setting) and recover only after fine-tuning on a small speaker-independent Taigi train split. Cascaded foundation ASR+LLM baselines are also reported and underperform end-to-end SSL classifiers. The dataset is released under CC BY 4.0.
Significance. The work fills a genuine gap: existing emergency/smart-home SLU corpora are almost exclusively high-resource languages and rarely target elderly speakers of primarily spoken languages. TaigiSpeech supplies a realistic, scenario-driven elderly Hokkien benchmark with speaker-independent splits, bootstrap CIs, and public release. The measured domain-mismatch result (Table 7) is load-bearing and directly supports the claim that mined drama data alone are insufficient, thereby justifying the new corpus. The dual mining strategies, while preliminary, offer a practical template for other unwritten languages. Strengths include transparent statistics (Tables 1, 4, 6, 8), multi-model evaluation, and explicit acknowledgment of speaker-count and AV-scale limitations.
major comments (2)
- §4.1 and Table 6: LLM (Gemini-3) pseudo-labels serve both as training targets for keyword mining and as the sole ground truth for the Drama test set. No human validation or inter-annotator agreement is reported for these labels. While the primary TaigiSpeech evaluation is independent, the magnitude of the reported domain gap (Drama C=5 vs. Taigi C=5) and the claimed value of mined pre-training rest partly on the quality of these labels; a modest human audit of a stratified sample would strengthen the central claim.
- §5.1 / Table 8: The Audio-Visual mining experiment uses only a 28 k-clip subset already filtered by keyword queries rather than the full 7 000-hour drama pool, and collapses to a binary emergency/non-emergency task. The near-random Taigi (C=2) transfer results are therefore difficult to interpret as a fair test of the multimodal approach; either a larger unfiltered pool or an explicit statement that AV mining remains exploratory would better calibrate the claim of “scalable data mining.”
minor comments (5)
- Table 2 and §2.2: A few related emergency corpora (e.g., more recent PERS or multi-lingual AAL sets) could be cited for completeness; the current survey is already useful but slightly incomplete.
- Figure 6: Confusion matrices are informative; adding per-class F1 or a short discussion of why LIGHT ON/OFF remain confusable after adaptation would help readers.
- §3.2 / Supplementary A.2–A.3: Scenario and video prompts are generated by Gemini/Veo; a brief note on human filtering criteria would improve reproducibility of the elicitation protocol.
- Minor typos: “T his finding” (Conclusion), inconsistent capitalization of intent names across tables, and occasional missing spaces after periods.
- §6.4: Cascaded ASR+LLM results would be more informative if the authors stated whether any prompt engineering or few-shot examples were used for the Qwen3-8B intent classifier.
Circularity Check
No significant circularity: empirical dataset release and transfer experiments with independent human-elicited TaigiSpeech evaluation.
full rationale
The paper is a dataset contribution plus preliminary mining and baseline experiments. Its central empirical claim (domain mismatch) is the measured accuracy drop when models trained on drama-mined data are evaluated on the fixed, independently collected TaigiSpeech elderly test set of 960 human-elicited utterances (Table 7: e.g. WavLM-large 92.36 % Drama C=5 o 70.00 % Taigi C=5). Drama labels are LLM pseudo-labels (Gemini-3) and thus measure consistency with that model, which the paper itself states; this does not make the Taigi numbers circular, because Taigi labels come from the scenario-driven recording protocol, not from the LLM or from any fitted parameter of the mining pipeline. No equation, uniqueness theorem, or ansatz is imported via self-citation to force a result; no parameter is fitted to a quantity that is then re-presented as a prediction. The work is therefore self-contained against its own external benchmark (the released TaigiSpeech recordings). Score 0 is appropriate.
Axiom & Free-Parameter Ledger
free parameters (2)
- top-k retrieval size for AV mining
- number of keywords per intent
axioms (4)
- domain assumption Mandarin subtitles of Taiwanese drama are sufficiently faithful translations for keyword matching and LLM intent verification
- domain assumption LLM (Gemini-3) pseudo-labels on contextualized subtitle windows constitute usable ground truth for both training and Drama test evaluation
- domain assumption Imagined scenario prompts plus optional silent videos elicit speech whose acoustic and lexical properties are representative of real emergency and home-assistant use
- standard math Standard self-supervised speech models (HuBERT, WavLM) and MatchboxNet architectures transfer to Taiwanese Hokkien after fine-tuning
read the original abstract
Speech technologies have advanced rapidly and serve diverse populations worldwide. However, many languages remain underrepresented due to limited resources. In this paper, we introduce \textbf{TaigiSpeech}, a real-world speech intent dataset in Taiwanese Taigi (aka Taiwanese Hokkien/Southern Min), which is a low-resource and primarily spoken language. The dataset is collected from older adults, comprising 21 speakers with a total of 3k utterances. It is designed for practical intent detection scenarios, including healthcare and home assistant applications. To address the scarcity of labeled data, we explore two data mining strategies with two levels of supervision: keyword match data mining with LLM pseudo labeling via an intermediate language and an audio-visual framework that leverages multimodal cues with minimal textual supervision. This design enables scalable dataset construction for low-resource and unwritten spoken languages. TaigiSpeech will be released under the CC BY 4.0 license to facilitate broad adoption and research on low-resource and unwritten languages. The project website and the dataset can be found on https://kwchang.org/taigispeech.
Forward citations
Cited by 1 Pith paper
-
Walking Through Uncertainty: An Empirical Study of Uncertainty Estimation for Audio-Aware Large Language Models
Semantic-level and verification-based uncertainty methods outperform token-level baselines for audio reasoning in ALLMs, but their relative performance on hallucination and unanswerable-question benchmarks is model- a...
Reference graph
Works this paper leans on
-
[1]
How- ever, many of them are not supported by commercial AI ser- vices
Introduction More than 7,000 languages are spoken worldwide [1]. How- ever, many of them are not supported by commercial AI ser- vices. A substantial portion of these languages lack a univer- sally standardized writing system, leading to inconsistent or- thographic resources, or are entirely unwritten [2]. As a result, they remain underrepresented in mode...
-
[2]
SOS CALL Calls for emergency help
-
[3]
Reports breathing difficulty
BREATH EMERG. Reports breathing difficulty
-
[4]
FALL HELP Indicates a fall and requests help
-
[5]
Non-Emergency Intents / Functional Commands
PAIN GENERAL Reports physical pain or discomfort. Non-Emergency Intents / Functional Commands
-
[6]
CALL CONTACT Requests to contact a person
-
[7]
LIGHT ON Requests to turn on the lights
-
[8]
LIGHT OFF Requests to turn off the lights
-
[9]
CANCEL Cancels a previously triggered alert. Total Utterances: 3,079 ������������ ������� ������������� ������ �� ��� ��� ��� ��� ���� ����� ������� ���� ������� ���� ������� ���� ������� �������� � �������� Figure 1:Primary language use among residents in Taiwan. El- ders (aged 65+) have 64.9% reporting Taiwanese as their pri- mary language. Data source:...
Pith/arXiv arXiv 2020
-
[10]
Related Works 2.1. Spoken Language Understanding Spoken language understanding (SLU) aims to extract semantic meaning from speech signals and serves as a core component in voice assistants and spoken dialogue systems [43]. A cen- tral subtask in SLU isintent recognition, which identifies the user’s underlying intention from spoken utterances and triggers ...
-
[11]
Participant Registration and Recording Interface Figure 2:Recording app interface
TaigiSpeech In this section, we describe the data collection procedure, partic- ipant recruitment process, scenario design, and basic statistics of the TaigiSpeech dataset 3.1. Participant Registration and Recording Interface Figure 2:Recording app interface. The English translations are shown for illustrative purposes in this paper. All recordings were c...
-
[12]
SOS CALL 387 7.78 3.28 50.16
-
[13]
FALL HELP 385 7.77 3.73 49.88
-
[14]
BREATH EMERG 385 8.22 3.30 52.71
-
[15]
PAIN GENERAL 385 7.91 3.73 50.79
-
[16]
CALL CONTACT 385 7.04 5.30 45.16
-
[17]
LIGHT ON 384 5.95 2.81 38.08
-
[18]
LIGHT OFF 384 5.67 2.35 36.28
-
[19]
Ahh! Fire! It’s on fire! Help!
CANCEL ALERT 384 6.85 3.17 43.85 Overall 3,079 7.15 3.66 366.91 Figure 3:Age distribution of speakers in the TaigiSpeech cor- pus. Our study focuses on elderly participants. Figure 4:The distribution of ambient noise level (dB). a mouse or touchscreen. 3.2. Scenario Design TaigiSpeech targets 8 intents, covering both emergency and non-emergency functional...
-
[20]
help”,“call the ambulance
Data Mining in-the-wild For many low-resource and unwritten languages, reliable ASR systems are often unavailable. Also, in most cases, speech data primarily exist in web-based sources (e.g., online videos), where recordings are noisy and loosely structured. These con- straints make large-scale supervised data collection impracti- cal, motivating the deve...
2074
-
[21]
Experimental Setup For pseudo-labeling in the Keyword Mining approach, we em- ploy Gemini-3 to generate intermediate annotations
Experiments 5.1. Experimental Setup For pseudo-labeling in the Keyword Mining approach, we em- ploy Gemini-3 to generate intermediate annotations. For audio- visual data mining, we adopt the PE-A V large [58] to ex- tract multimodal representations and identify candidate seg- ments without relying on paired textual supervision. Note that the PE-A V model ...
-
[22]
We summarize the key observations in the following sections
Preliminary Results The main results are presented in Table 7. We summarize the key observations in the following sections. 6.1. Domain mismatch A clear performance discrepancy is observed between evalua- tions on the drama dataset and on the real-world TaigiSpeech recordings. This trend consistently appears in both settings: Drama (C=5)�Taigi (C=5)under ...
-
[23]
Conclusion and Future Work In this work, we introduce TaigiSpeech, a Taiwanese Hokkien spoken intent dataset designed for elderly users in home- assistant and healthcare scenarios. The dataset com- prises 8 intent categories, including four emergency intents (SOS CALL, BREATH EMERGENCY, FALL HELP, and PAIN GENERAL) and four non-emergency functional intent...
-
[24]
This work was supported in part by the National Science and Technology Council (NSTC), Taiwan, under Grant No
Acknowledgments The authors would like to express their sincere gratitude to the Taiwanese Language and Culture Club of Keelung Community University for their invaluable support and assistance, and to the Taipei City Nangang Social Welfare Center for their advice and suggestions. This work was supported in part by the National Science and Technology Counc...
-
[25]
All authors take full respon- sibility for the content of this paper, including its accuracy, orig- inality, and conclusions
Generative AI Use Disclosure Generative AI tools were used solely for language editing and minor polishing of the manuscript. All authors take full respon- sibility for the content of this paper, including its accuracy, orig- inality, and conclusions. No generative AI tools were used to produce substantial portions of the scientific content, analysis, or results
-
[26]
W. R. Leben, “Languages of the World,” 02 2018. [Online]. Avail- able: https://oxfordre.com/linguistics/view/10.1093/acrefore/ 9780199384655.001.0001/acrefore-9780199384655-e-349
-
[27]
D. M. Eberhard, G. F. Simons, and A. J. Robinson, Eds.,Ethno- logue: Languages of the World, 29th ed. SIL Global, 2026
2026
-
[28]
Kwok,Southern Min: Comparative phonology and sub- grouping
B.-C. Kwok,Southern Min: Comparative phonology and sub- grouping. Routledge, 2018
2018
-
[29]
The influence of Southern Min on the Mandarin of Taiwan,
C. C. Kubler, “The influence of Southern Min on the Mandarin of Taiwan,”Anthropological Linguistics, pp. 156–176, 1985
1985
-
[30]
Language Usage for the Resident Nationals Aged 6 Years and Over - 2020 Population and Housing Census,
A. Directorate General of Budget and T. Statistics (DGBAS), “Language Usage for the Resident Nationals Aged 6 Years and Over - 2020 Population and Housing Census,” 2021
2020
-
[31]
Challenges in real-life emotion annotation and machine learning based detection,
L. Devillers, L. Vidrascu, and L. Lamel, “Challenges in real-life emotion annotation and machine learning based detection,”Neu- ral Networks, vol. 18, no. 4, pp. 407–422, 2005, emotion and Brain
2005
-
[32]
A Multimodal Corpus Recorded in a Health Smart Home,
A. Fleury, M. Vacher, F. Portet, P. Chahuara, and N. Noury, “A Multimodal Corpus Recorded in a Health Smart Home,” inLREC 2010, The International Conference on Language Resources and Evaluation, Valetta, Malta, May 2010, pp. 99–105
2010
-
[33]
The CARES corpus: a database of older adult actor simulated emergency dialogue for developing a personal emergency response system,
V . Young and A. Mihailidis, “The CARES corpus: a database of older adult actor simulated emergency dialogue for developing a personal emergency response system,”International Journal of Speech Technology, vol. 16, no. 1, pp. 55–73, 2013
2013
-
[34]
A French corpus of audio and multimodal interactions in a health smart home,
A. Fleury, M. Vacher, F. Portet, P. Chahuara, and N. Noury, “A French corpus of audio and multimodal interactions in a health smart home,”Journal on Multimodal User Interfaces, vol. 7, no. 1, pp. 93–109, 2013
2013
-
[35]
The Sweet-Home speech and multimodal corpus for home automation interaction,
M. Vacher, B. Lecouteux, P. Chahuara, F. Portet, B. Meillon, and N. Bonnefond, “The Sweet-Home speech and multimodal corpus for home automation interaction,” inLREC 2014, Reykjavik, Ice- land, May 2014, pp. 4499–4506
2014
-
[36]
An integrated system for voice command recognition and emergency detection based on audio signals,
E. Principi, S. Squartini, R. Bonfigli, G. Ferroni, and F. Pi- azza, “An integrated system for voice command recognition and emergency detection based on audio signals,”Expert Syst. Appl., vol. 42, no. 13, p. 5668–5683, Aug. 2015
2015
-
[37]
Speech commands: A dataset for limited-vocabulary speech recognition,
P. Warden, “Speech commands: A dataset for limited-vocabulary speech recognition,”arXiv preprint arXiv:1804.03209, 2018
Pith/arXiv arXiv 2018
-
[38]
Speech Model Pre-Training for End-to-End Spoken Language Understanding,
L. Lugosch, M. Ravanelli, P. Ignoto, V . S. Tomar, and Y . Bengio, “Speech Model Pre-Training for End-to-End Spoken Language Understanding,” inInterspeech 2019, 2019, pp. 814–818
2019
-
[39]
Context-Aware V oice-Based Interaction in Smart Home - V ocADom@A4H Corpus Collection and Empirical Assessment of Its Useful- ness,
F. Portet, S. Caffiau, F. Ringeval, M. Vacher, N. Bonne- fond, S. Rossato, B. Lecouteux, and T. Desot, “Context-Aware V oice-Based Interaction in Smart Home - V ocADom@A4H Corpus Collection and Empirical Assessment of Its Useful- ness,” in2019 IEEE Intl Conf on Dependable, Autonomic and Secure Computing, Intl Conf on Pervasive Intelligence and Computing, ...
2019
-
[40]
SLURP: A spoken language understanding resource package,
E. Bastianelli, A. Vanzo, P. Swietojanski, and V . Rieser, “SLURP: A spoken language understanding resource package,” inProceed- ings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2020, pp. 7252–7262
2020
-
[41]
Learning Asr-Robust Contex- tualized Embeddings for Spoken Language Understanding,
C.-W. Huang and Y .-N. Chen, “Learning Asr-Robust Contex- tualized Embeddings for Spoken Language Understanding,” in ICASSP 2020 - 2020 IEEE International Conference on Acous- tics, Speech and Signal Processing (ICASSP), 2020, pp. 8009– 8013
2020
-
[42]
EMSAssist: An End-to-End Mobile V oice Assis- tant at the Edge for Emergency Medical Services,
L. Jin, T. Liu, A. Haroon, R. Stoleru, M. Middleton, Z. Zhu, and T. Chaspari, “EMSAssist: An End-to-End Mobile V oice Assis- tant at the Edge for Emergency Medical Services,” ser. MobiSys ’23. New York, NY , USA: Association for Computing Machin- ery, 2023, p. 275–288
2023
-
[43]
VN-SLU: A Vietnamese Spoken Language Un- derstanding Dataset,
T. Tran, K. Le, N. D. Nguyen, M. Vu, H. Ngo, W. Park, and T. T. T. Nguyen, “VN-SLU: A Vietnamese Spoken Language Un- derstanding Dataset,” inInterspeech 2024, 2024, pp. 1335–1339
2024
-
[44]
Enhancing V oice Wake-Up for Dysarthria: Man- darin Dysarthria Speech Corpus Release and Customized System Design,
M. Gao, H. Chen, J. Du, X. Xu, H. Guo, H. Bu, J. Yang, M. Li, and C.-H. Lee, “Enhancing V oice Wake-Up for Dysarthria: Man- darin Dysarthria Speech Corpus Release and Customized System Design,” inInterspeech 2024, 2024, pp. 2465–2469
2024
-
[45]
Examining the role of living alone and loneliness in predicting health-related quality of life: results from the Healthy Aging Longitudinal Study in Taiwan (HALST),
H.-Y . Tseng, C.-Y . Lee, C.-S. Wu, I.-C. Wu, H.-Y . Chang, C.- C. Hsu, and C. A. Hsiung, “Examining the role of living alone and loneliness in predicting health-related quality of life: results from the Healthy Aging Longitudinal Study in Taiwan (HALST),” Quality of Life Research, vol. 33, no. 4, pp. 1015–1028, 2024
2024
-
[46]
Y . Lin, W. Chen, H. Yu, and I. Cho, “Spatial Distribution, Characterization, and Policy Opportunities for Taiwan’s Solo Elderly: A Big Data Approach,” inProceedings of DRS2024: Boston, C. Gray, E. Ciliotta Chehade, P. Hekkert, L. Forlano, P. Ciuccarelli, and P. Lloyd, Eds., Boston, USA, Jun. 2024. [Online]. Available: https://doi.org/10.21606/drs.2024.366
-
[47]
K.-M. Chen, S.-T. Wang, S.-R. Chao, and K. Kasirisir, “Link Between Social Relationships and Vulnerability Among Community-Dwelling Older Adults Living Alone in Taiwan,” Asian Social Work and Policy Review, vol. 19, no. 1, p. e70000, 2025, e70000 ASWP-May-2024-0051.R1. [Online]. Available: https://onlinelibrary.wiley.com/doi/abs/10.1111/aswp.70000
-
[48]
Digital health literacy and its deter- minants among community dwelling elderly people in Taiwan,
T. T. Tran, P. W. Chang, J.-M. Yang, T.-H. Chen, C.-T. Su, D. Levin-Zamir, O. Baron-Epel, E. Neter, S. F. Tsai, B. Lo, T. V . Duong, and S.-H. Yang, “Digital health literacy and its deter- minants among community dwelling elderly people in Taiwan,” DIGITAL HEALTH, vol. 10, p. 20552076241278926, 2024
2024
-
[49]
A smartphone model for post-acute care decreases all- cause mortality with improved left ventricular ejection fraction in patients hospitalized with heart failure in Taiwan,
S.-C. Weng, W.-W. Lin, J.-L. Huang, C.-Y . Chao, C.-Y . Hsu, and S.-Y . Lin, “A smartphone model for post-acute care decreases all- cause mortality with improved left ventricular ejection fraction in patients hospitalized with heart failure in Taiwan,”Maturitas, vol. 197, p. 108269, 2025. [Online]. Available: https://www. sciencedirect.com/science/article...
2025
-
[50]
From command to care: a scoping re- view on utilization of smart speakers by patients and providers,
R. Saripalle and R. Patel, “From command to care: a scoping re- view on utilization of smart speakers by patients and providers,” Mayo Clinic Proceedings: Digital Health, vol. 2, no. 2, pp. 207– 220, 2024
2024
-
[51]
Falls in older people: epidemiology, risk fac- tors and strategies for prevention,
L. Z. Rubenstein, “Falls in older people: epidemiology, risk fac- tors and strategies for prevention,”Age and ageing, vol. 35, no. suppl 2, pp. ii37–ii41, 2006
2006
-
[52]
The Growing Burden of Fall-Related Injuries Among Older Adults: A Seven-Year Study from a Tertiary Medical Center in Taiwan,
Y .-C. Tsai, S.-Y . Chen, Y .-M. Chen, E. P.-C. Huang, and F.-P. Lu, “The Growing Burden of Fall-Related Injuries Among Older Adults: A Seven-Year Study from a Tertiary Medical Center in Taiwan,”Annals of Geriatric Medicine and Research,
-
[53]
Available: http://www.e-agmr.org/journal/view
[Online]. Available: http://www.e-agmr.org/journal/view. php?number=1240
-
[54]
Concept of the term long lie: a scoping review,
J. Kubitza, I. T. Schneider, and B. Reuschenbach, “Concept of the term long lie: a scoping review,”European Review of Aging and Physical Activity, vol. 20, no. 1, p. 16, 2023
2023
-
[55]
SmartLohas: A Smart Assistive System for Elder People,
Y .-T. Tsai, C.-L. Fan, C.-L. Lo, and S.-H. Huang, “SmartLohas: A Smart Assistive System for Elder People,” in2017 14th Inter- national Symposium on Pervasive Systems, Algorithms and Net- works & 2017 11th International Conference on Frontier of Com- puter Science and Technology & 2017 Third International Sym- posium of Creative Computing (ISPAN-FCST-ISCC), 2017
2017
-
[56]
V oice Health As- sistant for Elderly with Multi-Language and Offline Support,
R. Chithra, R. Ackshaya, and T. Dhinakaran, “V oice Health As- sistant for Elderly with Multi-Language and Offline Support,” in 2025 10th International Conference on Smart Structures and Sys- tems (ICSSS), 2025, pp. 1–7
2025
-
[57]
Development of an automated speech recognition interface for personal emer- gency response systems,
M. Hamill, V . Young, J. Boger, and A. Mihailidis, “Development of an automated speech recognition interface for personal emer- gency response systems,”Journal of NeuroEngineering and Re- habilitation, vol. 6, no. 1, p. 26, 2009
2009
-
[58]
Evaluation of a context-aware voice interface for ambient assisted living: qualitative user study vs. quantitative system evaluation,
M. Vacher, S. Caffiau, F. Portet, B. Meillon, C. Roux, E. Elias, B. Lecouteux, and P. Chahuara, “Evaluation of a context-aware voice interface for ambient assisted living: qualitative user study vs. quantitative system evaluation,”ACM Transactions on Acces- sible Computing (TACCESS), vol. 7, no. 2, pp. 1–36, 2015
2015
-
[59]
Making emergency calls more accessible to older adults through a hands-free speech interface in the house,
M. Vacher, F. Aman, S. Rossato, F. Portet, and B. Lecouteux, “Making emergency calls more accessible to older adults through a hands-free speech interface in the house,”ACM Transactions on Accessible Computing (TACCESS), vol. 12, no. 2, pp. 1–25, 2019
2019
-
[60]
V oice-controlled intelligent personal assistants to sup- port aging in place,
K. O’Brien, A. Liggett, V . Ramirez-Zohfeld, P. Sunkara, and L. A. Lindquist, “V oice-controlled intelligent personal assistants to sup- port aging in place,”Journal of the American Geriatrics Society, vol. 68, no. 1, pp. 176–179, 2020
2020
-
[61]
Speech-massive: A multilingual speech dataset for slu and be- yond,
B. Lee, I. Calapodescu, M. Gaido, M. Negri, and L. Besacier, “Speech-massive: A multilingual speech dataset for slu and be- yond,” inProc. Interspeech 2024, 2024, pp. 817–821
2024
-
[62]
Catslu: The 1st chinese audio-textual spoken language understanding chal- lenge,
S. Zhu, Z. Zhao, T. Zhao, C. Zong, and K. Yu, “Catslu: The 1st chinese audio-textual spoken language understanding chal- lenge,” in2019 International Conference on Multimodal Interac- tion, 2019, pp. 521–525
2019
-
[63]
Su ´ısiann dataset,
`I-thuˆan, “Su ´ısiann dataset,” https://suisiann-dataset.ithuan.tw/, 2019
2019
-
[64]
Taiwanese across taiwan corpus and its applications,
Y .-F. Liao, J. S. Tsay, P. Kang, H.-L. Khoo, L.-K. Tan, L.- C. Chang, U.-G. Iunn, H.-L. Su, T.-G. Thiann, H.-K. Tiun et al., “Taiwanese across taiwan corpus and its applications,” in 2022 25th Conference of the Oriental COCOSDA International Committee for the Co-ordination and Standardisation of Speech Databases and Assessment Techniques (O-COCOSDA). IEE...
2022
-
[65]
Com- mon voice: A massively-multilingual speech corpus,
R. Ardila, M. Branson, K. Davis, M. Kohler, J. Meyer, M. Hen- retty, R. Morais, L. Saunders, F. Tyers, and G. Weber, “Com- mon voice: A massively-multilingual speech corpus,” inProceed- ings of the twelfth language resources and evaluation conference, 2020, pp. 4218–4222
2020
-
[66]
Ml-superb: Multilingual speech universal performance benchmark,
J. Shi, D. Berrebbi, W. Chen, E.-P. Hu, W.-P. Huang, H.-L. Chung, X. Chang, S.-W. Li, A. Mohamed, H.-y. Leeet al., “Ml-superb: Multilingual speech universal performance benchmark,” inProc. Interspeech 2023, 2023, pp. 884–888
2023
-
[67]
C.-Y . Yang, C.-C. Wang, L.-W. Chen, H.-S. Lee, H.-M. Wang, and B. Chen, “Tg-asr: Translation-guided learning with parallel gated cross attention for low-resource automatic speech recognition,” arXiv preprint arXiv:2602.22039, 2026
arXiv 2026
-
[68]
Evaluating self-supervised speech models on a taiwanese hokkien corpus,
Y .-H. Chou, K. Chang, M.-J. Wu, W. Ou, A. W.-H. Bi, C. Yang, B. Y . Chen, R.-W. Pai, P.-Y . Yeh, J.-P. Chianget al., “Evaluating self-supervised speech models on a taiwanese hokkien corpus,” in2023 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU). IEEE, 2023, pp. 1–7
2023
-
[69]
Espnet-slu: Advancing spoken language understanding through espnet,
S. Arora, S. Dalmia, P. Denisov, X. Chang, Y . Ueda, Y . Peng, Y . Zhang, S. Kumar, K. Ganesan, B. Yanet al., “Espnet-slu: Advancing spoken language understanding through espnet,” in ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2022
2022
-
[70]
The atis spoken language systems pilot corpus,
C. T. Hemphill, J. J. Godfrey, and G. R. Doddington, “The atis spoken language systems pilot corpus,” inSpeech and Natural Language: Proceedings of a Workshop Held at Hidden Valley, Pennsylvania, June 24-27, 1990, 1990
1990
-
[71]
Slue: New benchmark tasks for spoken language un- derstanding evaluation on natural speech,
S. Shon, A. Pasad, F. Wu, P. Brusco, Y . Artzi, K. Livescu, and K. J. Han, “Slue: New benchmark tasks for spoken language un- derstanding evaluation on natural speech,” inICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2022, pp. 7927–7931
2022
-
[72]
Self-supervised speech representation learning: A review,
A. Mohamed, H.-y. Lee, L. Borgholt, J. D. Havtorn, J. Edin, C. Igel, K. Kirchhoff, S.-W. Li, K. Livescu, L. Maaløeet al., “Self-supervised speech representation learning: A review,”IEEE Journal of Selected Topics in Signal Processing, vol. 16, no. 6, pp. 1179–1210, 2022
2022
-
[73]
Hubert: Self-supervised speech represen- tation learning by masked prediction of hidden units,
W.-N. Hsu, B. Bolte, Y .-H. H. Tsai, K. Lakhotia, R. Salakhutdi- nov, and A. Mohamed, “Hubert: Self-supervised speech represen- tation learning by masked prediction of hidden units,”IEEE/ACM transactions on audio, speech, and language processing, vol. 29, pp. 3451–3460, 2021
2021
-
[74]
Wavlm: Large-scale self- supervised pre-training for full stack speech processing,
S. Chen, C. Wang, Z. Chen, Y . Wu, S. Liu, Z. Chen, J. Li, N. Kanda, T. Yoshioka, X. Xiaoet al., “Wavlm: Large-scale self- supervised pre-training for full stack speech processing,”IEEE Journal of Selected Topics in Signal Processing, vol. 16, no. 6, pp. 1505–1518, 2022
2022
-
[75]
A large-scale evaluation of speech foundation models,
S.-w. Yang, H.-J. Chang, Z. Huang, A. T. Liu, C.-I. Lai, H. Wu, J. Shi, X. Chang, H.-S. Tsai, W.-C. Huanget al., “A large-scale evaluation of speech foundation models,”IEEE/ACM Transac- tions on Audio, Speech, and Language Processing, vol. 32, 2024
2024
-
[76]
Robust speech recognition via large-scale weak supervision,
A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever, “Robust speech recognition via large-scale weak supervision,” inInternational conference on machine learning. PMLR, 2023, pp. 28 492–28 518
2023
-
[77]
X. Shi, X. Wang, Z. Guo, Y . Wang, P. Zhang, X. Zhang, Z. Guo, H. Hao, Y . Xi, B. Yanget al., “Qwen3-asr technical report,”arXiv preprint arXiv:2601.21337, 2026
Pith/arXiv arXiv 2026
-
[78]
An official American Thoracic Society statement: update on the mechanisms, assessment, and management of dyspnea,
M. B. Parshall, R. M. Schwartzstein, L. Adams, R. B. Banzett, H. L. Manning, J. Bourbeau, P. M. Calverley, A. G. Gift, A. Harver, S. C. Lareau, D. A. Mahler, P. M. Meek, D. E. O’Donnell, and American Thoracic Society Committee on Dyspnea, “An official American Thoracic Society statement: update on the mechanisms, assessment, and management of dyspnea,”Ame...
2012
-
[79]
Gemini: A Family of Highly Capable Multimodal Models,
G. Team, “Gemini: A Family of Highly Capable Multimodal Models,” 2025. [Online]. Available: https://arxiv.org/abs/2312. 11805
2025
-
[80]
Veo 3: Advancing High-Fidelity Generative Video Modeling,
G. Team and D. Team, “Veo 3: Advancing High-Fidelity Generative Video Modeling,” Google DeepMind, Tech. Rep., 2025, accessed: 2026-02-23. [Online]. Available: https://storage. googleapis.com/deepmind-media/veo/Veo-3-Tech-Report.pdf
2025
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.