Pith. sign in

REVIEW 2 major objections 5 minor 1 cited by

TaigiSpeech supplies a real elderly Taiwanese Hokkien intent corpus and shows that models trained only on mined drama speech lose substantial accuracy on it.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-13 20:49 UTC pith:4OAJCZXP

load-bearing objection Solid elderly Taigi intent corpus plus a clean domain-mismatch result; mining is useful scaffolding, not the main claim. the 2 major comments →

arxiv 2603.21478 v2 pith:4OAJCZXP submitted 2026-03-23 cs.CL cs.LGeess.AS

TaigiSpeech: A Low-Resource Real-World Speech Intent Dataset and Preliminary Results with Scalable Data Mining In-the-Wild

classification cs.CL cs.LGeess.AS
keywords spoken language understandinglow-resource languageintent recognitionTaiwanese Hokkienelderly speechdata mininghome assistantspeech dataset
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper introduces TaigiSpeech, a 3,079-utterance spoken-intent dataset recorded from 21 older adults speaking Taiwanese Hokkien in healthcare and home-assistant scenarios. It argues that low-resource, primarily spoken languages still lack realistic evaluation resources, so systems built only from web-mined media cannot be trusted for elderly users. Two scalable mining methods are tested: keyword matching on Mandarin drama subtitles followed by LLM pseudo-labeling, and weakly supervised audio-visual retrieval. When models trained on the mined data are tested on the new elderly recordings, accuracy falls sharply, confirming a domain mismatch that the released dataset is designed to expose and help close.

Core claim

Models trained solely on intent segments mined from Taiwanese drama videos suffer large accuracy drops when evaluated on real-world elderly TaigiSpeech recordings, demonstrating a domain mismatch that makes a purpose-built elderly corpus necessary for low-resource spoken intent recognition.

What carries the argument

TaigiSpeech (the 8-intent elderly corpus) together with the two mining pipelines—keyword-match mining plus LLM pseudo-labeling via Mandarin subtitles, and PE-AV audio-visual retrieval—used to create the pre-training data against which the domain gap is measured.

Load-bearing premise

The claim rests on treating LLM pseudo-labels generated from Mandarin drama subtitle windows as accurate enough ground truth for both the mined training pool and the drama test set.

What would settle it

Independent human re-annotation of a random sample of the LLM-pseudo-labeled drama segments would show low agreement with the automatic labels, or models trained on those labels would perform near chance even on held-out drama speech.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Realistic elderly-speech benchmarks become a prerequisite before home-assistant or emergency systems can be trusted for low-resource languages.
  • Keyword-match mining with an intermediate high-resource language yields stronger subsequent adaptation than pure audio-visual mining under current multimodal encoders.
  • End-to-end SSL speech classifiers outperform cascaded foundation ASR-plus-LLM pipelines on this task while using far fewer parameters.
  • A few hundred in-domain elderly utterances are enough to recover most of the accuracy lost to domain mismatch.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Comparable domain gaps are likely for other primarily oral or unwritten languages when training data are mined from broadcast drama or television.
  • The persistent gap between lightweight MatchboxNet models and SSL models points to a need for better distillation if elderly intent systems must run on edge devices.
  • Mining from more naturalistic elderly speech sources, rather than drama alone, could shrink the mismatch without requiring additional manual labels.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. The paper introduces TaigiSpeech, a spoken-intent corpus of 3,079 utterances from 21 elderly Taiwanese Hokkien speakers (ages 54–78) covering eight intents (four emergency, four functional home-assistant commands). It further explores two in-the-wild mining pipelines—keyword matching on Mandarin drama subtitles followed by LLM pseudo-labeling, and weakly supervised audio-visual retrieval with PE-AV—and evaluates lightweight MatchboxNet and SSL models (HuBERT, WavLM). Table 7 shows a clear domain gap: models trained only on mined drama data drop substantially on the held-out TaigiSpeech test set (e.g., WavLM-large 92.36 % → 70.00 % under the 5-class keyword setting) and recover only after fine-tuning on a small speaker-independent Taigi train split. Cascaded foundation ASR+LLM baselines are also reported and underperform end-to-end SSL classifiers. The dataset is released under CC BY 4.0.

Significance. The work fills a genuine gap: existing emergency/smart-home SLU corpora are almost exclusively high-resource languages and rarely target elderly speakers of primarily spoken languages. TaigiSpeech supplies a realistic, scenario-driven elderly Hokkien benchmark with speaker-independent splits, bootstrap CIs, and public release. The measured domain-mismatch result (Table 7) is load-bearing and directly supports the claim that mined drama data alone are insufficient, thereby justifying the new corpus. The dual mining strategies, while preliminary, offer a practical template for other unwritten languages. Strengths include transparent statistics (Tables 1, 4, 6, 8), multi-model evaluation, and explicit acknowledgment of speaker-count and AV-scale limitations.

major comments (2)
  1. §4.1 and Table 6: LLM (Gemini-3) pseudo-labels serve both as training targets for keyword mining and as the sole ground truth for the Drama test set. No human validation or inter-annotator agreement is reported for these labels. While the primary TaigiSpeech evaluation is independent, the magnitude of the reported domain gap (Drama C=5 vs. Taigi C=5) and the claimed value of mined pre-training rest partly on the quality of these labels; a modest human audit of a stratified sample would strengthen the central claim.
  2. §5.1 / Table 8: The Audio-Visual mining experiment uses only a 28 k-clip subset already filtered by keyword queries rather than the full 7 000-hour drama pool, and collapses to a binary emergency/non-emergency task. The near-random Taigi (C=2) transfer results are therefore difficult to interpret as a fair test of the multimodal approach; either a larger unfiltered pool or an explicit statement that AV mining remains exploratory would better calibrate the claim of “scalable data mining.”
minor comments (5)
  1. Table 2 and §2.2: A few related emergency corpora (e.g., more recent PERS or multi-lingual AAL sets) could be cited for completeness; the current survey is already useful but slightly incomplete.
  2. Figure 6: Confusion matrices are informative; adding per-class F1 or a short discussion of why LIGHT ON/OFF remain confusable after adaptation would help readers.
  3. §3.2 / Supplementary A.2–A.3: Scenario and video prompts are generated by Gemini/Veo; a brief note on human filtering criteria would improve reproducibility of the elicitation protocol.
  4. Minor typos: “T his finding” (Conclusion), inconsistent capitalization of intent names across tables, and occasional missing spaces after periods.
  5. §6.4: Cascaded ASR+LLM results would be more informative if the authors stated whether any prompt engineering or few-shot examples were used for the Qwen3-8B intent classifier.

Circularity Check

0 steps flagged

No significant circularity: empirical dataset release and transfer experiments with independent human-elicited TaigiSpeech evaluation.

full rationale

The paper is a dataset contribution plus preliminary mining and baseline experiments. Its central empirical claim (domain mismatch) is the measured accuracy drop when models trained on drama-mined data are evaluated on the fixed, independently collected TaigiSpeech elderly test set of 960 human-elicited utterances (Table 7: e.g. WavLM-large 92.36 % Drama C=5 o 70.00 % Taigi C=5). Drama labels are LLM pseudo-labels (Gemini-3) and thus measure consistency with that model, which the paper itself states; this does not make the Taigi numbers circular, because Taigi labels come from the scenario-driven recording protocol, not from the LLM or from any fitted parameter of the mining pipeline. No equation, uniqueness theorem, or ansatz is imported via self-citation to force a result; no parameter is fitted to a quantity that is then re-presented as a prediction. The work is therefore self-contained against its own external benchmark (the released TaigiSpeech recordings). Score 0 is appropriate.

Axiom & Free-Parameter Ledger

2 free parameters · 4 axioms · 0 invented entities

As an empirical dataset-and-benchmark paper the load-bearing premises are standard speech-processing assumptions plus a few domain-specific choices about labeling and mining; no new physical entities or free parameters are fitted to produce the central domain-gap claim.

free parameters (2)
  • top-k retrieval size for AV mining
    Authors select top 2 000 emergency and bottom 2 000 non-emergency clips from the 28 k pool; the threshold is chosen by hand and affects the AV training distribution.
  • number of keywords per intent
    At least 10 Mandarin keywords per intent plus ~100 daily-life keywords are manually constructed; coverage directly determines the mined set size.
axioms (4)
  • domain assumption Mandarin subtitles of Taiwanese drama are sufficiently faithful translations for keyword matching and LLM intent verification
    Invoked throughout Section 4.1; without it the keyword-mining pipeline has no pivot language.
  • domain assumption LLM (Gemini-3) pseudo-labels on contextualized subtitle windows constitute usable ground truth for both training and Drama test evaluation
    Stated in Section 5.1 and footnote 2; the Drama accuracy numbers rest on this.
  • domain assumption Imagined scenario prompts plus optional silent videos elicit speech whose acoustic and lexical properties are representative of real emergency and home-assistant use
    Section 3.2; the claim that TaigiSpeech is a realistic benchmark depends on this ecological-validity assumption.
  • standard math Standard self-supervised speech models (HuBERT, WavLM) and MatchboxNet architectures transfer to Taiwanese Hokkien after fine-tuning
    Implicit in the experimental design of Section 5; no new architecture is claimed.

pith-pipeline@v1.1.0-grok45 · 25358 in / 2519 out tokens · 29962 ms · 2026-07-13T20:49:28.528790+00:00 · methodology

0 comments
read the original abstract

Speech technologies have advanced rapidly and serve diverse populations worldwide. However, many languages remain underrepresented due to limited resources. In this paper, we introduce \textbf{TaigiSpeech}, a real-world speech intent dataset in Taiwanese Taigi (aka Taiwanese Hokkien/Southern Min), which is a low-resource and primarily spoken language. The dataset is collected from older adults, comprising 21 speakers with a total of 3k utterances. It is designed for practical intent detection scenarios, including healthcare and home assistant applications. To address the scarcity of labeled data, we explore two data mining strategies with two levels of supervision: keyword match data mining with LLM pseudo labeling via an intermediate language and an audio-visual framework that leverages multimodal cues with minimal textual supervision. This design enables scalable dataset construction for low-resource and unwritten spoken languages. TaigiSpeech will be released under the CC BY 4.0 license to facilitate broad adoption and research on low-resource and unwritten languages. The project website and the dataset can be found on https://kwchang.org/taigispeech.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Walking Through Uncertainty: An Empirical Study of Uncertainty Estimation for Audio-Aware Large Language Models

    eess.AS 2026-04 unverdicted novelty 7.0

    Semantic-level and verification-based uncertainty methods outperform token-level baselines for audio reasoning in ALLMs, but their relative performance on hallucination and unanswerable-question benchmarks is model- a...

Reference graph

Works this paper leans on

96 extracted references · 2 canonical work pages · cited by 1 Pith paper

  1. [1]

    How- ever, many of them are not supported by commercial AI ser- vices

    Introduction More than 7,000 languages are spoken worldwide [1]. How- ever, many of them are not supported by commercial AI ser- vices. A substantial portion of these languages lack a univer- sally standardized writing system, leading to inconsistent or- thographic resources, or are entirely unwritten [2]. As a result, they remain underrepresented in mode...

  2. [2]

    SOS CALL Calls for emergency help

  3. [3]

    Reports breathing difficulty

    BREATH EMERG. Reports breathing difficulty

  4. [4]

    FALL HELP Indicates a fall and requests help

  5. [5]

    Non-Emergency Intents / Functional Commands

    PAIN GENERAL Reports physical pain or discomfort. Non-Emergency Intents / Functional Commands

  6. [6]

    CALL CONTACT Requests to contact a person

  7. [7]

    LIGHT ON Requests to turn on the lights

  8. [8]

    LIGHT OFF Requests to turn off the lights

  9. [9]

    CANCEL Cancels a previously triggered alert. Total Utterances: 3,079 ������������ ������� ������������� ������ �� ��� ��� ��� ��� ���� ����� ������� ���� ������� ���� ������� ���� ������� �������� � �������� Figure 1:Primary language use among residents in Taiwan. El- ders (aged 65+) have 64.9% reporting Taiwanese as their pri- mary language. Data source:...

  10. [10]

    Related Works 2.1. Spoken Language Understanding Spoken language understanding (SLU) aims to extract semantic meaning from speech signals and serves as a core component in voice assistants and spoken dialogue systems [43]. A cen- tral subtask in SLU isintent recognition, which identifies the user’s underlying intention from spoken utterances and triggers ...

  11. [11]

    Participant Registration and Recording Interface Figure 2:Recording app interface

    TaigiSpeech In this section, we describe the data collection procedure, partic- ipant recruitment process, scenario design, and basic statistics of the TaigiSpeech dataset 3.1. Participant Registration and Recording Interface Figure 2:Recording app interface. The English translations are shown for illustrative purposes in this paper. All recordings were c...

  12. [12]

    SOS CALL 387 7.78 3.28 50.16

  13. [13]

    FALL HELP 385 7.77 3.73 49.88

  14. [14]

    BREATH EMERG 385 8.22 3.30 52.71

  15. [15]

    PAIN GENERAL 385 7.91 3.73 50.79

  16. [16]

    CALL CONTACT 385 7.04 5.30 45.16

  17. [17]

    LIGHT ON 384 5.95 2.81 38.08

  18. [18]

    LIGHT OFF 384 5.67 2.35 36.28

  19. [19]

    Ahh! Fire! It’s on fire! Help!

    CANCEL ALERT 384 6.85 3.17 43.85 Overall 3,079 7.15 3.66 366.91 Figure 3:Age distribution of speakers in the TaigiSpeech cor- pus. Our study focuses on elderly participants. Figure 4:The distribution of ambient noise level (dB). a mouse or touchscreen. 3.2. Scenario Design TaigiSpeech targets 8 intents, covering both emergency and non-emergency functional...

  20. [20]

    help”,“call the ambulance

    Data Mining in-the-wild For many low-resource and unwritten languages, reliable ASR systems are often unavailable. Also, in most cases, speech data primarily exist in web-based sources (e.g., online videos), where recordings are noisy and loosely structured. These con- straints make large-scale supervised data collection impracti- cal, motivating the deve...

  21. [21]

    Experimental Setup For pseudo-labeling in the Keyword Mining approach, we em- ploy Gemini-3 to generate intermediate annotations

    Experiments 5.1. Experimental Setup For pseudo-labeling in the Keyword Mining approach, we em- ploy Gemini-3 to generate intermediate annotations. For audio- visual data mining, we adopt the PE-A V large [58] to ex- tract multimodal representations and identify candidate seg- ments without relying on paired textual supervision. Note that the PE-A V model ...

  22. [22]

    We summarize the key observations in the following sections

    Preliminary Results The main results are presented in Table 7. We summarize the key observations in the following sections. 6.1. Domain mismatch A clear performance discrepancy is observed between evalua- tions on the drama dataset and on the real-world TaigiSpeech recordings. This trend consistently appears in both settings: Drama (C=5)�Taigi (C=5)under ...

  23. [23]

    Conclusion and Future Work In this work, we introduce TaigiSpeech, a Taiwanese Hokkien spoken intent dataset designed for elderly users in home- assistant and healthcare scenarios. The dataset com- prises 8 intent categories, including four emergency intents (SOS CALL, BREATH EMERGENCY, FALL HELP, and PAIN GENERAL) and four non-emergency functional intent...

  24. [24]

    This work was supported in part by the National Science and Technology Council (NSTC), Taiwan, under Grant No

    Acknowledgments The authors would like to express their sincere gratitude to the Taiwanese Language and Culture Club of Keelung Community University for their invaluable support and assistance, and to the Taipei City Nangang Social Welfare Center for their advice and suggestions. This work was supported in part by the National Science and Technology Counc...

  25. [25]

    All authors take full respon- sibility for the content of this paper, including its accuracy, orig- inality, and conclusions

    Generative AI Use Disclosure Generative AI tools were used solely for language editing and minor polishing of the manuscript. All authors take full respon- sibility for the content of this paper, including its accuracy, orig- inality, and conclusions. No generative AI tools were used to produce substantial portions of the scientific content, analysis, or results

  26. [26]

    Languages of the World,

    W. R. Leben, “Languages of the World,” 02 2018. [Online]. Avail- able: https://oxfordre.com/linguistics/view/10.1093/acrefore/ 9780199384655.001.0001/acrefore-9780199384655-e-349

  27. [27]

    D. M. Eberhard, G. F. Simons, and A. J. Robinson, Eds.,Ethno- logue: Languages of the World, 29th ed. SIL Global, 2026

  28. [28]

    Kwok,Southern Min: Comparative phonology and sub- grouping

    B.-C. Kwok,Southern Min: Comparative phonology and sub- grouping. Routledge, 2018

  29. [29]

    The influence of Southern Min on the Mandarin of Taiwan,

    C. C. Kubler, “The influence of Southern Min on the Mandarin of Taiwan,”Anthropological Linguistics, pp. 156–176, 1985

  30. [30]

    Language Usage for the Resident Nationals Aged 6 Years and Over - 2020 Population and Housing Census,

    A. Directorate General of Budget and T. Statistics (DGBAS), “Language Usage for the Resident Nationals Aged 6 Years and Over - 2020 Population and Housing Census,” 2021

  31. [31]

    Challenges in real-life emotion annotation and machine learning based detection,

    L. Devillers, L. Vidrascu, and L. Lamel, “Challenges in real-life emotion annotation and machine learning based detection,”Neu- ral Networks, vol. 18, no. 4, pp. 407–422, 2005, emotion and Brain

  32. [32]

    A Multimodal Corpus Recorded in a Health Smart Home,

    A. Fleury, M. Vacher, F. Portet, P. Chahuara, and N. Noury, “A Multimodal Corpus Recorded in a Health Smart Home,” inLREC 2010, The International Conference on Language Resources and Evaluation, Valetta, Malta, May 2010, pp. 99–105

  33. [33]

    The CARES corpus: a database of older adult actor simulated emergency dialogue for developing a personal emergency response system,

    V . Young and A. Mihailidis, “The CARES corpus: a database of older adult actor simulated emergency dialogue for developing a personal emergency response system,”International Journal of Speech Technology, vol. 16, no. 1, pp. 55–73, 2013

  34. [34]

    A French corpus of audio and multimodal interactions in a health smart home,

    A. Fleury, M. Vacher, F. Portet, P. Chahuara, and N. Noury, “A French corpus of audio and multimodal interactions in a health smart home,”Journal on Multimodal User Interfaces, vol. 7, no. 1, pp. 93–109, 2013

  35. [35]

    The Sweet-Home speech and multimodal corpus for home automation interaction,

    M. Vacher, B. Lecouteux, P. Chahuara, F. Portet, B. Meillon, and N. Bonnefond, “The Sweet-Home speech and multimodal corpus for home automation interaction,” inLREC 2014, Reykjavik, Ice- land, May 2014, pp. 4499–4506

  36. [36]

    An integrated system for voice command recognition and emergency detection based on audio signals,

    E. Principi, S. Squartini, R. Bonfigli, G. Ferroni, and F. Pi- azza, “An integrated system for voice command recognition and emergency detection based on audio signals,”Expert Syst. Appl., vol. 42, no. 13, p. 5668–5683, Aug. 2015

  37. [37]

    Speech commands: A dataset for limited-vocabulary speech recognition,

    P. Warden, “Speech commands: A dataset for limited-vocabulary speech recognition,”arXiv preprint arXiv:1804.03209, 2018

  38. [38]

    Speech Model Pre-Training for End-to-End Spoken Language Understanding,

    L. Lugosch, M. Ravanelli, P. Ignoto, V . S. Tomar, and Y . Bengio, “Speech Model Pre-Training for End-to-End Spoken Language Understanding,” inInterspeech 2019, 2019, pp. 814–818

  39. [39]

    Context-Aware V oice-Based Interaction in Smart Home - V ocADom@A4H Corpus Collection and Empirical Assessment of Its Useful- ness,

    F. Portet, S. Caffiau, F. Ringeval, M. Vacher, N. Bonne- fond, S. Rossato, B. Lecouteux, and T. Desot, “Context-Aware V oice-Based Interaction in Smart Home - V ocADom@A4H Corpus Collection and Empirical Assessment of Its Useful- ness,” in2019 IEEE Intl Conf on Dependable, Autonomic and Secure Computing, Intl Conf on Pervasive Intelligence and Computing, ...

  40. [40]

    SLURP: A spoken language understanding resource package,

    E. Bastianelli, A. Vanzo, P. Swietojanski, and V . Rieser, “SLURP: A spoken language understanding resource package,” inProceed- ings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2020, pp. 7252–7262

  41. [41]

    Learning Asr-Robust Contex- tualized Embeddings for Spoken Language Understanding,

    C.-W. Huang and Y .-N. Chen, “Learning Asr-Robust Contex- tualized Embeddings for Spoken Language Understanding,” in ICASSP 2020 - 2020 IEEE International Conference on Acous- tics, Speech and Signal Processing (ICASSP), 2020, pp. 8009– 8013

  42. [42]

    EMSAssist: An End-to-End Mobile V oice Assis- tant at the Edge for Emergency Medical Services,

    L. Jin, T. Liu, A. Haroon, R. Stoleru, M. Middleton, Z. Zhu, and T. Chaspari, “EMSAssist: An End-to-End Mobile V oice Assis- tant at the Edge for Emergency Medical Services,” ser. MobiSys ’23. New York, NY , USA: Association for Computing Machin- ery, 2023, p. 275–288

  43. [43]

    VN-SLU: A Vietnamese Spoken Language Un- derstanding Dataset,

    T. Tran, K. Le, N. D. Nguyen, M. Vu, H. Ngo, W. Park, and T. T. T. Nguyen, “VN-SLU: A Vietnamese Spoken Language Un- derstanding Dataset,” inInterspeech 2024, 2024, pp. 1335–1339

  44. [44]

    Enhancing V oice Wake-Up for Dysarthria: Man- darin Dysarthria Speech Corpus Release and Customized System Design,

    M. Gao, H. Chen, J. Du, X. Xu, H. Guo, H. Bu, J. Yang, M. Li, and C.-H. Lee, “Enhancing V oice Wake-Up for Dysarthria: Man- darin Dysarthria Speech Corpus Release and Customized System Design,” inInterspeech 2024, 2024, pp. 2465–2469

  45. [45]

    Examining the role of living alone and loneliness in predicting health-related quality of life: results from the Healthy Aging Longitudinal Study in Taiwan (HALST),

    H.-Y . Tseng, C.-Y . Lee, C.-S. Wu, I.-C. Wu, H.-Y . Chang, C.- C. Hsu, and C. A. Hsiung, “Examining the role of living alone and loneliness in predicting health-related quality of life: results from the Healthy Aging Longitudinal Study in Taiwan (HALST),” Quality of Life Research, vol. 33, no. 4, pp. 1015–1028, 2024

  46. [46]

    Spatial Distribution, Characterization, and Policy Opportunities for Taiwan’s Solo Elderly: A Big Data Approach,

    Y . Lin, W. Chen, H. Yu, and I. Cho, “Spatial Distribution, Characterization, and Policy Opportunities for Taiwan’s Solo Elderly: A Big Data Approach,” inProceedings of DRS2024: Boston, C. Gray, E. Ciliotta Chehade, P. Hekkert, L. Forlano, P. Ciuccarelli, and P. Lloyd, Eds., Boston, USA, Jun. 2024. [Online]. Available: https://doi.org/10.21606/drs.2024.366

  47. [47]

    Link Between Social Relationships and Vulnerability Among Community-Dwelling Older Adults Living Alone in Taiwan,

    K.-M. Chen, S.-T. Wang, S.-R. Chao, and K. Kasirisir, “Link Between Social Relationships and Vulnerability Among Community-Dwelling Older Adults Living Alone in Taiwan,” Asian Social Work and Policy Review, vol. 19, no. 1, p. e70000, 2025, e70000 ASWP-May-2024-0051.R1. [Online]. Available: https://onlinelibrary.wiley.com/doi/abs/10.1111/aswp.70000

  48. [48]

    Digital health literacy and its deter- minants among community dwelling elderly people in Taiwan,

    T. T. Tran, P. W. Chang, J.-M. Yang, T.-H. Chen, C.-T. Su, D. Levin-Zamir, O. Baron-Epel, E. Neter, S. F. Tsai, B. Lo, T. V . Duong, and S.-H. Yang, “Digital health literacy and its deter- minants among community dwelling elderly people in Taiwan,” DIGITAL HEALTH, vol. 10, p. 20552076241278926, 2024

  49. [49]

    A smartphone model for post-acute care decreases all- cause mortality with improved left ventricular ejection fraction in patients hospitalized with heart failure in Taiwan,

    S.-C. Weng, W.-W. Lin, J.-L. Huang, C.-Y . Chao, C.-Y . Hsu, and S.-Y . Lin, “A smartphone model for post-acute care decreases all- cause mortality with improved left ventricular ejection fraction in patients hospitalized with heart failure in Taiwan,”Maturitas, vol. 197, p. 108269, 2025. [Online]. Available: https://www. sciencedirect.com/science/article...

  50. [50]

    From command to care: a scoping re- view on utilization of smart speakers by patients and providers,

    R. Saripalle and R. Patel, “From command to care: a scoping re- view on utilization of smart speakers by patients and providers,” Mayo Clinic Proceedings: Digital Health, vol. 2, no. 2, pp. 207– 220, 2024

  51. [51]

    Falls in older people: epidemiology, risk fac- tors and strategies for prevention,

    L. Z. Rubenstein, “Falls in older people: epidemiology, risk fac- tors and strategies for prevention,”Age and ageing, vol. 35, no. suppl 2, pp. ii37–ii41, 2006

  52. [52]

    The Growing Burden of Fall-Related Injuries Among Older Adults: A Seven-Year Study from a Tertiary Medical Center in Taiwan,

    Y .-C. Tsai, S.-Y . Chen, Y .-M. Chen, E. P.-C. Huang, and F.-P. Lu, “The Growing Burden of Fall-Related Injuries Among Older Adults: A Seven-Year Study from a Tertiary Medical Center in Taiwan,”Annals of Geriatric Medicine and Research,

  53. [53]

    Available: http://www.e-agmr.org/journal/view

    [Online]. Available: http://www.e-agmr.org/journal/view. php?number=1240

  54. [54]

    Concept of the term long lie: a scoping review,

    J. Kubitza, I. T. Schneider, and B. Reuschenbach, “Concept of the term long lie: a scoping review,”European Review of Aging and Physical Activity, vol. 20, no. 1, p. 16, 2023

  55. [55]

    SmartLohas: A Smart Assistive System for Elder People,

    Y .-T. Tsai, C.-L. Fan, C.-L. Lo, and S.-H. Huang, “SmartLohas: A Smart Assistive System for Elder People,” in2017 14th Inter- national Symposium on Pervasive Systems, Algorithms and Net- works & 2017 11th International Conference on Frontier of Com- puter Science and Technology & 2017 Third International Sym- posium of Creative Computing (ISPAN-FCST-ISCC), 2017

  56. [56]

    V oice Health As- sistant for Elderly with Multi-Language and Offline Support,

    R. Chithra, R. Ackshaya, and T. Dhinakaran, “V oice Health As- sistant for Elderly with Multi-Language and Offline Support,” in 2025 10th International Conference on Smart Structures and Sys- tems (ICSSS), 2025, pp. 1–7

  57. [57]

    Development of an automated speech recognition interface for personal emer- gency response systems,

    M. Hamill, V . Young, J. Boger, and A. Mihailidis, “Development of an automated speech recognition interface for personal emer- gency response systems,”Journal of NeuroEngineering and Re- habilitation, vol. 6, no. 1, p. 26, 2009

  58. [58]

    Evaluation of a context-aware voice interface for ambient assisted living: qualitative user study vs. quantitative system evaluation,

    M. Vacher, S. Caffiau, F. Portet, B. Meillon, C. Roux, E. Elias, B. Lecouteux, and P. Chahuara, “Evaluation of a context-aware voice interface for ambient assisted living: qualitative user study vs. quantitative system evaluation,”ACM Transactions on Acces- sible Computing (TACCESS), vol. 7, no. 2, pp. 1–36, 2015

  59. [59]

    Making emergency calls more accessible to older adults through a hands-free speech interface in the house,

    M. Vacher, F. Aman, S. Rossato, F. Portet, and B. Lecouteux, “Making emergency calls more accessible to older adults through a hands-free speech interface in the house,”ACM Transactions on Accessible Computing (TACCESS), vol. 12, no. 2, pp. 1–25, 2019

  60. [60]

    V oice-controlled intelligent personal assistants to sup- port aging in place,

    K. O’Brien, A. Liggett, V . Ramirez-Zohfeld, P. Sunkara, and L. A. Lindquist, “V oice-controlled intelligent personal assistants to sup- port aging in place,”Journal of the American Geriatrics Society, vol. 68, no. 1, pp. 176–179, 2020

  61. [61]

    Speech-massive: A multilingual speech dataset for slu and be- yond,

    B. Lee, I. Calapodescu, M. Gaido, M. Negri, and L. Besacier, “Speech-massive: A multilingual speech dataset for slu and be- yond,” inProc. Interspeech 2024, 2024, pp. 817–821

  62. [62]

    Catslu: The 1st chinese audio-textual spoken language understanding chal- lenge,

    S. Zhu, Z. Zhao, T. Zhao, C. Zong, and K. Yu, “Catslu: The 1st chinese audio-textual spoken language understanding chal- lenge,” in2019 International Conference on Multimodal Interac- tion, 2019, pp. 521–525

  63. [63]

    Su ´ısiann dataset,

    `I-thuˆan, “Su ´ısiann dataset,” https://suisiann-dataset.ithuan.tw/, 2019

  64. [64]

    Taiwanese across taiwan corpus and its applications,

    Y .-F. Liao, J. S. Tsay, P. Kang, H.-L. Khoo, L.-K. Tan, L.- C. Chang, U.-G. Iunn, H.-L. Su, T.-G. Thiann, H.-K. Tiun et al., “Taiwanese across taiwan corpus and its applications,” in 2022 25th Conference of the Oriental COCOSDA International Committee for the Co-ordination and Standardisation of Speech Databases and Assessment Techniques (O-COCOSDA). IEE...

  65. [65]

    Com- mon voice: A massively-multilingual speech corpus,

    R. Ardila, M. Branson, K. Davis, M. Kohler, J. Meyer, M. Hen- retty, R. Morais, L. Saunders, F. Tyers, and G. Weber, “Com- mon voice: A massively-multilingual speech corpus,” inProceed- ings of the twelfth language resources and evaluation conference, 2020, pp. 4218–4222

  66. [66]

    Ml-superb: Multilingual speech universal performance benchmark,

    J. Shi, D. Berrebbi, W. Chen, E.-P. Hu, W.-P. Huang, H.-L. Chung, X. Chang, S.-W. Li, A. Mohamed, H.-y. Leeet al., “Ml-superb: Multilingual speech universal performance benchmark,” inProc. Interspeech 2023, 2023, pp. 884–888

  67. [67]

    Tg-asr: Translation-guided learning with parallel gated cross attention for low-resource automatic speech recognition,

    C.-Y . Yang, C.-C. Wang, L.-W. Chen, H.-S. Lee, H.-M. Wang, and B. Chen, “Tg-asr: Translation-guided learning with parallel gated cross attention for low-resource automatic speech recognition,” arXiv preprint arXiv:2602.22039, 2026

  68. [68]

    Evaluating self-supervised speech models on a taiwanese hokkien corpus,

    Y .-H. Chou, K. Chang, M.-J. Wu, W. Ou, A. W.-H. Bi, C. Yang, B. Y . Chen, R.-W. Pai, P.-Y . Yeh, J.-P. Chianget al., “Evaluating self-supervised speech models on a taiwanese hokkien corpus,” in2023 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU). IEEE, 2023, pp. 1–7

  69. [69]

    Espnet-slu: Advancing spoken language understanding through espnet,

    S. Arora, S. Dalmia, P. Denisov, X. Chang, Y . Ueda, Y . Peng, Y . Zhang, S. Kumar, K. Ganesan, B. Yanet al., “Espnet-slu: Advancing spoken language understanding through espnet,” in ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2022

  70. [70]

    The atis spoken language systems pilot corpus,

    C. T. Hemphill, J. J. Godfrey, and G. R. Doddington, “The atis spoken language systems pilot corpus,” inSpeech and Natural Language: Proceedings of a Workshop Held at Hidden Valley, Pennsylvania, June 24-27, 1990, 1990

  71. [71]

    Slue: New benchmark tasks for spoken language un- derstanding evaluation on natural speech,

    S. Shon, A. Pasad, F. Wu, P. Brusco, Y . Artzi, K. Livescu, and K. J. Han, “Slue: New benchmark tasks for spoken language un- derstanding evaluation on natural speech,” inICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2022, pp. 7927–7931

  72. [72]

    Self-supervised speech representation learning: A review,

    A. Mohamed, H.-y. Lee, L. Borgholt, J. D. Havtorn, J. Edin, C. Igel, K. Kirchhoff, S.-W. Li, K. Livescu, L. Maaløeet al., “Self-supervised speech representation learning: A review,”IEEE Journal of Selected Topics in Signal Processing, vol. 16, no. 6, pp. 1179–1210, 2022

  73. [73]

    Hubert: Self-supervised speech represen- tation learning by masked prediction of hidden units,

    W.-N. Hsu, B. Bolte, Y .-H. H. Tsai, K. Lakhotia, R. Salakhutdi- nov, and A. Mohamed, “Hubert: Self-supervised speech represen- tation learning by masked prediction of hidden units,”IEEE/ACM transactions on audio, speech, and language processing, vol. 29, pp. 3451–3460, 2021

  74. [74]

    Wavlm: Large-scale self- supervised pre-training for full stack speech processing,

    S. Chen, C. Wang, Z. Chen, Y . Wu, S. Liu, Z. Chen, J. Li, N. Kanda, T. Yoshioka, X. Xiaoet al., “Wavlm: Large-scale self- supervised pre-training for full stack speech processing,”IEEE Journal of Selected Topics in Signal Processing, vol. 16, no. 6, pp. 1505–1518, 2022

  75. [75]

    A large-scale evaluation of speech foundation models,

    S.-w. Yang, H.-J. Chang, Z. Huang, A. T. Liu, C.-I. Lai, H. Wu, J. Shi, X. Chang, H.-S. Tsai, W.-C. Huanget al., “A large-scale evaluation of speech foundation models,”IEEE/ACM Transac- tions on Audio, Speech, and Language Processing, vol. 32, 2024

  76. [76]

    Robust speech recognition via large-scale weak supervision,

    A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever, “Robust speech recognition via large-scale weak supervision,” inInternational conference on machine learning. PMLR, 2023, pp. 28 492–28 518

  77. [77]

    Qwen3-asr technical report,

    X. Shi, X. Wang, Z. Guo, Y . Wang, P. Zhang, X. Zhang, Z. Guo, H. Hao, Y . Xi, B. Yanget al., “Qwen3-asr technical report,”arXiv preprint arXiv:2601.21337, 2026

  78. [78]

    An official American Thoracic Society statement: update on the mechanisms, assessment, and management of dyspnea,

    M. B. Parshall, R. M. Schwartzstein, L. Adams, R. B. Banzett, H. L. Manning, J. Bourbeau, P. M. Calverley, A. G. Gift, A. Harver, S. C. Lareau, D. A. Mahler, P. M. Meek, D. E. O’Donnell, and American Thoracic Society Committee on Dyspnea, “An official American Thoracic Society statement: update on the mechanisms, assessment, and management of dyspnea,”Ame...

  79. [79]

    Gemini: A Family of Highly Capable Multimodal Models,

    G. Team, “Gemini: A Family of Highly Capable Multimodal Models,” 2025. [Online]. Available: https://arxiv.org/abs/2312. 11805

  80. [80]

    Veo 3: Advancing High-Fidelity Generative Video Modeling,

    G. Team and D. Team, “Veo 3: Advancing High-Fidelity Generative Video Modeling,” Google DeepMind, Tech. Rep., 2025, accessed: 2026-02-23. [Online]. Available: https://storage. googleapis.com/deepmind-media/veo/Veo-3-Tech-Report.pdf

Showing first 80 references.