{"id":"1088b972-5f1e-4d90-91dd-bc2cafcc8cac","arxiv_id":"1908.08971","paper_version":2,"verdict":"UNVERDICTED","confidence":"HIGH","novelty_score":2.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"A community event report describing how a diverse group discussed using neural-network ASR to transcribe endangered languages; no new technical findings are presented.","lead":"This paper is a report on a 2018 retreat where linguists, computer scientists, and native speakers discussed using automatic speech recognition for endangered languages. It describes the event and the promise of tools like Persephone, but presents no new experimental results or technical findings.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Feasibility claim rests on two tonal-language case studies; generalization to the retreat's severely under-resourced languages is untested.","rationale":"The reader marked the paper UNVERDICTED because it is a meeting report with no new technical results and a feasibility claim lacking supporting evidence. My stress-test identifies the same load-bearing assumption: the generalization from two tonal-language case studies to the broad set of endangered languages represented at the retreat. I cannot find a reason to move away from UNVERDICTED: the paper's scientific contribution is nil, but its value as community documentation is not compromised by the concern. The concern is real but does not change the verdict; it reinforces the reader's assessment. I am not manufacturing a new issue, and I do not see an internal inconsistency that would warrant REJECT. The paper explicitly frames the retreat as a discussion, not a proof, so the overgeneralization is a weakness in the rhetoric, not a fatal flaw. Thus, UNCHANGED is appropriate.","tokens_in":3664,"tokens_out":5563,"duration_ms":53137,"concrete_test":"Train Persephone on a non-tonal language from the retreat with very limited data, e.g., Bardi or Mocho, using the same protocol as Adams et al. 2018 (phonemic transcription, 30 minutes per speaker if available, or all available data), and measure phoneme error rate (PER). Compare against the reported ~20% PER for Yongning Na and the qualitative Chatino result. If PER for a non-tonal or highly endangered language exceeds e.g. 30%, the generalization is unsupported and the 'no theoretical obstacle' claim needs qualification. Alternatively, inspect Adams et al. 2018 for whether evaluation is speaker-independent; if not, re-evaluate with held-out speakers.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central feasibility claim (Section 2: 'there is no theoretical obstacle... largely a matter of having a professional software engineer develop the tool') is supported only by Persephone results on Yongning Na (20% error) and Chatino (described as 'performed well', without metrics) from Adams et al. 2018. The retreat explicitly includes languages with 4-5 speakers (Bardi, Djambarrpuyngu) or one speaker (Mocho), from non-tonal families (Pama-Nyungan, Mayan, etc.). The cited results are for tonal languages with specific experimental conditions (phonemic transcription, likely a single speaker, 30 minutes of data per speaker). There is no analysis or evidence that these results transfer to languages with extreme data sparsity, different phonologies, or complex morphology. Consequently, the claim that only software engineering stands between current ASR and practical deployment for endangered languages is an overgeneralization from two case studies. This is load-bearing because the paper's purpose is to motivate deployment of ASR; if the underlying technology fails on many retreat languages, the conclusion would be undermined. However, the paper is a meeting report and does not provide a rigorous scientific argument, so this weakness does not change the already-UNVERDICTED status.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper is a meeting report on an August 2018 retreat in Quechee, Vermont, that brought together computer scientists, linguists, native speakers, and language activists to discuss the use of automatic speech recognition (ASR), especially neural-network methods, for transcribing endangered languages. It describes two earlier collaborations, on Chatino (led by the first author, Hilaria Cruz) and on Yongning Na (with Alexis Michaud and Oliver Adams), which used the open-source Persephone toolkit and reported promising results including a 20% error rate on Na and reasonable accuracy on Chatino with as little as thirty minutes of data per speaker. The paper's central conclusion, stated in Section 2, is that 'there is no theoretical obstacle' to building an interface that would let a linguist upload speech and transcriptions for model training, and that the main barrier is essentially the effort of a professional software engineer. The manuscript also emphasizes the importance of involving native speakers in ASR development and recounts the personal motivations behind the first author's work. No new experiments, data, or quantitative analyses are reported.","tokens_in":3894,"tokens_out":4622,"duration_ms":47990,"significance":"If the feasibility claim is accepted, the paper identifies a concrete and relatively accessible path toward integrating ASR into language documentation workflows, with the potential to increase transcription consistency and free linguists to act as editors rather than transcribers. The paper's emphasis on native-speaker leadership and its detailed account of community interactions are genuinely valuable for a field where technology development has often been driven by outsiders. The paper also honestly credits prior open-source work (Persephone, the Chatino corpus) and names specific people and projects. However, the load-bearing feasibility claim is supported only by anecdotal experience with two tonal languages and is not backed by any new analysis, error bars, or systematic review of the conditions under which ASR works for severely under-resourced languages. As a meeting report, the manuscript is informative, but as a scientific argument for the 'no theoretical obstacle' assertion, it is under-supported.","major_comments":[{"comment":"The assertion that \"there is no theoretical obstacle\" to creating an upload-and-train ASR interface is overgeneralized from evidence on exactly two tonal languages (Yongning Na and Chatino), yet the retreat itself included languages from non-tonal families (Pama-Nyungan, Nyulnyulan, Mayan) with as few as one to five remaining speakers (Bardi, Djambarrpuyngu, Mocho). The paper provides no analysis explaining why results from a tonal language with 30 minutes of per-speaker data should transfer to these extremely sparse, typologically different settings, so the conclusion in Section 2 is not supported by the evidence presented.","section":"Section 2"},{"comment":"The reported \"20% error rate\" for Yongning Na and \"reasonable accuracy\" for Chatino with thirty minutes of data are cited to Adams et al. (2018) but are not contextualized with the evaluation metric (phoneme error rate versus word error rate), the number of test speakers, the size or composition of the test set, or the overlap between training and test conditions. Without these details, the results cannot be independently assessed and therefore cannot carry the weight of the paper's central feasibility claim in Section 2.","section":"Section 1"},{"comment":"The paper claims that the remaining barrier is \"largely a matter of having a professional software engineer develop the tool,\" but it also states that Persephone is currently \"only accessible to computer scientists\" and has only preliminary support for ELAN files. This internal tension suggests that the gap is not merely engineering but also includes model customization to the linguistic and orthographic properties of each language, user training, community-specific data governance, and the need for error-analysis tools, none of which are addressed; the paper therefore understates the obstacles between the cited laboratory successes and practical deployment.","section":"Section 2"}],"minor_comments":[{"comment":"The narrative inconsistently alternates between \"the first author\" and the first-person pronoun \"I\"; for a paper with two authors, the intended referent should be made unambiguous, for example by using \"the first author\" throughout or by naming the author at each transition.","section":"Entire manuscript"},{"comment":"The in-text citation \"Cavar et al. 2016\" does not match the diacritic used in the reference list entry \"Ćavar et al. 2016\"; please standardize the spelling so that the citation and reference list agree.","section":"References"},{"comment":"The phrase the \"event was a resounding success\" is subjective; since the paper otherwise describes concrete activities, a short list of tangible outcomes (such as the plan for a web API, an identified set of pilot languages, or a follow-up meeting) would give readers a more informative basis for evaluating the retreat's productivity.","section":"Section 2"},{"comment":"The phrase \"the bottle neck\" should be written as the single word \"bottleneck\" for consistency with standard terminology, and the abstract would benefit from explicitly stating that the paper is a meeting report rather than a research article.","section":"Abstract and Section 1"}],"recommendation":"major_revision","confidential_remarks":"This manuscript is more of a community-report/position piece than a research article, which is acceptable if the venue welcomes such contributions. The main revision needed is to reframe the 'no theoretical obstacle' claim as a hypothesis or a call for research, with a candid discussion of the limited empirical basis and the typological and data-sparsity conditions under which the cited results were obtained. Alternatively, the paper could be repositioned as a proposal for a community-driven tool-building effort, which would make the feasibility claim a motivating goal rather than an established fact. The self-citation to Adams et al. (2018) is appropriate, but the current paper does not add new evidence, so the claim should be scaled accordingly."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know this paper is what it says it is: a report on a 2018 retreat that brought native speakers, field linguists, and NLP people together around ASR for endangered languages. There is no new method, no new data, no experiment. The one genuinely new thing is the meeting itself — the first of its kind convened by a native speaker — and the paper documents that event faithfully. As community record and advocacy, it has real value: it names the bottleneck, gives concrete numbers from a native speaker's transcription workload, and records the resources and speakers present. I believe the authors are not overselling what happened at the retreat; the tone is earnest and specific.\n\nThe technical core is thin. The load-bearing sentence is in Section 2: 'there is no theoretical obstacle... largely a matter of having a professional software engineer develop the tool.' That conclusion is supported only by the Persephone results on Yongning Na and Chatino from Adams et al. 2018 — two tonal languages, with a 20% error rate on Na and 'performed well' on Chatino, with no numbers given here. The retreat itself included languages with four or five speakers (Bardi, Djambarrpuyngu) and one speaker (Mocho), from non-tonal families. The jump from two tonal case studies to 'no theoretical obstacle' for all endangered languages is an overgeneralization. The stress-test note has this right. That said, the paper does not pretend to be a rigorous evaluation; it is a meeting report. The claim is aspirational rather than load-bearing in the scientific sense, because the paper's purpose is to motivate collaboration, not to prove a theorem. Still, a more careful sentence like 'for the languages tested so far' would have been more accurate and cost nothing.\n\nThe citations are honest and mostly self-referential but appropriately so: Adams et al. 2018 is the primary source for the performance claims, and it is a real LREC paper. I don't see citation inflation or invented entities. The paper's limitations are not hidden — it is openly a narrative, and it acknowledges that Persephone is not yet user-friendly.\n\nWho is this for? People working in language documentation and revitalization, especially those deciding whether to fund or participate in similar cross-disciplinary meetings. A researcher looking for a technical contribution will be disappointed. A community organizer or funding agency will find it a useful, if brief, model for what such gatherings can do.\n\nMy take: it deserves a serious referee in the sense that a journal like Language Documentation & Conservation should consider it — it is a legitimate field note, not a research article. I would not cite it for any technical claim, but I might cite it as evidence of community-building momentum. If I were editing, I'd send it to a reviewer who can assess whether the feasibility claim should be softened, and then publish it as a report.\n\nRecommendation: engage with it, but as a community document, not a scientific result.","headline":"A honest meeting report with one overreaching feasibility claim; worth knowing about for the community, but not a research result.","tokens_in":4352,"tokens_out":717,"would_cite":false,"duration_ms":9562,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Practical automatic transcription for endangered languages is within reach, this workshop report contends.","keywords":["automatic speech recognition","endangered languages","language documentation","low-resource languages","neural networks","Persephone toolkit","tonal languages","language revitalization"],"falsifier":"Take one of the most severely endangered languages named at the retreat, for example a Mocho Maya variety with a single speaker, train Persephone on the same minimal data regime of roughly thirty minutes per speaker, and measure the error rate on held-out speech. If the model fails to train or the error rate is far above the 20 percent reported for Yongning Na, the paper's 'no theoretical obstacle' conclusion would need to be weakened.","tokens_in":3476,"feed_emoji":"🗣️","tokens_out":4111,"duration_ms":42715,"temperature":0.7,"pith_summary":"This paper reports on a 2018 retreat that brought together computer scientists, linguists, native speakers, and language activists to discuss automatic speech recognition for endangered languages. It argues that neural-network ASR, particularly the Persephone toolkit, has reached the point where practical transcription of under-resourced languages is possible: a 20 percent error rate on Yongning Na and credible performance on Chatino with thirty minutes of data per speaker. On that basis it concludes the main remaining barrier is not science but engineering, building an interface linguists can use without programming skills. A sympathetic reader cares because automated transcription would attack the time bottleneck in language documentation, especially for languages with very few speakers.","feed_headline":"Speech recognition for endangered languages is within reach","feed_subtitle":"A workshop report says neural ASR already works on tiny tonal corpora; the missing piece is user-friendly software.","key_machinery":"The central object is Persephone, an open-source neural-network ASR toolkit designed for low-resource languages. In the paper's account it is the existence proof: trained on Yongning Na it returned a 20 percent error rate, and it performed well on Chatino, a tonal language, with as little as thirty minutes of data per speaker. It is what turns the paper's conclusion from a wish into a concrete claim.","core_discovery":"The paper's central claim is that there is no theoretical obstacle to creating a practical interface that lets a linguist upload speech and transcriptions and train an ASR model; it is largely a matter of having a professional software engineer build the tool. The report treats Persephone's results on two endangered tonal languages as evidence that neural networks already work on very small corpora, and it concludes that automated transcription can become a normal part of language documentation, changing the linguist's role from transcriber to editor and improving consistency.","pith_inferences":["Beyond the paper, the same logic suggests that even tiny archives of old recordings could become training sets, potentially producing prototype transcribers for hundreds of languages that currently have no ASR resources.","The paper leaves implicit that the strongest test is not another tonal language but a non-tonal or typologically distant language; a reader should not assume the two strong examples cover the world's range of sound systems.","The retreat model itself, convened by a native speaker, could be replicated as a way to align natural-language-processing research with community needs, though the paper does not attempt to measure how much this improved the technical outcomes."],"forward_implications":["If this is right, automated transcription will shift a field linguist's role from transcriber to editor, catching errors and polishing output rather than typing words from scratch.","If this is right, corpora of just thirty minutes per speaker may be enough to bootstrap a workable recognizer for a single speaker, which matters for languages with only a handful of speakers left.","If this is right, a professionally engineered upload-and-train interface would make ASR usable by linguists without programming backgrounds.","If this is right, automated drafts can improve transcription consistency across a documentation project and give researchers fresh insight into the language under study."],"supporting_citations":[{"why":"Supplies the central empirical result: Persephone achieved a 20 percent error rate on Yongning Na and reasonable accuracy for a single speaker with thirty minutes of data, and performed well on Chatino.","marker":"Adams et al. 2018"},{"why":"Reports the design of Na recordings for ASR training and the integration of Persephone into a language documentation workflow, grounding the claim that automated transcription can be deployed.","marker":"Michaud et al. 2018"},{"why":"Describes the creation of the first Chatino speech corpus for ASR training, which is the resource used to test Persephone on a second tonal language.","marker":"Cavar et al. 2016"},{"why":"Provides the contrast case of a dominant language where transcription labor is cheap and abundant, clarifying why endangered languages need automated tools.","marker":"Hernández Mena & Herrera 2017"}],"fun_headline_variants":["Neural ASR works on tiny corpora; missing piece is user-friendly software","No barrier to ASR for endangered languages; software is the gap","Endangered languages: neural ASR ready, software missing","ASR for endangered languages: just needs good software","Neural ASR already works on small corpora; build the tool"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The entire feasibility claim rests on assuming the strong results seen on two tonal languages, Yongning Na and Chatino, will generalize across the many unrelated and severely under-resourced languages of the world, including languages with only one or a handful of speakers.","fun_headline_variants_meta":{"raw":{"variants":["Neural ASR works on tiny corpora; missing piece is user-friendly software","No barrier to ASR for endangered languages; software is the gap","Endangered languages: neural ASR ready, software missing","ASR for endangered languages: just needs good software","Neural ASR already works on small corpora; build the tool"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000183,"raw_usage":{"total_tokens":1170,"prompt_tokens":653,"completion_tokens":517,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":269,"completion_tokens_details":{"reasoning_tokens":427}},"tokens_in":269,"tokens_out":517,"duration_ms":5027,"temperature":1.0,"reasoning_tokens":427,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T11:24:12.145482+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take one of the most severely endangered languages named at the retreat, for example a Mocho Maya variety with a single speaker, train Persephone on the same minimal data regime of roughly thirty minutes per speaker, and measure the error rate on held-out speech. If the model fails to train or the error rate is far above the 20 percent reported for Yongning Na, the paper's 'no theoretical obstacle' conclusion would need to be weakened.","supporting_citations":[],"review_version":1}