{"id":"aa237c45-a2bb-43de-b92d-f71a84605e76","arxiv_id":"2412.15497","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"The paper reports a 200-language machine-translated suicide lexicon and a pilot five-language human evaluation, but releases no data and contains inconsistent evaluation scores.","lead":"The authors machine-translated a 50-item English suicide-related word list into 200 languages and asked a small number of native speakers to rate 5 of the resulting lists. The paper also proposes ethical guidelines and a community website, but the translated lists, codebook, and website are not actually available in the preprint.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 5-language pilot cannot support the 200-language claim: Finnish's near-chance Context score and the paper's own cultural-localism argument show the seed-transfer assumption is unverified.","rationale":"The reader's weakest assumption identifies seed-list transferability and pilot representativeness as the load-bearing premise, and my reading agrees. Section 2's argument that suicide language is culturally local and metaphorical directly contradicts the Section 3.1 method of mechanically translating a 50-item English Twitter-derived list into 200 languages. The only quantitative evidence for transfer is Table 1's pilot scores, which cover 5 languages with 2 annotators and no uncertainty or agreement metrics. Crucially, Finnish's Context score of 0.48 is near chance for a binary measure, suggesting the translation is not culturally appropriate even for a high-resource language. The additional inconsistency in Table 1 (Catalan Spelling Errors = 3.7 on a scale where 0 is best, contradicting the text) further undermines the auditability of these scores. Because the dictionaries and codebook are not released, the pilot is the only evidence for the 200-language resource, and it does not support the claimed usability. The verdict should be unchanged from the reader's REJECT; the paper would need broader, transparent evaluation and released artifacts to support its central claim.","tokens_in":14876,"tokens_out":10009,"duration_ms":86022,"concrete_test":"Re-analyze the existing pilot data: compute per-language Wilson 95% confidence intervals for the binary Culture and Context proportions using the number of entries and raters reported in Section 4.3, and report inter-annotator agreement (e.g., Cohen's kappa). If the Finnish Context interval includes 0.5, the pilot provides no evidence of cultural appropriateness for that language. Then extend the evaluation to a stratified sample of at least 20 languages from the Flores list spanning language families and high/low resource tiers; if Culture/Context scores vary widely or fall below 0.7 for several languages, the 5-language pilot cannot support the 200-language claim. Also request the raw annotation sheets to audit Table 1's Catalan Spelling Errors value of 3.7 against the metric definition in Section 4.2.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim of a usable 200-language suicide lexicon rests on two nested assumptions: (i) the 50-item English seed list from O'Dea et al. (2015) can be meaningfully transferred by automatic translation, and (ii) human scores on 5 pilot languages with at least 2 self-selected annotators each can represent the remaining 195 languages. Section 2 argues against (i) by showing that suicide language is metaphorical, context-dependent, and local, e.g., 'my suicide letter' is not idiomatic German 'Mein Selbstmordbrief' but 'Mein Abschiedsbrief'. The pilot data do not rescue (ii): Table 1 gives Finnish Culture=0.68 and Context=0.48 on a binary 0/1 scale, which is statistically indistinguishable from chance with roughly 50 entries and 2 raters; the paper treats this as a localized finding rather than evidence that the transfer approach fails even for a high-resource language. No confidence intervals, inter-annotator agreement, or stratified sampling are reported. The reliability of even these pilot numbers is further undermined by Table 1's Catalan Spelling Errors value of 3.7 when Section 4.2 defines 0 as the best score and the text says Galician has the highest error rate. Because the dictionaries, codebook, and website are not released, these five-language scores are the only quantitative support for the 200-language resource, and they are insufficient.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces the 'Lexicography Saves Lives Project', whose three contributions are (1) a set of ethical guidelines for developing suicide-related multilingual resources, (2) automatic translation of a 50-item English seed lexicon compiled by O'Dea et al. (2015) into 200 languages using a multilingual NLLB-style translation model, and (3) a public website for community participation and feedback. The authors report a pilot human evaluation of five translated dictionaries (Catalan, Danish, Finnish, Galician, German) using seven variables: Adequacy, Fluency, Spelling Errors, Cultural Acceptability, Context, Alternative Translations, and Contributions in local language. They also present sample alternative translations and community-contributed terms, and outline future plans for improving the quality metrics and soliciting more evaluations.","tokens_in":15105,"tokens_out":5540,"duration_ms":39503,"significance":"If the translated dictionaries and evaluation protocols were released and the reported scores were reliable, this work would be a valuable step toward filling the gap in suicide-related lexical resources for low-resource languages. The ethical discussion, particularly the treatment of linguistic misrepresentation and distributive injustice, is thoughtful and raises important considerations for the NLP community. The community-participation oriented website is a promising mechanism for involving native speakers. However, the significance is currently limited because the central resource is not accessible (the website is a placeholder and the codebook is withheld), and the quantitative evaluation is a small pilot with internal inconsistencies that undermine confidence in the stated results.","major_comments":[{"comment":"The Catalan Spelling Errors score of 3.7 contradicts the metric definition in Section 4.2, where the best score for Spelling Errors is 0, and it contradicts the text in Section 4.4 that Galician has the highest spelling-error score. This internal inconsistency raises doubts about the integrity of the reported numbers and must be corrected, with a clear statement of the scale used and what the Catalan value actually represents.","section":"Table 1 and Section 4.2/4.4"},{"comment":"The pilot evaluation of only five languages with at least two self-selected annotators per language, with no inter-annotator agreement, no confidence intervals, and no report of how many of the 50 entries were actually rated, cannot support the implicit claim that the 200-language lexicon is usable. In particular, Finnish Context=0.48 on a binary scale is statistically indistinguishable from chance, and the paper treats this as a localized finding rather than as evidence that the seed-transfer approach may fail even for a high-resource language.","section":"Section 4.3 and Table 1"},{"comment":"The paper's own argument in Section 2 that suicide-related language is metaphorical, context-dependent, and deeply local directly undermines the decision to translate a fixed English seed list into 200 languages without a per-language validation mechanism. The pilot results, such as the alternative translations in Table 2 and the low Finnish Context score, confirm this tension, yet the paper does not integrate it into the central claims and leaves the other 195 languages completely unvalidated.","section":"Section 2 vs. Section 3"},{"comment":"The central resource is not released: the website is described as 'made available upon publication' at 'www.dummy.com', and the codebook is likewise withheld. Since the contribution is a translated lexicon plus an evaluation protocol, the absence of any accessible artifact prevents verification of the claimed 200-language translations and of the evaluation methodology.","section":"Section 5 and footnotes 5/6"},{"comment":"The description of the translation system is technically confused: Flores 101 is an evaluation benchmark, not a translation system, yet the text says 'We use the Fairseq research toolkit with the transformer-based pre-trained language model (PLM) baseline' that translates into 200 languages. The authors likely mean the NLLB-200 model, but the exact model name, checkpoint, and decoding settings are not specified, which hinders reproducibility.","section":"Section 3.1"}],"minor_comments":[{"comment":"The name 'Finish' is used instead of 'Finnish' in the text, in Table 1, and in Table 3; this should be corrected throughout.","section":"Throughout"},{"comment":"The list of 50 words and phrases contains duplicates: 'not worth living' appears twice and 'take my own life' appears twice, so the stated count of 50 and the enumerated items should be reconciled.","section":"Section 3, list of seed entries"},{"comment":"Several cross-references are left unresolved: the contributions section refers to 'section ??' for the ethical guidelines, and the Appendix is referenced as 'Section ??'; these should be replaced with actual section numbers.","section":"Introduction and Appendix"},{"comment":"The reference for Munzner in Section 5 is incomplete, listing only 'T. Munzner' without a full bibliographic entry; this should be added.","section":"References"},{"comment":"The paper interchangeably uses 'Flores 101' and 'Flores200' / NLLB; the relationship between the language list, the benchmark, and the translation model should be clarified, including the exact number of languages actually translated.","section":"Section 3.1"},{"comment":"Section 4.1 introduces seven variables (five quantitative and two free-text), but Section 4.2 says 'we use quantitative metrics based on the 5 variables proposed in section 4.1'; the wording should be aligned to avoid ambiguity about which variables are quantified.","section":"Section 4.1 and Section 4.2"},{"comment":"The table header and text mention contributions only in 'Danish and Finish'; after fixing the typo, clarify whether contributions in other pilot languages were solicited but not obtained, and if so, why.","section":"Table 3"},{"comment":"The appendix table contains apparent typos and inconsistencies (e.g., 'Chokew' instead of Chokwe, and inconsistent script entries such as 'Arabic low' with a missing script value); the language and script list should be carefully checked against the NLLB/Flores list.","section":"Appendix, language list"}],"recommendation":"major_revision","confidential_remarks":"The paper is honest about the preliminary nature of the evaluation, but the internal contradiction in Table 1 and the unavailability of the resource make it impossible to verify the central claims at this stage. The ethical discussion and the idea of community participation are valuable, but the manuscript currently reads as a project proposal rather than a finished research contribution. A major revision that releases the artifacts, corrects the reported data, and substantially strengthens or appropriately limits the evaluation claims could make it publishable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This paper has a genuinely thoughtful ethical section and a plausible goal, but as submitted it does not deliver on its central claim of a usable 200-language suicide-related lexicon.\n\nWhat's actually new: extending the O'Dea/Sawhney English lexicon to 200 languages via NLLB/Flores, and proposing a pilot evaluation protocol that includes cultural acceptability and context variables. That is a legitimate extension of prior single-language work (e.g., Lee et al. for Korean). The free-text collection of alternative translations and local contributions is a sensible idea.\n\nWhat it does well: Section 2 makes a strong case that suicide language is metaphorical, context-dependent, and culturally local, connecting this to meaning/reference and distributive injustice. The discussion of legal contexts where suicide attempts are criminalized is important and often missing in this literature. The paper is also honest about its own limitations in the conclusion.\n\nSoft spots are substantial. The claimed resource is not available: the website is a placeholder, the codebook is withheld until publication, and the only quantitative support is a 5-language pilot, all high-resource European languages, with at least two self-selected annotators each. No inter-annotator agreement, no confidence intervals. Finnish's Culture 0.68 and Context 0.48 are near chance on a binary 0/1 scale with roughly 50 entries, which the paper treats as a localized finding rather than evidence that the transfer approach is shaky even for a high-resource language. The paper's own Section 2 argument cuts against the seed-transfer assumption for the other 195 languages. There is also a clear internal inconsistency: Catalan's Spelling Errors score of 3.7, where 0 is the best score, while the text says Galician has the highest error rate. That undermines trust in the reported numbers.\n\nCitation pattern is fine, circularity burden is low, and the paper is not incoherent on its own terms. But the empirical contribution is not substantiated as written. The pilot needs low-resource languages, released artifacts, and a corrected table before the resource claim can be evaluated.\n\nI would not cite this in its current form, but it deserves serious referee attention rather than a desk reject, because the ethical framing and the criticism of naive MT-based lexicon transfer are valuable for the field. A generous but demanding review could help the authors turn this into a useful position paper with a credible resource. The stress-test note's concern is valid and should be passed on.","headline":"Good ethics discussion and a sensible idea, but the 200-language resource is unsubstantiated: missing artifacts, a 5-language pilot that can't carry the weight, and a table that contradicts itself.","tokens_in":15682,"tokens_out":2495,"would_cite":false,"duration_ms":20545,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper reports a pipeline that translates a 50-item English suicide-ideation lexicon into 200 languages and scores five of the resulting dictionaries with human ratings for adequacy, fluency, cultural acceptability, and…","keywords":["suicide ideation detection","multilingual lexicon","machine translation","low-resource languages","human evaluation","cultural acceptability","Flores-101","No Language Left Behind"],"falsifier":"Take any of the 195 unevaluated languages and have two or more native speakers score the translated dictionary with the paper's culture and context questions; if most entries score below 0.5 on either scale, the claim that the 200-language set is a usable resource does not generalize beyond the pilot languages.","tokens_in":14655,"feed_emoji":"🌐","tokens_out":8280,"duration_ms":66378,"temperature":0.7,"pith_summary":"Suicide-ideation detection is almost entirely built on English and Western cultural contexts, which leaves low-resource languages without usable vocabularies and risks representing suicide language as a word-for-word checklist. This paper tries to close that gap by translating a 50-item English lexicon of suicidal-ideation phrases into 200 languages with a multilingual machine-translation system, then scoring five of the translated dictionaries through human evaluation. The authors introduce five quantitative variables—adequacy, fluency, spelling errors, cultural acceptability, and suicide-context fit—plus free-text fields where native speakers can add alternative translations and local metaphors. Their pilot results show four of five lexicons scoring above 3.0 on adequacy and fluency, with culture and context scores lowest for Finnish, which they read as evidence that the translated dictionaries are a usable start but still require local correction. They also provide ethics guidelines and a public website intended to turn the dictionaries into a community-maintained resource.","feed_headline":"Suicide-ideation lexicon expanded to 200 languages","feed_subtitle":"Human raters scored five translations; four were adequate and fluent, but local metaphors still need community review.","key_machinery":"The load-bearing artifact is the 50-item seed lexicon of English suicidal-ideation words and phrases, which the authors translate with a transformer-based multilingual machine-translation model (No Language Left Behind) trained over the Flores-101 benchmark's 200-language list, run through the Fairseq toolkit. The evaluation machinery is a five-variable human-rating protocol—adequacy, fluency, spelling errors, cultural acceptability, and context—averaged per dictionary, plus free-text fields for alternative translations and contributions in the local language. This protocol is what lets the paper claim quality in cultural and suicide-specific terms, not just fluency.","core_discovery":"The central claim is that a resource gap can be addressed responsibly: a 50-item English seed lexicon, originally built from Twitter data, can be machine-translated into 200 languages, and the resulting dictionaries can be assessed with a small human-evaluation protocol rather than trusted blindly. The paper reports that for the five pilot languages—Catalan, Danish, Finnish (listed as 'Finish' in the table), Galician, and German—adequacy and fluency are 'high overall,' with four of five lexicons above 3.0 on both measures, while culture scores range from 0.68 (Finnish) to 0.98 (Danish) and context scores from 0.48 (Finnish) to 0.96 (Galician). It treats the uneven culture and context scores as confirmation that machine translation alone cannot capture the local, metaphorical ways people talk about suicide, and therefore pairs the dictionaries with community contribution and alternative-translation collection.","pith_inferences":["The pilot's alternative-translation examples suggest that even a fluent machine translation frequently misses the euphemisms and idioms that mark real suicide-related speech; if that pattern holds across the 195 unevaluated languages, coverage of colloquial suicidal language is likely weaker than the adequacy scores alone imply.","One direct test of the resource would be to plug the translated dictionary into a lexicon-based detector for the same language and compare precision against a detector built from locally collected suicide-related posts; the paper reports no downstream detection experiment, so the practical utility of the lexicons is unmeasured.","The same evaluation template could be applied to sensitive vocabulary beyond suicide—for example, intimate-partner violence or substance-use language—where literal translation also carries cultural risk.","Because only five of 200 dictionaries have human scores and each had at least two self-selected annotators, the remaining dictionaries should be treated as drafts awaiting community review rather than validated resources."],"forward_implications":["Future suicide-ideation detection work in any of the 200 languages can start from a common seed lexicon instead of re-translating English lists ad hoc.","The five-variable scoring protocol gives a low-cost way to check whether a translated lexicon is adequate, fluent, culturally acceptable, and relevant to suicidal ideation.","The public website creates a mechanism for native speakers to correct mistranslations and add local metaphors, making the dictionaries living resources rather than static outputs.","For the five pilot languages, the scores indicate that the translated dictionaries are mostly usable but not complete: Finnish in particular scores low on culture and context, so local review is required even for well-resourced languages."],"supporting_citations":[{"why":"Supplies the 50-item English seed list of suicidal-ideation words and phrases that the entire translation pipeline starts from.","marker":"O’dea et al. (2015)"},{"why":"Connects the seed lexicon to suicide-ideation feature extraction and transmits the list the paper adopts.","marker":"Sawhney et al. (2018)"},{"why":"Provides the No Language Left Behind multilingual machine-translation system and the 200-language evaluation list.","marker":"Team et al. (2022b)"},{"why":"Defines the Flores-101 evaluation benchmark that structures the translation task and language coverage.","marker":"Goyal et al. (2021)"},{"why":"Provides the Fairseq toolkit used to run the transformer-based multilingual model.","marker":"Ott et al. (2019)"},{"why":"Supplies the transformer architecture underlying the pre-trained multilingual translation model.","marker":"Vaswani et al. (2017)"},{"why":"Establishes the adequacy and fluency evaluation variables the paper adapts for its human-rating protocol.","marker":"Castilho et al. (2018)"},{"why":"Grounds the argument that low-resource machine translation needs cultural and linguistic context beyond literal transfer.","marker":"Ortega and Church (2023)"},{"why":"Supports the claim that suicidal language is not universal and therefore needs culture- and context-specific evaluation.","marker":"Kirtley et al. (2022)"}],"fun_headline_variants":["Suicide lexicon reaches 200 languages, but culture scores lag","Beyond translation: suicide lexicon in 200 languages needs locals","LSL project: suicide lexicon in 200 languages, human-rated","Translating suicide terms to 200 languages isn't the final step","Suicide-talk lexicon: 200 languages, yet local context is key"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper's conclusion assumes that a 50-item English seed list, plus ratings from at least two self-selected annotators per language on only five languages, can stand in for suicide-related language in all 200 target languages.","fun_headline_variants_meta":{"raw":{"variants":["Suicide lexicon reaches 200 languages, but culture scores lag","Beyond translation: suicide lexicon in 200 languages needs locals","LSL project: suicide lexicon in 200 languages, human-rated","Translating suicide terms to 200 languages isn't the final step","Suicide-talk lexicon: 200 languages, yet local context is key"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000267,"raw_usage":{"total_tokens":1607,"prompt_tokens":929,"completion_tokens":678,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":545,"completion_tokens_details":{"reasoning_tokens":589}},"tokens_in":545,"tokens_out":678,"duration_ms":6202,"temperature":1.0,"reasoning_tokens":589,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T11:22:35.962053+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take any of the 195 unevaluated languages and have two or more native speakers score the translated dictionary with the paper's culture and context questions; if most entries score below 0.5 on either scale, the claim that the 200-language set is a usable resource does not generalize beyond the pilot languages.","supporting_citations":[],"review_version":1}