Pith. sign in

REVIEW 5 major objections 8 minor 12 references

Lexicography Saves Lives (LSL): Automatically Translating Suicide-Related Language

T0 review · 5 major / 8 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read This paper reports a pipeline that translates a 50-item English suicide-ideation lexicon into 200 languages and scores five of the resulting dictionaries with human ratings for adequacy, fluency, cultural acceptability, and…

desk verdict Good ethics discussion and a sensible idea, but the 200-language resource is unsubstantiated: missing artifacts, a 5-language pilot that can't carry the weight, and a table that contradicts itself. read the letter →

arxiv 2412.15497 v1 pith:L4XTHOAR submitted 2024-12-20 cs.CL cs.AI

classification cs.CLcs.AI
keywords suicideideationdetectionmultilinguallexiconmachinetranslationlow-resourcelanguageshumanevaluationculturalacceptabilityFlores-101NoLanguageLeftBehind
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Suicide-ideation detection is almost entirely built on English and Western cultural contexts, which leaves low-resource languages without usable vocabularies and risks representing suicide language as a word-for-word checklist. This paper tries to close that gap by translating a 50-item English lexicon of suicidal-ideation phrases into 200 languages with a multilingual machine-translation system, then scoring five of the translated dictionaries through human evaluation. The authors introduce five quantitative variables—adequacy, fluency, spelling errors, cultural acceptability, and suicide-context fit—plus free-text fields where native speakers can add alternative translations and local metaphors. Their pilot results show four of five lexicons scoring above 3.0 on adequacy and fluency, with culture and context scores lowest for Finnish, which they read as evidence that the translated dictionaries are a usable start but still require local correction. They also provide ethics guidelines and a public website intended to turn the dictionaries into a community-maintained resource.

What carries the argument

The load-bearing artifact is the 50-item seed lexicon of English suicidal-ideation words and phrases, which the authors translate with a transformer-based multilingual machine-translation model (No Language Left Behind) trained over the Flores-101 benchmark's 200-language list, run through the Fairseq toolkit. The evaluation machinery is a five-variable human-rating protocol—adequacy, fluency, spelling errors, cultural acceptability, and context—averaged per dictionary, plus free-text fields for alternative translations and contributions in the local language. This protocol is what lets the paper claim quality in cultural and suicide-specific terms, not just fluency.

What would settle it

Take any of the 195 unevaluated languages and have two or more native speakers score the translated dictionary with the paper's culture and context questions; if most entries score below 0.5 on either scale, the claim that the 200-language set is a usable resource does not generalize beyond the pilot languages.

Watch

Extended reading notes

Core claim

The central claim is that a resource gap can be addressed responsibly: a 50-item English seed lexicon, originally built from Twitter data, can be machine-translated into 200 languages, and the resulting dictionaries can be assessed with a small human-evaluation protocol rather than trusted blindly. The paper reports that for the five pilot languages—Catalan, Danish, Finnish (listed as 'Finish' in the table), Galician, and German—adequacy and fluency are 'high overall,' with four of five lexicons above 3.0 on both measures, while culture scores range from 0.68 (Finnish) to 0.98 (Danish) and context scores from 0.48 (Finnish) to 0.96 (Galician). It treats the uneven culture and context scores as confirmation that machine translation alone cannot capture the local, metaphorical ways people talk about suicide, and therefore pairs the dictionaries with community contribution and alternative-translation collection.

Load-bearing premise

The paper's conclusion assumes that a 50-item English seed list, plus ratings from at least two self-selected annotators per language on only five languages, can stand in for suicide-related language in all 200 target languages.

Editorial extensions

If this is right

  • Future suicide-ideation detection work in any of the 200 languages can start from a common seed lexicon instead of re-translating English lists ad hoc.
  • The five-variable scoring protocol gives a low-cost way to check whether a translated lexicon is adequate, fluent, culturally acceptable, and relevant to suicidal ideation.
  • The public website creates a mechanism for native speakers to correct mistranslations and add local metaphors, making the dictionaries living resources rather than static outputs.
  • For the five pilot languages, the scores indicate that the translated dictionaries are mostly usable but not complete: Finnish in particular scores low on culture and context, so local review is required even for well-resourced languages.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The pilot's alternative-translation examples suggest that even a fluent machine translation frequently misses the euphemisms and idioms that mark real suicide-related speech; if that pattern holds across the 195 unevaluated languages, coverage of colloquial suicidal language is likely weaker than the adequacy scores alone imply.
  • One direct test of the resource would be to plug the translated dictionary into a lexicon-based detector for the same language and compare precision against a detector built from locally collected suicide-related posts; the paper reports no downstream detection experiment, so the practical utility of the lexicons is unmeasured.
  • The same evaluation template could be applied to sensitive vocabulary beyond suicide—for example, intimate-partner violence or substance-use language—where literal translation also carries cultural risk.
  • Because only five of 200 dictionaries have human scores and each had at least two self-selected annotators, the remaining dictionaries should be treated as drafts awaiting community review rather than validated resources.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 8 minor

Summary. The paper introduces the 'Lexicography Saves Lives Project', whose three contributions are (1) a set of ethical guidelines for developing suicide-related multilingual resources, (2) automatic translation of a 50-item English seed lexicon compiled by O'Dea et al. (2015) into 200 languages using a multilingual NLLB-style translation model, and (3) a public website for community participation and feedback. The authors report a pilot human evaluation of five translated dictionaries (Catalan, Danish, Finnish, Galician, German) using seven variables: Adequacy, Fluency, Spelling Errors, Cultural Acceptability, Context, Alternative Translations, and Contributions in local language. They also present sample alternative translations and community-contributed terms, and outline future plans for improving the quality metrics and soliciting more evaluations.

Significance. If the translated dictionaries and evaluation protocols were released and the reported scores were reliable, this work would be a valuable step toward filling the gap in suicide-related lexical resources for low-resource languages. The ethical discussion, particularly the treatment of linguistic misrepresentation and distributive injustice, is thoughtful and raises important considerations for the NLP community. The community-participation oriented website is a promising mechanism for involving native speakers. However, the significance is currently limited because the central resource is not accessible (the website is a placeholder and the codebook is withheld), and the quantitative evaluation is a small pilot with internal inconsistencies that undermine confidence in the stated results.

major comments (5)
  1. [Table 1 and Section 4.2/4.4] The Catalan Spelling Errors score of 3.7 contradicts the metric definition in Section 4.2, where the best score for Spelling Errors is 0, and it contradicts the text in Section 4.4 that Galician has the highest spelling-error score. This internal inconsistency raises doubts about the integrity of the reported numbers and must be corrected, with a clear statement of the scale used and what the Catalan value actually represents.
  2. [Section 4.3 and Table 1] The pilot evaluation of only five languages with at least two self-selected annotators per language, with no inter-annotator agreement, no confidence intervals, and no report of how many of the 50 entries were actually rated, cannot support the implicit claim that the 200-language lexicon is usable. In particular, Finnish Context=0.48 on a binary scale is statistically indistinguishable from chance, and the paper treats this as a localized finding rather than as evidence that the seed-transfer approach may fail even for a high-resource language.
  3. [Section 2 vs. Section 3] The paper's own argument in Section 2 that suicide-related language is metaphorical, context-dependent, and deeply local directly undermines the decision to translate a fixed English seed list into 200 languages without a per-language validation mechanism. The pilot results, such as the alternative translations in Table 2 and the low Finnish Context score, confirm this tension, yet the paper does not integrate it into the central claims and leaves the other 195 languages completely unvalidated.
  4. [Section 5 and footnotes 5/6] The central resource is not released: the website is described as 'made available upon publication' at 'www.dummy.com', and the codebook is likewise withheld. Since the contribution is a translated lexicon plus an evaluation protocol, the absence of any accessible artifact prevents verification of the claimed 200-language translations and of the evaluation methodology.
  5. [Section 3.1] The description of the translation system is technically confused: Flores 101 is an evaluation benchmark, not a translation system, yet the text says 'We use the Fairseq research toolkit with the transformer-based pre-trained language model (PLM) baseline' that translates into 200 languages. The authors likely mean the NLLB-200 model, but the exact model name, checkpoint, and decoding settings are not specified, which hinders reproducibility.
minor comments (8)
  1. [Throughout] The name 'Finish' is used instead of 'Finnish' in the text, in Table 1, and in Table 3; this should be corrected throughout.
  2. [Section 3, list of seed entries] The list of 50 words and phrases contains duplicates: 'not worth living' appears twice and 'take my own life' appears twice, so the stated count of 50 and the enumerated items should be reconciled.
  3. [Introduction and Appendix] Several cross-references are left unresolved: the contributions section refers to 'section ??' for the ethical guidelines, and the Appendix is referenced as 'Section ??'; these should be replaced with actual section numbers.
  4. [References] The reference for Munzner in Section 5 is incomplete, listing only 'T. Munzner' without a full bibliographic entry; this should be added.
  5. [Section 3.1] The paper interchangeably uses 'Flores 101' and 'Flores200' / NLLB; the relationship between the language list, the benchmark, and the translation model should be clarified, including the exact number of languages actually translated.
  6. [Section 4.1 and Section 4.2] Section 4.1 introduces seven variables (five quantitative and two free-text), but Section 4.2 says 'we use quantitative metrics based on the 5 variables proposed in section 4.1'; the wording should be aligned to avoid ambiguity about which variables are quantified.
  7. [Table 3] The table header and text mention contributions only in 'Danish and Finish'; after fixing the typo, clarify whether contributions in other pilot languages were solicited but not obtained, and if so, why.
  8. [Appendix, language list] The appendix table contains apparent typos and inconsistencies (e.g., 'Chokew' instead of Chokwe, and inconsistent script entries such as 'Arabic low' with a missing script value); the language and script list should be carefully checked against the NLLB/Flores list.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the reported lexicon quality scores are direct human judgments of fixed NLLB translations; no parameter is fitted and no prediction is derived from the evaluation inputs.

full rationale

The paper's derivation chain is: (1) take a 50-item English seed lexicon from O'Dea et al. (2015); (2) translate it into 200 languages with a fixed NLLB/Flores-101 model; (3) ask native speakers to rate five of the resulting lexicons on adequacy, fluency, spelling errors, cultural acceptability, and context; (4) average those ratings with x = sum(E)/N (Eq. 1). The reported numbers in Table 1 are arithmetic summaries of the annotator responses themselves, not outputs of a model fitted to those responses, so there is no sense in which the scores are forced by the inputs by construction. The self-citations (Schoene et al. 2023; Ortega and Church 2023) are used as related work and as methodological pointers to translation-evaluation practice; they do not supply a uniqueness theorem or a premise that entails the measured scores. The paper's own caveats are limitation statements, not circular moves: Section 4.3 calls the study a 'pilot evaluation' with 'at least 2 participants' per language, and Section 6 says the authors 'will improve the methodology behind the lexicon quality scores and identify a means of measuring overall quality.' Those admissions undercut the strength of the empirical claim but do not make the claim self-referential. The remaining concerns (small annotator pool, near-chance Finnish Context score, unreleased resources) are external-validity or reproducibility issues, which are outside the circularity construct. Consequently, no circular step can be quoted, and the appropriate score is 0.

Assumptions & free parameters 0 free parameters · 5 assumptions · 0 invented entities

No numerical free parameters are introduced; the paper instead relies on methodological and domain assumptions about the seed lexicon, the MT model, and the human evaluation. No new theoretical entities are postulated.

assumptions (5)
  • domain assumption The NLLB-based multilingual model produces adequate translations for all 200 target languages from English.
    The paper relies on this pretrained model without verifying output quality for the 195 languages that were not human-evaluated (Section 3.1).
  • domain assumption The 50-item English seed lexicon (O'Dea et al., 2015) is an appropriate source for 200 culturally distinct languages.
    The lexicon is used as the sole translation input, despite the paper's own argument that suicide-related meaning is local and metaphorical (Section 3, Section 2).
  • domain assumption Untrained native speakers can reliably judge cultural acceptability and suicide-related context with binary yes/no questions.
    The pilot evaluation asks non-experts to answer culture and context questions, with no reliability check (Section 4.1, 4.3).
  • domain assumption Averaging a small number of human ratings yields a meaningful quality score for a lexicon.
    Equation 1 averages evaluations without modeling annotator variance, sampling bias, or item-level uncertainty (Section 4.2).
  • domain assumption The philosophical meaning/reference distinction is an adequate basis for the ethical argument about linguistic misrepresentation.
    Used to support the claim that literal translation cannot capture suicide-related meaning (Section 2).

how reviews work

0 comments
Cite this review

Pith. "Pith review of Lexicography Saves Lives (LSL): Automatically Translating Suicide-Related Language." pith.science (2026). https://pith.science/paper/L4XTHOAR

@misc{pith2026241215497,
  author       = {Pith},
  title        = {Pith review of: Lexicography Saves Lives (LSL): Automatically Translating Suicide-Related Language},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/L4XTHOAR}},
  note         = {Machine review of arXiv:2412.15497}
}
read the original abstract

Recent years have seen a marked increase in research that aims to identify or predict risk, intention or ideation of suicide. The majority of new tasks, datasets, language models and other resources focus on English and on suicide in the context of Western culture. However, suicide is global issue and reducing suicide rate by 2030 is one of the key goals of the UN's Sustainable Development Goals. Previous work has used English dictionaries related to suicide to translate into different target languages due to lack of other available resources. Naturally, this leads to a variety of ethical tensions (e.g.: linguistic misrepresentation), where discourse around suicide is not present in a particular culture or country. In this work, we introduce the 'Lexicography Saves Lives Project' to address this issue and make three distinct contributions. First, we outline ethical consideration and provide overview guidelines to mitigate harm in developing suicide-related resources. Next, we translate an existing dictionary related to suicidal ideation into 200 different languages and conduct human evaluations on a subset of translated dictionaries. Finally, we introduce a public website to make our resources available and enable community participation.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

12 extracted references · 8 canonical work pages

  1. [4]

    Philosophy & Technology , 31:669–684

    Ethics and artificial intelligence: suicide pre- vention on facebook. Philosophy & Technology , 31:669–684. Naman Goyal, Cynthia Gao, Vishrav Chaudhary, Peng- Jen Chen, Guillaume Wenzek, Da Ju, Sanjana Kr- ishnan, Marc’Aurelio Ranzato, Francisco Guzmán, and Angela Fan. 2021. The flores-101 evaluation benchmark for low-resource and multilingual ma- chine t...

  2. [8]

    PeerJ, 3:e1455

    Creating a chinese suicide dictionary for iden- tifying suicide risk on social media. PeerJ, 3:e1455. Laura Martinengo, Louise Van Galen, Elaine Lum, Mar- tin Kowalski, Mythily Subramaniam, and Josip Car

  3. [9]

    BMC medicine, 17(1):1–12

    Suicide prevention and depression apps’ sui- cide risk assessment and management: a systematic assessment of adherence to clinical guidelines. BMC medicine, 17(1):1–12. Thomas H McCoy, Victor M Castro, Ashlee M Rober- son, Leslie A Snapper, and Roy H Perlis. 2016. Im- proving prediction of suicide and accidental death after discharge from general hospital...

  4. [10]

    Internet Inter- ventions, 2(2):183–188

    Detecting suicidality on twitter. Internet Inter- ventions, 2(2):183–188. National Institutes of Health et al. 2016. Policy on good clinical practice training for nih awardees involved in nih-funded clinical trials. NOT-OD-16-1482017. Elena Okhapkina, Valentin Okhapkin, and Oleg Kazarin

  5. [12]

    Advances in Neural Information Processing Systems, 35:27730–27744

    Training language models to follow instruc- tions with human feedback. Advances in Neural Information Processing Systems, 35:27730–27744. Hilary Putnam. 1981. Reason, truth and history , vol- ume 3. Cambridge University Press. Diana Ramírez-Cifuentes, Ana Freire, Ricardo Baeza-Yates, Joaquim Puntí, Pilar Medina-Bravo, Diego Alejandro Velazquez, Josep Mari...

  6. [662]

    suicidal ideation in dementia: as- sociations with neuropsychiatric symptoms and sub- type diagnosis

    IEEE. Varsha D Badal and Colin A Depp. 2022. Natural lan- guage processing of medical records: new under- standing of suicide ideation by dementia subtypes: Commentary on “suicidal ideation in dementia: as- sociations with neuropsychiatric symptoms and sub- type diagnosis” by naismith et al. International Psy- chogeriatrics, 34(4):319–321. Loïc Barrault, ...

  7. [2014]

    Mining twitter for suicide prevention. In Natu- ral Language Processing and Information Systems: 19th International Conference on Applications of Nat- ural Language to Information Systems, NLDB 2014, Montpellier, France, June 18-20, 2014. Proceedings 19, pages 250–253. Springer. Ghelmar Astoveza, Randolph Jay P Obias, Roi Jed L Palcon, Ramon L Rodriguez, ...

  8. [2015]

    In 2015 37th annual international conference of the IEEE engineering in Medicine and biology society (EMBC), pages 7316–7319

    The use of technology in suicide prevention. In 2015 37th annual international conference of the IEEE engineering in Medicine and biology society (EMBC), pages 7316–7319. IEEE. Daeun Lee, Migyeong Kang, Minji Kim, and Jinyoung Han. 2022. Detecting suicidality with a contextual graph neural network. In Proceedings of the eighth workshop on computational li...

Show all 12 references
  1. [2017]

    In 2017 31st International Con- ference on Advanced Information Networking and Ap- plications Workshops (WAINA), pages 87–92

    Adaptation of information retrieval methods for identifying of destructive informational influence in social networks. In 2017 31st International Con- ference on Advanced Information Networking and Ap- plications Workshops (WAINA), pages 87–92. IEEE. World Health Organization ...

  2. [2018]

    Machine translation, 32(3):255–278

    Evaluating mt for massive open online courses: A multifaceted comparison between pbsmt and nmt systems. Machine translation, 32(3):255–278. Karen L Celedonia, Marcelo Corrales Compagnucci, Timo Minssen, and Michael Lowery Wilson. 2021. Legal, ethical, and wider implications of...

  3. [2019]

    Smriti Jha, Gerry Chan, Rita Orji, et al

    Frontiers in psychology, 12:637547. Smriti Jha, Gerry Chan, Rita Orji, et al. 2023. Iden- tification of risk factors for suicide and insights for developing suicide prevention technologies: A sys- tematic review and meta-analysis. Human Behavior and Emerging Technologies, 2023...

  4. [2022]

    Neural Computing and Applications, 34(13):10309–10319

    Suicidal ideation and mental disorder detection with attentive relation networks. Neural Computing and Applications, 34(13):10309–10319. Shaoxiong Ji, Shirui Pan, Xue Li, Erik Cambria, Guodong Long, and Zi Huang. 2020. Suicidal ideation detection: A review of machine learning ...

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.