Pith. sign in

REVIEW 4 major objections 5 minor 78 references

Multilingual and Explainable Text Detoxification with Parallel Corpora

T0 review · 4 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read This paper claims that manually curated parallel corpora in German, Chinese, Arabic, Hindi, and Amharic, combined with a cluster-aware Chain-of-Thought prompt, make text detoxification both more multilingual and more explainable.

desk verdict The five-language parallel detoxification corpora are a real resource worth citing; the CoT method claim is not supported by the current evaluation and needs per-language validation of the toxicity classifier, significance testing, and held-out cluster selection. read the letter →

arxiv 2412.11691 v1 pith:2CIZYQW5 submitted 2024-12-16 cs.CL cs.AI

classification cs.CLcs.AI
keywords textdetoxificationstyletransfermultilingualparallelcorporachain-of-thoughtpromptingexplainableAItoxiclanguageLLMParaDetox
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tries to establish two related things: that parallel text detoxification data can be produced by hand for German, Chinese, Arabic, Hindi, and Amharic, and that explaining toxicity with an LLM can be turned into a better detoxification prompt. The authors collect and release manually curated toxic-to-neutral sentence pairs for the five new languages, then use GPT-4 to label every sentence in the resulting nine-language corpus with a toxicity level, tone, language style, implied sentiment, and negative connotations. Those labels are clustered into three repair strategies per language, and the cluster identity is fed into a Chain-of-Thought prompt that first diagnoses the input and then rewrites it using a representative example. In the paper's automatic evaluation, this cluster-conditioned Chain-of-Thought prompt reports the highest style-transfer accuracy among compared detoxification systems and the highest average joint score, which combines non-toxicity, content preservation, and fluency. If the result holds, detoxification becomes a more predictable, explainable operation for languages that previously had no parallel training data.

What carries the argument

The load-bearing mechanism is the cluster-conditioned Chain-of-Thought (CoT) prompt. First, GPT-4 extracts five descriptive features (toxicity level, tone, language type, implied sentiment, negative connotations) for toxic and detoxified sentences; the toxic sentences' feature vectors are validated against native speakers, then one-hot encoded and clustered by k-means into three groups per language. The three clusters correspond to repair strategies—removing profanities, rephrasing condescending or biased language, and lightly adjusting informal text—and each cluster carries a human-readable explanation and a representative parallel pair. When a new toxic sentence arrives, the prompt asks the LLM to estimate its features, assign it to a cluster, and detoxify it using that cluster's explanation and example. This turns the explainability analysis into a conditioning signal for generation, which the paper says reduces hallucination and makes the edit more targeted.

What would settle it

Give native speakers a blind sample of 200 GPT-4 CoT outputs and 200 GPT-4 few-shot outputs per language and ask them to judge toxicity and meaning preservation. If human raters do not rate CoT outputs as non-toxic more often, or if the per-language ranking changes, the reported STA-based Joint advantage is an artifact of the toxicity checker. A direct pilot: measure the checker's agreement with human toxicity judgments on the human detoxification references for Chinese and Amharic, where the checker's scores diverge most.

Watch

Extended reading notes

Core claim

The paper's central claim is that supervised text detoxification can be extended to German, Chinese, Arabic, Hindi, and Amharic by manually curating parallel toxic-to-neutral pairs, and that explaining toxicity with GPT-4 can be turned into a better detoxification prompt. The authors collect 400 training and 600 test sentence pairs for each new language, then ask GPT-4 to label sentences across the nine languages with toxicity level, tone, language style, implied sentiment, and negative connotations, with native-speaker validation reported at 98% agreement. The toxic sentences' labels are one-hot encoded and clustered by k-means into three per-language detoxification strategies, and a Chain-of-Thought prompt first assigns a new toxic sentence to a cluster and then rewrites it using the cluster's explanation and a representative example. The paper reports that this method achieves the highest Style Transfer Accuracy (the share of outputs judged non-toxic by the toxicity classifier) among compared approaches and the highest average Joint score, an aggregate of non-toxicity, content preservation, and fluency; it interprets this as evidence that cluster knowledge reduces hallucination and yields more precise edits.

Load-bearing premise

The load-bearing assumption is that the automatic toxicity checker used to score outputs is a fair judge of non-toxicity in all nine languages; the paper's own Table 9 shows Chinese human detoxification references receiving only 0.266 on that check, so if the checker is biased, the reported rankings of methods, including the CoT advantage, would not be trustworthy.

Editorial extensions

If this is right

  • The new 400/600 train/test splits for German, Hindi, Amharic, Arabic, and Chinese give subsequent work a standard benchmark for supervised detoxification in languages with no prior parallel corpus.
  • The cluster-conditioned CoT prompt can be applied without fine-tuning: given any toxic sentence, the model diagnoses the edit type from the cluster and rewrites accordingly, reducing hallucinated content.
  • The Delete baseline's strong showing in Chinese, Arabic, and Amharic implies that for some languages, removing toxic tokens is a competitive fallback when no good paraphrase model exists.
  • The descriptive feature analysis provides a reusable map of each language's toxic lexicon, for instance animal insults in Hindi and Amharic and refugee-related wordplay in German, that can inform language-specific moderation and generation.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: because the STA checker disagrees sharply with human detoxification references in some languages (Chinese human references get STA 0.266), the reported Joint-score gaps could be re-ranked by a human-validated toxicity measure; the CoT method's advantage should be treated as pending that check.
  • Editorial inference: the describe-and-cluster recipe is not specific to toxicity; applying feature extraction plus k-means over repair strategies to formality transfer or sentiment transfer could produce the same kind of conditioning signal, and existing parallel datasets would allow that test.
  • Editorial inference: the cluster explanations in the paper are English-only, so varying the language and phrasing of the cluster instructions, or the number of clusters, is an untested dimension that could matter more in lower-resource languages such as Amharic.
  • Editorial inference: a direct human evaluation of fluency and content preservation on the test outputs would settle whether the STA advantage reflects genuinely better detoxifications or merely a shift toward the classifier's notion of non-toxicity.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper extends parallel text detoxification corpora to German, Chinese, Arabic, Hindi, and Amharic, reports a detailed annotation pipeline for each language, and uses GPT-4 to annotate descriptive features of toxic and non-toxic sentences across nine languages. On the basis of this feature analysis, the authors build per-language K-means clusters and propose a Chain-of-Thought (CoT) prompting method that identifies a test sentence's cluster and supplies a representative detoxification example in the prompt. The central empirical claim is that this cluster-conditioned CoT method achieves the best average joint score J = STA x SIM x ChrF1 and the best STA scores against several unsupervised, supervised, and few-shot baselines.

Significance. If the claims hold, the resource contribution is genuinely valuable: manually curated parallel detoxification data for five non-European languages, with public data and code, extends the reach of supervised detoxification substantially. The descriptive feature analysis across nine languages is also a useful reference for cross-lingual work on toxicity. However, the empirical comparison that supports the CoT method is currently not trustworthy. The STA toxicity classifier used in the joint metric is never validated per language, and internal tables show it disagrees sharply with human gold labels for Chinese and Amharic. In addition, the clusters and prompt examples used in the CoT method are derived from the full 1,000-pair set, including the 600 test instances, so the evaluation is not a clean test of generalization. These issues are fixable with additional experiments and reporting, but they are load-bearing for the paper's central claim.

major comments (4)
  1. [Section 5, Tables 4, 9, 10] The STA component of the joint score is computed by an XLM-R-large classifier fine-tuned on 5,000 subsampled examples per language, but the paper reports no per-language validation, accuracy, F1, or calibration of this classifier. The internal results show that the classifier is not consistently measuring non-toxicity: in Table 9, Chinese human detoxified references receive STA 0.266, meaning most gold non-toxic rewrites are deemed toxic, while GPT-4 CoT on the same language receives STA 0.716; in Table 10, the Duplicate baseline for Amharic receives STA 0.426, so unchanged toxic inputs are called non-toxic nearly half the time. Because J is multiplicative, such a biased STA directly distorts every comparison in Table 4, including the claimed CoT advantage over few-shot prompting. The 98% expert agreement reported in Section 4.1 concerns GPT-4's descriptive feature annotations and not the validity of the STA classifier, so it does not mitigate this concern.
  2. [Sections 3.6, 4.1, 4.5, Appendix A.3] The CoT method is evaluated on test instances that were used to construct the method. Feature extraction and K-means clustering are run on all 1,000 pairs per language, and the 600 test sentences are a subset of these 1,000 pairs. The cluster definitions, the choice of K=3, and the representative cluster examples embedded in the CoT prompt are therefore informed by the same inputs and references that are later scored in Table 4. This is a form of test-set leakage: the comparison does not measure how the method would perform on unseen inputs. The authors should derive clusters from the 400 training pairs only, or otherwise exclude the test set from all prompt and cluster construction, and then re-run the evaluation.
  3. [Section 7, Table 4, Appendix C.5] The reported advantage of GPT-4 CoT over few-shot prompting is very small on average (0.331 vs 0.324) and negative for English (0.326 vs 0.475), yet no significance tests, confidence intervals, or multiple-run variance are reported. Appendix C.5 states that GPT-4 was run with default hyperparameters including temperature=1.0, and no number of repeated inferences is given. With a single stochastic run per input, the observed differences may be sampling noise. Paired bootstrap tests or multiple runs with confidence intervals are needed before claiming that CoT 'achieved the highest scores across all approaches'.
  4. [Section 4.1, Table 7, Appendix E] The feature-extraction analysis relies on GPT-4 annotations for the full 1,000-pair set, and the paper states that experts agreed with GPT-4 in 98% of cases. However, the size of the expert-reviewed sample, the number of annotators per language, and the agreement measure are not reported. Since the method's clusters and prompts are built directly on these annotations, the reliability of this 98% figure matters for the validity of the CoT approach. The authors should specify the validation protocol and report per-language agreement.
minor comments (5)
  1. [Section 3, Annotators Compensations] The text reports 'C20 per hour' and 'C7.65 above the minimum wage'; the 'C' appears to be a rendering error for the euro sign, which should be corrected.
  2. [Section 3.3.2, Amharic Annotation Process] 'Two annotators ... were evolved in the main annotation' should read 'were involved in the main annotation'.
  3. [Section 3.1.2, German Annotation Process] The text says each sample was transcribed by only one annotator, while Table 1 lists two annotators per sentence for German; the apparent discrepancy between the prose and the table should be clarified.
  4. [Appendix C.5] The phrase 'top_k=0.0' is not a standard GPT-4 API parameter in the same sense as temperature and may confuse readers; the paper should either explain the sampling configuration precisely or omit the unsupported parameter.
  5. [Section 2, Related Work] The related-work section would benefit from a short paragraph positioning the new corpora against the MultiParaDetox and TextDetox CLEF-2024 shared task, since the data are already described as the basis of that task in footnote 3.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity; the corpus construction, metric computation, and CoT comparison are not equivalent to their inputs by construction.

full rationale

The paper's central contributions—new manually curated parallel detoxification corpora (Sections 3.1–3.5), GPT-4-based descriptive feature analysis with native-speaker validation (Section 4), and the Chain-of-Thoughts prompting method (Section 4.5)—do not reduce, by the paper's own equations or definitions, to their inputs. The corpora are curated from external toxic-language datasets and are publicly released; the STA classifier is fine-tuned on 5,000 subsampled toxicity-classification examples per language explicitly 'not used for ParaDetox data collection' (Section 5), so the J-score comparisons in Tables 4, 8–10 are not fitted from the detoxified outputs being scored. The CoT method uses GPT-4-derived cluster prompts, but the detoxified sentences are generated by GPT-4 and are not computed from the cluster fit; the cluster label is a conditioning input, and the claimed improvement is an empirical comparison against few-shot prompting on the same test set. The self-citations (Dementieva et al. 2024a for the toxicity definition; Logacheva et al. 2022 for the evaluation pipeline; Dementieva et al. 2024b for the shared task) point to published, externally available resources and are not load-bearing. Concerns raised by reviewers—the unvalidated per-language STA signal (e.g., Table 9 gives Chinese human references STA 0.266) and the use of all 1,000 pairs per language, including test pairs, to build the CoT clusters (Sections 4.1 and 4.5)—are correctness, calibration, and potential test-set-leakage risks, not by-construction equivalences between claimed predictions and fitted inputs. The paper itself acknowledges limitations, including reliance on closed-source GPT-4 and English-only cluster explanations. Therefore no significant circularity is present.

Assumptions & free parameters 3 free parameters · 3 assumptions · 0 invented entities

The method's cluster definitions and K=3 are derived from GPT-4 feature annotations on the full 1,000-pair set, so the test split is not held out from the analysis that defines the prompting strategy. The validity of the GPT-4 feature annotations is assumed without independent gold-standard feature labels, and the STA classifier's language-consistency is undermined by the Chinese human-reference score.

free parameters (3)
  • Number of clusters K = 3 per language
    Chosen by hyperparameter experiments in Section 4.5 on the same data used for evaluation; no held-out validation reported.
  • Chinese toxic-score filtering threshold = 0.978
    Threshold applied in Section 3.5.1 to select candidate sentences from TOXICN; hand-chosen and affects the composition of the Chinese dataset.
  • Chinese curation thresholds = 1 to 5 toxic words, toxic word ratio < 0.5, length 3 to 50 words
    Heuristic filtering criteria in Section 3.5.1 used to create the Chinese candidate pool; these choices shape the dataset and are not derived from a principled optimization.
assumptions (3)
  • domain assumption GPT-4 feature annotations accurately characterize toxicity and detoxification in all nine languages.
    Section 4.1 states native speakers validated them, reporting 98% agreement, but no validation protocol, sample size, or disagreement analysis is provided. The same GPT-4 outputs are used to define the clusters for the CoT method.
  • domain assumption The STA classifier trained on 5,000 subsampled examples per language reliably measures style transfer accuracy.
    Section 5 and Appendix B describe the classifier; it is used for all J scores. The Chinese human-reference STA of 0.266 (Table 9) suggests the classifier does not align with human judgments for that language.
  • domain assumption Toxicity is defined as vulgar or profane language excluding deep insults and hate, following Dementieva et al. (2024a).
    Section 3 introduces this definition; it shapes what was annotated and analyzed, and excludes categories of abuse that other toxicity datasets include.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Multilingual and Explainable Text Detoxification with Parallel Corpora." pith.science (2026). https://pith.science/paper/2CIZYQW5

@misc{pith2026241211691,
  author       = {Pith},
  title        = {Pith review of: Multilingual and Explainable Text Detoxification with Parallel Corpora},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/2CIZYQW5}},
  note         = {Machine review of arXiv:2412.11691}
}
read the original abstract

Even with various regulations in place across countries and social media platforms (Government of India, 2021; European Parliament and Council of the European Union, 2022, digital abusive speech remains a significant issue. One potential approach to address this challenge is automatic text detoxification, a text style transfer (TST) approach that transforms toxic language into a more neutral or non-toxic form. To date, the availability of parallel corpora for the text detoxification task (Logachevavet al., 2022; Atwell et al., 2022; Dementievavet al., 2024a) has proven to be crucial for state-of-the-art approaches. With this work, we extend parallel text detoxification corpus to new languages -- German, Chinese, Arabic, Hindi, and Amharic -- testing in the extensive multilingual setup TST baselines. Next, we conduct the first of its kind an automated, explainable analysis of the descriptive features of both toxic and non-toxic sentences, diving deeply into the nuances, similarities, and differences of toxicity and detoxification across 9 languages. Finally, based on the obtained insights, we experiment with a novel text detoxification method inspired by the Chain-of-Thoughts reasoning approach, enhancing the prompting process through clustering on relevant descriptive attributes.

Figures

Figures reproduced from arXiv: 2412.11691 by the authors.

Figure 1
Figure 1. Examples of the desired texts detoxification [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. In this work, we extend parallel text detoxification data to new languages as well as provide explainability [PITH_FULL_IMAGE:figures/full_fig_p002_2.png] view at source ↗
Figure 3
Figure 3. Extracted with GPT-4 toxicity levels and top descriptive features per toxic and non-toxic parts in the [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figures from the paper (11 more)
Figure 4
Figure 4. Figure 4: Top-5 extracted keywords from toxic parts. [PITH_FULL_IMAGE:figures/full_fig_p007_4.png]
Figure 5
Figure 5. Figure 5: Text detoxification with CoT: analyze the [PITH_FULL_IMAGE:figures/full_fig_p008_5.png]
Figure 6
Figure 6. Figure 6: Comparison of toxic and non-toxic texts lengths distributions per each language. [PITH_FULL_IMAGE:figures/full_fig_p020_6.png]
Figure 7
Figure 7. Figure 7: Levenshtein distances between toxic and non-toxic parts distribution. [PITH_FULL_IMAGE:figures/full_fig_p020_7.png]
Figure 8
Figure 8. Figure 8: Descriptive words of the different features in the toxic training part for all languages. (a) Tone (b) Language Type (c) Sentiment (d) All together [PITH_FULL_IMAGE:figures/full_fig_p022_8.png]
Figure 9
Figure 9. Figure 9: Descriptive words of the different features in the toxic test part for all languages [PITH_FULL_IMAGE:figures/full_fig_p022_9.png]
Figure 10
Figure 10. Figure 10: The PCA projection of the toxic sentences cluster based on their descriptive features and detoxification types [PITH_FULL_IMAGE:figures/full_fig_p023_10.png]
Figure 11
Figure 11. Figure 11: Examples of parallel detoxified pairs from HiParaDetox. [PITH_FULL_IMAGE:figures/full_fig_p026_11.png]
Figure 12
Figure 12. Figure 12: Examples of parallel detoxified pairs from AmParaDetox. [PITH_FULL_IMAGE:figures/full_fig_p027_12.png]
Figure 13
Figure 13. Figure 13: Examples of parallel detoxified pairs from ZhParaDetox. [PITH_FULL_IMAGE:figures/full_fig_p028_13.png]
Figure 14
Figure 14. Figure 14: Examples of parallel detoxified pairs from ArParaDetox. [PITH_FULL_IMAGE:figures/full_fig_p028_14.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

78 extracted references · 35 canonical work pages

  1. [1]

    Katherine Atwell, Sabit Hassan, and Malihe Alikhani. 2022. https://aclanthology.org/2022.coling-1.530 APPDIA: A discourse-aware transformer-based style transfer model for offensive social media conversations . In Proceedings of the 29th International Conference on Computational Linguistics, COLING 2022, Gyeongju, Republic of Korea, October 12-17, 2022 , p...

  2. [2]

    Abinew Ali Ayele, Skadi Dinter, Tadesse Destaw Belay, Tesfa Tegegne Asfaw, Seid Muhie Yimam, and Chris Biemann. 2022. https://ieeexplore.ieee.org/document/9971189 The 5Js in Ethiopia: Amharic hate speech data annotation using Toloka Crowdsourcing Platform . In Proceedings of the 4th International Conference on Information and Communication Technology for ...

  3. [3]

    Abinew Ali Ayele, Seid Muhie Yimam, Tadesse Destaw Belay, Tesfa Asfaw, and Chris Biemann. 2023. https://aclanthology.org/2023.ranlp-1.6 Exploring A mharic hate speech data collection and classification approaches . In Proceedings of the 14th International Conference on Recent Advances in Natural Language Processing, pages 49--59, Varna, Bulgaria. INCOMA L...

  4. [4]

    Anatoly Belchikov. 2019. Russian language toxic comments. https://www.kaggle.com/blackmoon/russian-language-toxic-comments https://www.kaggle.com/blackmoon/russian-language-toxic-comments. Accessed: 2023-12-14

  5. [5]

    Kateryna Bobrovnyk. 2019 a . https://ena.lpnu.ua:8443/server/api/core/bitstreams/c4c645c1-f465-4895-98dd-765f862cf186/content Automated building and analysis of ukrainian twitter corpus for toxic text detection . In COLINS 2019. Volume II: Workshop

  6. [6]

    Kateryna Bobrovnyk. 2019 b . The dictionary of ukrainian obscene words. https://github.com/saganoren/obscene-ukr https://github.com/saganoren/obscene-ukr. Accessed: 2024-12-12

  7. [7]

    Aditya Bohra, Deepanshu Vijay, Vinay Singh, Syed Sarfaraz Akhtar, and Manish Shrivastava. 2018. https://doi.org/10.18653/v1/W18-1105 A dataset of H indi- E nglish code-mixed social media text for hate speech detection . In Proceedings of the Second Workshop on Computational Modeling of People ' s Opinions, Personality, and Emotions in Social Media , pages...

  8. [8]

    Eleftheria Briakou, Di Lu, Ke Zhang, and Joel Tetreault. 2021. https://doi.org/10.18653/v1/2021.naacl-main.256 Ol \'a , bonjour, salve! XFORMAL : A benchmark for multilingual formality style transfer . In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 31...

Show all 78 references
  1. [9]

    Keith Carlson, Allen Riddell, and Daniel Rockmore. 2018. https://royalsocietypublishing.org/doi/10.1098/rsos.171920 Evaluating prose style transfer with the bible . Royal Society open science, 5(10):171920

  2. [10]

    Jennifer Cobbe. 2021. Algorithmic censorship by social platforms: Power and resistance. Philosophy & Technology, 34(4):739--766

  3. [11]

    Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzm \' a n, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020. https://doi.org/10.18653/V1/2020.ACL-MAIN.747 Unsupervised cross-lingual representation learning...

  4. [12]

    Costa - juss \` a , James Cross, Onur C elebi, Maha Elbayad, Kenneth Heafield, Kevin Heffernan, Elahe Kalbassi, Janice Lam, Daniel Licht, Jean Maillard, Anna Y

    Marta R. Costa - juss \` a , James Cross, Onur C elebi, Maha Elbayad, Kenneth Heafield, Kevin Heffernan, Elahe Kalbassi, Janice Lam, Daniel Licht, Jean Maillard, Anna Y. Sun, Skyler Wang, Guillaume Wenzek, Al Youngblood, Bapi Akula, Lo \" c Barrault, Gabriel Mejia Gonzalez, Pr...

  5. [13]

    Costa - juss \` a , Mariano Coria Meglioli, Pierre Andrews, David Dale, Prangthip Hansanti, Elahe Kalbassi, Alexandre Mourachko, Christophe Ropers, and Carleigh Wood

    Marta R. Costa - juss \` a , Mariano Coria Meglioli, Pierre Andrews, David Dale, Prangthip Hansanti, Elahe Kalbassi, Alexandre Mourachko, Christophe Ropers, and Carleigh Wood. 2024. https://doi.org/10.48550/ARXIV.2401.05060 Mutox: Universal multilingual audio-based toxicity da...

  6. [14]

    David Dale, Anton Voronov, Daryna Dementieva, Varvara Logacheva, Olga Kozlova, Nikita Semenov, and Alexander Panchenko. 2021. https://doi.org/10.18653/v1/2021.emnlp-main.629 Text detoxification using large pre-trained neural models . In Proceedings of the 2021 Conference on Em...

  7. [15]

    Daryna Dementieva, Nikolay Babakov, and Alexander Panchenko. 2024 a . https://doi.org/10.18653/v1/2024.naacl-short.12 M ulti P ara D etox: Extending text detoxification with parallel data to new languages . In Proceedings of the 2024 Conference of the North American Chapter of...

  8. [16]

    Krotova, Nikita Semenov, Tatiana Shavrina, and Alexander Panchenko

    Daryna Dementieva, Varvara Logacheva, Irina Nikishina, Alena Fenogenova, David Dale, I. Krotova, Nikita Semenov, Tatiana Shavrina, and Alexander Panchenko. 2022. https://api.semanticscholar.org/CorpusID:253169495 RUSSE-2022: Findings of the First Russian Detoxification Shared ...

  9. [17]

    Daryna Dementieva, Daniil Moskovskiy, Nikolay Babakov, Abinew Ali Ayele, Naquee Rizwan, Frolian Schneider, Xintog Wang, Seid Muhie Yimam, Dmitry Ustalov, Elisei Stakovskii, Alisa Smirnova, Ashraf Elnagar, Animesh Mukherjee, and Alexander Panchenko. 2024 b . https://ceur-ws.org...

  10. [18]

    Daryna Dementieva, Daniil Moskovskiy, David Dale, and Alexander Panchenko. 2023. https://doi.org/10.18653/v1/2023.ijcnlp-main.70 Exploring methods for cross-lingual text style transfer: The case of text detoxification . In Proceedings of the 13th International Joint Conference...

  11. [19]

    Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. https://doi.org/10.18653/v1/N19-1423 BERT : Pre-training of deep bidirectional transformers for language understanding . In Proceedings of the 2019 Conference of the North A merican Chapter of the Associat...

  12. [20]

    European Parliament and Council of the European Union . 2022. https://eur-lex.europa.eu/eli/reg/2022/2065/oj Regulation (eu) 2022/2065 of the european parliament and of the council of 19 october 2022 on a single market for digital services (digital services act) and amending d...

  13. [21]

    Fangxiaoyu Feng, Yinfei Yang, Daniel Cer, Naveen Arivazhagan, and Wei Wang. 2022. https://doi.org/10.18653/V1/2022.ACL-LONG.62 Language-agnostic BERT sentence embedding . In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long...

  14. [22]

    Griffin Floto, Mohammad Mahdi Abdollah Pour, Parsa Farinneya, Zhenwei Tang, Ali Pesaranghader, Manasa Bharadwaj, and Scott Sanner. 2023. https://doi.org/10.18653/v1/2023.findings-acl.478 D iffu D etox: A mixed diffusion model for text detoxification . In Findings of the Associ...

  15. [23]

    Robert James Gabriel. 2023. English full list of bad words and top swear words banned by google. https://github.com/coffee-and-fun/google-profanity-words/blob/main/data/en.txt https://github.com/coffee-and-fun/google-profanity-words/blob/main/data/en.txt. Accessed: 2024-12-12

  16. [24]

    Gongane, Mousami V

    Vaishali U. Gongane, Mousami V. Munot, and Alwin D. Anuse. 2024. https://doi.org/10.1007/S42001-024-00248-9 A survey of explainable AI techniques for detection of fake news and hate speech on social media platforms . J. Comput. Soc. Sci., 7(1):587--623

  17. [25]

    Government of India . 2021. https://www.meity.gov.in/writereaddata/files/Intermediary_Guidelines_and_Digital_Media_Ethics_Code_Rules-2021.pdf Information technology (intermediary guidelines and digital media ethics code) rules, 2021 . Ministry of Electronics and Information Te...

  18. [26]

    Hatem Haddad, Hala Mulki, and Asma Oueslati. 2019. https://doi.org/10.1007/978-3-030-32959-4\_18 T-HSAB: A tunisian hate speech and abusive dataset . In Arabic Language Processing: From Theory to Practice - 7th International Conference, ICALP 2019, Nancy, France, October 16-17...

  19. [27]

    Skyler Hallinan, Alisa Liu, Yejin Choi, and Maarten Sap. 2023. https://doi.org/10.18653/v1/2023.acl-short.21 Detoxifying text with M a RC o: Controllable revision with experts and anti-experts . In Proceedings of the 61st Annual Meeting of the Association for Computational Lin...

  20. [28]

    Nhat Hoang, Xuan Long Do, Duc Anh Do, Duc Anh Vu, and Anh Tuan Luu. 2024. https://doi.org/10.18653/v1/2024.naacl-long.359 T o XCL : A unified framework for toxic speech detection and explanation . In Proceedings of the 2024 Conference of the North American Chapter of the Assoc...

  21. [29]

    Zachary Horvitz, Ajay Patel, Chris Callison - Burch, Zhou Yu, and Kathleen R. McKeown. 2024. https://doi.org/10.1609/AAAI.V38I16.29780 Paraguide: Guided diffusion paraphrasers for plug-and-play textual style transfer . In Thirty-Eighth AAAI Conference on Artificial Intelligenc...

  22. [30]

    Imbwaga, Nagaratna B

    Joan L. Imbwaga, Nagaratna B. Chittaragi, and Shashidhar G. Koolagudi. 2024. https://doi.org/10.1007/S10772-024-10135-3 Explainable hate speech detection using LIME . Int. J. Speech Technol., 27(3):793--815

  23. [31]

    Aiqi Jiang, Xiaohan Yang, Yang Liu, and Arkaitz Zubiaga. 2022. https://doi.org/10.1016/J.OSNEM.2021.100182 SWSR: A chinese dataset and lexicon for online sexism detection . Online Soc. Networks Media, 27:100182

  24. [32]

    Jigsaw. 2017. Toxic comment classification challenge. https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge. Accessed: 2024-03-18

  25. [33]

    Di Jin, Zhijing Jin, Zhiting Hu, Olga Vechtomova, and Rada Mihalcea. 2022. https://doi.org/10.1162/coli_a_00426 Deep learning for text style transfer: A survey . Computational Linguistics, 48(1):155--205

  26. [34]

    Md Tawkat Islam Khondaker, Muhammad Abdul - Mageed, and Laks V. S. Lakshmanan. 2024. https://doi.org/10.48550/ARXIV.2402.15951 Greenllama: A framework for detoxification with explanations . CoRR, abs/2402.15951

  27. [35]

    Enes Kulenovi \'c . 2023. https://link.springer.com/article/10.1007/s10677-022-10336-2 Should democracies ban hate speech? hate speech laws and counterspeech . Ethical Theory and Moral Practice, 26(4):511--532

  28. [36]

    Teyun Kwon and Anandha Gopalan. 2021. https://arxiv.org/abs/2112.00819 CO-STAR: conceptualisation of stereotypes for analysis and reasoning . CoRR, abs/2112.00819

  29. [37]

    Chak Tou Leong, Yi Cheng, Jiashuo Wang, Jian Wang, and Wenjie Li. 2023. https://doi.org/10.18653/V1/2023.EMNLP-MAIN.269 Self-detoxifying language models via toxification reversal . In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, EMNLP...

  30. [38]

    Juncen Li, Robin Jia, He He, and Percy Liang. 2018. https://doi.org/10.18653/V1/N18-1169 Delete, retrieve, generate: a simple approach to sentiment and style transfer . In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Lin...

  31. [39]

    Yinhan Liu, Jiatao Gu, Naman Goyal, Xian Li, Sergey Edunov, Marjan Ghazvininejad, Mike Lewis, and Luke Zettlemoyer. 2020. https://doi.org/10.1162/TACL\_A\_00343 Multilingual denoising pre-training for neural machine translation . Trans. Assoc. Comput. Linguistics, 8:726--742

  32. [40]

    Varvara Logacheva, Daryna Dementieva, Sergey Ustyantsev, Daniil Moskovskiy, David Dale, Irina Krotova, Nikita Semenov, and Alexander Panchenko. 2022. https://doi.org/10.18653/v1/2022.acl-long.469 P ara D etox: Detoxification with parallel data . In Proceedings of the 60th Annu...

  33. [41]

    Junyu Lu, Bo Xu, Xiaokun Zhang, Changrong Min, Liang Yang, and Hongfei Lin. 2023. https://aclanthology.org/2023.acl-long.898 Facilitating fine-grained detection of C hinese toxic language: Hierarchical taxonomy, resources, and benchmarks . In Proceedings of the 61st Annual Mee...

  34. [42]

    Lundberg and Su - In Lee

    Scott M. Lundberg and Su - In Lee. 2017. https://proceedings.neurips.cc/paper/2017/hash/8a20a8621978632d76c43dfd28b67767-Abstract.html A unified approach to interpreting model predictions . In Advances in Neural Information Processing Systems 30: Annual Conference on Neural In...

  35. [43]

    Thomas Mandl, Sandip Modha, Prasenjit Majumder, Daksh Patel, Mohana Dave, Chintak Mandlia, and Aditya Patel. 2019. https://doi.org/10.1145/3368567.3368584 Overview of the hasoc track at fire 2019: Hate speech and offensive content identification in indo-european languages . In...

  36. [44]

    Binny Mathew, Punyajoy Saha, Seid Muhie Yimam, Chris Biemann, Pawan Goyal, and Animesh Mukherjee. 2021. https://doi.org/10.1609/AAAI.V35I17.17745 Hatexplain: A benchmark dataset for explainable hate speech detection . In Thirty-Fifth AAAI Conference on Artificial Intelligence,...

  37. [45]

    Jos \' e Mar \' a Molero, Jorge P \' e rez - Mart \' n, \' A lvaro Rodrigo, and Anselmo Pe \ n as. 2023. https://doi.org/10.1109/ACCESS.2023.3310244 Offensive language detection in spanish social media: Testing from bag-of-words to transformers models . IEEE Access , 11:95639--95652

  38. [46]

    Edoardo Mosca, Daryna Dementieva, Tohid Ebrahim Ajdari, Maximilian Kummeth, Kirill Gringauz, Yutong Zhou, and Georg Groh. 2023. https://doi.org/10.18653/v1/2023.ijcnlp-demo.7 IFAN : An explainability-focused interaction framework for humans and NLP models . In Proceedings of t...

  39. [47]

    Daniil Moskovskiy, Daryna Dementieva, and Alexander Panchenko. 2022. https://doi.org/10.18653/v1/2022.acl-srw.26 Exploring cross-lingual text detoxification with large multilingual language models. In Proceedings of the 60th Annual Meeting of the Association for Computational ...

  40. [48]

    Hamdy Mubarak, Kareem Darwish, Walid Magdy, Tamer Elsayed, and Hend Al-Khalifa. 2020. https://aclanthology.org/2020.osact-1.7 Overview of OSACT 4 A rabic offensive language detection shared task . In Proceedings of the 4th Workshop on Open-Source Arabic Corpora and Processing ...

  41. [49]

    Niklas Muennighoff, Thomas Wang, Lintang Sutawika, Adam Roberts, Stella Biderman, Teven Le Scao, M. Saiful Bari, Sheng Shen, Zheng Xin Yong, Hailey Schoelkopf, Xiangru Tang, Dragomir Radev, Alham Fikri Aji, Khalid Almubarak, Samuel Albanie, Zaid Alyafeai, Albert Webson, Edward...

  42. [50]

    Ojha, and Ond r ej Du s ek

    Sourabrata Mukherjee, Akanksha Bansal, Pritha Majumdar, Atul Kr. Ojha, and Ond r ej Du s ek. 2023. https://doi.org/10.18653/v1/2023.banglalp-1.5 Low-resource text style transfer for B angla: Data & models . In Proceedings of the First Workshop on Bangla Language Processing (BL...

  43. [51]

    Ojha, Akanksha Bansal, Deepak Alok, John P

    Sourabrata Mukherjee, Atul Kr. Ojha, Akanksha Bansal, Deepak Alok, John P. McCrae, and Ondrej Dusek. 2024 a . https://doi.org/10.48550/ARXIV.2405.20805 Multilingual text style transfer: Datasets & models for indian languages . CoRR, abs/2405.20805

  44. [52]

    Ojha, and Ondrej Dusek

    Sourabrata Mukherjee, Atul Kr. Ojha, and Ondrej Dusek. 2024 b . https://doi.org/10.48550/ARXIV.2406.05885 Are large language models actually good at text style transfer? CoRR, abs/2406.05885

  45. [53]

    Hala Mulki and Bilal Ghanem. 2021. https://aclanthology.org/2021.wanlp-1.16 Let-mi: An A rabic L evantine T witter dataset for misogynistic language . In Proceedings of the Sixth Arabic Natural Language Processing Workshop, pages 154--163, Kyiv, Ukraine (Virtual). Association ...

  46. [54]

    Hala Mulki, Hatem Haddad, Chedi Bechikh Ali, and Halima Alshabani. 2019. https://doi.org/10.18653/v1/W19-3512 L - HSAB : A L evantine T witter dataset for hate speech and abusive language . In Proceedings of the Third Workshop on Abusive Language Online, pages 111--118, Floren...

  47. [55]

    Cicero Nogueira dos Santos, Igor Melnyk, and Inkit Padhi. 2018. https://doi.org/10.18653/v1/P18-2031 Fighting offensive language on social media with unsupervised text style transfer . In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (...

  48. [56]

    OpenAI. 2022. https://openai.com/blog/chatgpt Chatgpt: Optimizing language models for dialogue . Accessed: 2024-05-31

  49. [57]

    Juan Carlos Pereira - Kohatsu, Lara Quijano S \' a nchez, Federico Liberatore, and Miguel Camacho - Collados. 2019. https://doi.org/10.3390/S19214654 Detecting and monitoring hate speech in twitter . Sensors, 19(21):4654

  50. [58]

    Juan Manuel P \'e rez, Dami \'a n Ariel Furman, Laura Alonso Alemany, and Franco M. Luque. 2022. https://aclanthology.org/2022.lrec-1.785 R o BERT uito: a pre-trained language model for social media text in S panish . In Proceedings of the Thirteenth Language Resources and Eva...

  51. [59]

    Matt Post. 2018. https://doi.org/10.18653/V1/W18-6319 A call for clarity in reporting BLEU scores . In Proceedings of the Third Conference on Machine Translation: Research Papers, WMT 2018, Belgium, Brussels, October 31 - November 1, 2018 , pages 186--191. Association for Comp...

  52. [60]

    Shuofei Qiao, Yixin Ou, Ningyu Zhang, Xiang Chen, Yunzhi Yao, Shumin Deng, Chuanqi Tan, Fei Huang, and Huajun Chen. 2023. https://doi.org/10.18653/v1/2023.acl-long.294 Reasoning with language model prompting: A survey . In Proceedings of the 61st Annual Meeting of the Associat...

  53. [61]

    Sudha Rao and Joel Tetreault. 2018. https://doi.org/10.18653/v1/N18-1012 Dear sir or madam, may I introduce the GYAFC dataset: Corpus, benchmarks and metrics for formality style transfer . In Proceedings of the 2018 Conference of the North A merican Chapter of the Association ...

  54. [62]

    why should I trust you?

    Marco T \' u lio Ribeiro, Sameer Singh, and Carlos Guestrin. 2016. https://doi.org/10.1145/2939672.2939778 "why should I trust you?": Explaining the predictions of any classifier . In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data M...

  55. [63]

    Julian Risch, Anke Stoll, Lena Wilms, and Michael Wiegand. 2021. https://aclanthology.org/2021.germeval-1.1 Overview of the germeval 2021 shared task on the identification of toxic, engaging, and fact-claiming comments . In Proceedings of the GermEval 2021 Shared Task on the I...

  56. [64]

    Bj \"o rn Ross, Michael Rist, Guillermo Carbonell, Benjamin Cabrera, Nils Kurowsky, and Michael Wojatzki. 2016. https://d-nb.info/1119886848/34#page=12 Measuring the Reliability of Hate Speech Annotations: The Case of the European Refugee Crisis . In Proceedings of NLP4CMC III...

  57. [65]

    Sarthak Roy, Ashish Harshvardhan, Animesh Mukherjee, and Punyajoy Saha. 2023. https://doi.org/10.18653/v1/2023.findings-emnlp.407 Probing LLM s for hate speech detection: strengths and vulnerabilities . In Findings of the Association for Computational Linguistics: EMNLP 2023, ...

  58. [66]

    Aleksandr Semiletov. 2020. Toxic Russian Comments: Labelled comments from the popular Russian social network . https://www.kaggle.com/alexandersemiletov/toxic-russian-comments https://www.kaggle.com/alexandersemiletov/toxic-russian-comments. Accessed: 2023-12-14

  59. [67]

    Zheyuan Ryan Shi, Claire Wang, and Fei Fang. 2020. https://arxiv.org/abs/2001.01818 Artificial intelligence for social good: A survey . CoRR, abs/2001.01818

  60. [68]

    Inc Shutterstock. 2020. List of dirty, naughty, obscene, and otherwise bad words. https://github.com/LDNOOBW/List-of-Dirty-Naughty-Obscene-and-Otherwise-Bad-Words https://github.com/LDNOOBW/List-of-Dirty-Naughty-Obscene-and-Otherwise-Bad-Words. Accessed: 2024-12-12

  61. [69]

    Chandan Singh, Jeevana Priya Inala, Michel Galley, Rich Caruana, and Jianfeng Gao. 2024. https://doi.org/10.48550/ARXIV.2402.01761 Rethinking interpretability in the era of large language models . CoRR, abs/2402.01761

  62. [70]

    Yuqing Tang, Chau Tran, Xian Li, Peng - Jen Chen, Naman Goyal, Vishrav Chaudhary, Jiatao Gu, and Angela Fan. 2020. https://arxiv.org/abs/2008.00401 Multilingual translation with extensible multilingual pretraining and finetuning . CoRR, abs/2008.00401

  63. [71]

    Mariona Taul \' e , Montserrat Nofre, V \' ctor Bargiela, and Xavier Bonet Casals. 2024. https://doi.org/10.1007/S10579-023-09711-X Newscom-tox: a corpus of comments on news articles annotated for toxicity in spanish . Lang. Resour. Evaluation, 58(4):1115--1155

  64. [72]

    Rachel Ung. 2023. https://waseda.repo.nii.ac.jp/record/2000931/files/t5121FG17.pdf Formality Style Transfer between Japanese and English . Ph.D. thesis, Waseda University

  65. [73]

    Shang Wang, Tianqing Zhu, Bo Liu, Ming Ding, Xu Guo, Dayong Ye, Wanlei Zhou, and Philip S. Yu. 2024. https://doi.org/10.48550/ARXIV.2406.07973 Unique security and privacy threats of large language model: A comprehensive survey . CoRR, abs/2406.07973

  66. [74]

    Michael Wiegand, Melanie Siegel, and Josef Ruppenhofer. 2018. https://epub.oeaw.ac.at/?arp=0x003a10d2 Overview of the GermEval 2018 Shared Task on the Identification of Offensive Language . In Proceedings of GermEval 2018, 14th Conference on Natural Language Processing (KONVEN...

  67. [75]

    Yifan Yao, Jinhao Duan, Kaidi Xu, Yuanfang Cai, Eric Sun, and Yue Zhang. 2023. https://doi.org/10.48550/ARXIV.2312.02003 A survey on large language model (LLM) security and privacy: The good, the bad, and the ugly . CoRR, abs/2312.02003

  68. [76]

    Chiyu Zhang, Honglong Cai, Yuezhang Li, Yuexin Wu, Le Hou, and Muhammad Abdul-Mageed. 2024. https://doi.org/10.18653/v1/2024.naacl-srw.21 Distilling text style transfer with self-explanation from LLM s . In Proceedings of the 2024 Conference of the North American Chapter of th...

  69. [77]

    online" 'onlinestring :=

    ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list...

  70. [78]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.