Pith. sign in

REVIEW 3 major objections 6 minor 30 references

BIDWESH: A Bangla Regional Based Hate Speech Detection Dataset

T0 review · 3 major / 6 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read A new 9,183-instance dataset targets hate speech in three Bangla dialects, built by adapting the standard-Bangla BD-SHS corpus.

desk verdict Real gap, plausible idea, but the manuscript as submitted doesn't establish the dataset's validity: the statistics don't sum, the labels appear inherited, and the data isn't accessible. read the letter →

arxiv 2507.16183 v1 pith:KXN2KLLE submitted 2025-07-22 cs.CL

classification cs.CL
keywords hatespeechdetectionBangladialectsdatasetlow-resourceNLPoffensivelanguagedialectadaptationBD-SHSannotation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Regional dialects of Bangla—Barishal, Noakhali, and Chittagong—carry offensive idioms that standard-language hate speech detectors miss. The paper's claim is that BIDWESH, a parallel corpus of 9,183 manually translated sentences built from the BD-SHS benchmark, fills that gap as the first multi-dialect Bangla hate speech dataset. Each of 3,061 source comments was rendered into all three dialects by native speakers, then validated and annotated with hate/non-hate, hate type, and target labels aligned to the original schema. If the corpus holds up, it gives moderators and NLP researchers a way to build and test dialect-sensitive detection rather than relying on standard Bangla alone.

What carries the argument

The load-bearing object is a parallel seed-and-translate corpus: 3,061 balanced comments from the BD-SHS dataset in standard Bangla, each rewritten by native speakers into the Barishal, Noakhali, and Chittagong dialects so that every dialectal sentence is aligned to the same source text and the same labels. This design turns an existing standard-Bangla hate speech benchmark into a dialectal one without collecting new social-media comments, and it is what enables direct cross-dialect comparison of offensive language. The annotation schema (binary hate, hate type, target group) is inherited from BD-SHS and manually verified for each dialectal version, giving every sentence a label and every hate instance a type and a target.

What would settle it

Take a random sample of, say, 300 dialectal sentences from BIDWESH, have native speakers who have never seen the standard-Bangla sources annotate them for hate presence, hate type, and target, and compare their labels against the inherited ones; if agreement falls well below the level expected for the original BD-SHS annotations, the inherited-label assumption fails.

Watch

Extended reading notes

Core claim

The paper presents BIDWESH as the first multi-dialect Bangla hate speech dataset: a balanced, manually translated parallel corpus of 9,183 sentences spanning Barishal, Noakhali, and Chittagong. The corpus is built from 3,061 comments taken from the BD-SHS benchmark, with each comment rewritten into all three dialects by native speakers and passed through five validation stages (re-evaluation, unrelated-text removal, sentence verification, emoji removal, and annotation). Each dialectal sentence is labeled for hate presence (hate versus non-hate), and each hate instance additionally receives a hate-type label (slander, gender, religion, call to violence, including multi-label combinations such as callToViolence_slander) and a target label (individual, male, female, group, including combinations such as male_female). The hate/non-hate split is 1,513 versus 1,548 per dialect, and each dialect contributes exactly one third of the corpus. The design deliberately inherits its annotation schema from BD-SHS so that the dialectal versions remain directly comparable to the standard-Bangla source.

Load-bearing premise

The load-bearing premise is that translating a sentence from standard Bangla into a regional dialect does not change whether it is hateful, who it targets, or what kind of hate it expresses, because every dialectal version keeps the labels of its standard-Bangla source.

Editorial extensions

If this is right

  • Dialect-aware hate speech detectors can be trained and evaluated separately for Barishal, Noakhali, and Chittagong, rather than treating all Bangla as one language variety.
  • Because the corpus is parallel, models can be tested for cross-dialect transfer, and dialect-normalization or translation systems can be trained to map regional hate expressions back to standard Bangla.
  • The balanced 1,513/1,548 hate/non-hate split and the multi-label type and target annotations allow separate benchmarks for binary hate detection, hate-type classification, and target identification on the same resource.
  • Content moderation pipelines in Bangladesh gain a reference set of dialect-specific slurs and idioms, addressing the paper's stated problem that regional hate speech is systematically under-detected.
  • The 13 type categories and 7 target categories expose overlap among slander, gender, religion, and calls to violence in actual dialectal comments, which simpler binary datasets cannot capture.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A direct test of the corpus is to sample dialectal sentences, have independent native speakers label them without seeing the standard-Bangla source, and measure agreement with the inherited labels; disagreement would reveal where translation shifted offensiveness, target, or hate type.
  • Because the design is parallel across dialects, the same seed sentence appears three times, which makes it possible to quantify how much a hate expression changes across dialects and to build dialect normalization or translation systems from the corpus.
  • The same inherited-label pipeline could be applied to other under-resourced Bangla dialects, such as Sylhet, Mymensingh, or Rangpur, and to other standard-language hate datasets, though the validation stages would need to be repeated for each new dialect.
  • A practical extension is measuring how standard-Bangla hate classifiers degrade on the dialectal test sets; if degradation is large, the dataset directly supports the paper's motivation that regional hate speech is currently under-detected.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper introduces BIDWESH, a claimed first multi-dialect Bangla hate speech dataset, constructed by translating 3,061 samples from the BD-SHS corpus into three regional dialects (Barishal, Noakhali, Chittagong), yielding 9,183 instances. Each instance is labeled for binary hate presence, hate speech type (13 categories with multi-label combinations), and target group (7 categories), with a roughly balanced hate/non-hate split. The authors describe a multi-stage validation pipeline, translator demographics, and data statistics, and state the dataset is publicly available on Mendeley. The central claim is that BIDWESH provides a valid, balanced, and first-of-its-kind resource for dialect-aware hate speech detection in Bangla.

Significance. If the dataset is valid and accessible, it would fill a genuine gap: existing Bangla hate speech resources (e.g., BD-SHS, BANTH) target standard or transliterated Bangla, not regional dialects. The paper's positive features include the use of native-speaker translators from each region, a parallel-corpus design that enables dialectal comparison, a balanced binary split, and a multi-level annotation scheme that goes beyond simple binary hate labels. The authors also cite prior dialectal resources (VASHANTOR, ANUBHUTI) and position their contribution relative to that landscape. However, the significance is conditional on the validity of the dialectal labels and on the availability of the data; the manuscript as written does not yet establish these, and the internal statistics contain inconsistencies that undermine confidence in the dataset's construction.

major comments (3)
  1. [§3.2, §3.4, §4] The central validity question is whether the hate labels, types, and targets are established for the dialectal text rather than inherited from the source. Section 3.2 states that translated outputs were 'validated to ensure alignment with the original annotation schema of the BD-SHS dataset,' and Section 4 states that annotation was performed by the same translators who produced the dialectal text. There is no independent re-annotation, no inter-annotator agreement measure, and no adjudication process reported. Because the same people who translated the text also labeled it, and because they were explicitly validating against the source schema, the labels are effectively carried over from Standard Bangla with the assumption that translation preserved hate force, type, and target. That assumption is load-bearing: if a dialectal rendering changes the offensiveness or referent of a sentence, the dataset labels are wrong for that dialectal version. The manuscript should report a sample-based re-annotation by independent annotators (and ideally per dialect), with agreement statistics, or otherwise demonstrate that label transfer is justified.
  2. [§5.3, Table 5, §5.4, Table 6] The reported type and target statistics are internally inconsistent. Table 4 reports 1,513 hate instances. Table 5's type counts sum to 1,702, not 1,513, and the listed percentages sum to approximately 111%. Table 6's target counts sum to 1,693, not 1,513, with percentages summing to approximately 112%. If type and target labels are multi-label (as the 13 type categories and the presence of combinations such as 'Gender_slander' suggest), then the row 'Total 1,513 100%' is wrong and the percentages should be presented with respect to instances or labels explicitly. If the labels are intended to be mutually exclusive, then the counts themselves are arithmetically wrong. Either way, the statistics as presented cannot be used to support the claimed 'stable data distribution' and balanced corpus, and the authors should correct the tables and clarify the annotation semantics.
  3. [§8] The dataset is the core contribution, but Section 8 provides only a hyperlink to a Mendeley landing page with a descriptive title, not a direct download link, DOI, or access instructions. Without the data, the existence, the exact label distributions, and the dialectal examples in Table 3 cannot be verified, and the internal inconsistencies noted above cannot be resolved by inspection. The paper should provide a resolver-stable identifier (e.g., a DOI) or an explicit link to the downloadable file, and ideally a data card or sample of the release.
minor comments (6)
  1. [Abstract] In the abstract, 'Existing datasets and systems' capitalizes the first word of a sentence fragment after a comma; please lowercase 'existing'.
  2. [§3.1] The selection of the 3,061-sample subset from BD-SHS is described only as 'balanced selection.' The authors should state the sampling procedure (e.g., random, stratified) and whether the 1,513/1,548 split was constructed intentionally or emerged from the sample.
  3. [§3.4.3] The description of manual spelling verification acknowledges the lack of standardized orthography but does not specify how disagreements among native speakers were resolved; a brief statement on adjudication would clarify the validation process.
  4. [Table 3] The transliteration in Table 3 is difficult to interpret and may contain rendering artifacts; ensure a clear convention is documented for transliteration and script representation.
  5. [Figures 4 and 5] The category labels on the horizontal axes are truncated (e.g., 'callToViolence_sl', 'Gender_religion_'), making the figures unusable in standalone form; please rotate labels or use a legend.
  6. [§7] The ethics statement says the data 'does not raise ethical concerns' because it is publicly available, but hate speech datasets often carry risks even when derived from public sources (e.g., reidentification of speakers, usage restrictions). The authors should at least acknowledge these considerations.

Circularity Check

0 steps flagged · score 2.0 of 10

No load-bearing circular step; inherited BD-SHS labels and non-load-bearing self-citations are the only mild concerns.

full rationale

BIDWESH is a translation dataset rather than a derived prediction: no fitted parameter, equation, or model output is claimed, so there is no derivation chain that can collapse into its own inputs. The construction does rely on BD-SHS labels for the three dialectal variants: Section 3.2 says translated outputs were validated 'to ensure alignment with the original annotation schema of the BD-SHS dataset,' and Section 3.4.1 says reviewers ensured 'the dialectal adaptations preserved the original meaning and hate speech classification.' That is a genuine auditability and validity concern, since the dialectal labels are not independently established and the internal totals in Tables 5 and 6 do not sum cleanly to the stated 1,513 hate instances, but it is not circularity under the definition used here: the source labels are external inputs, not outputs generated by the paper's own argument. The self-citations (Vashantor, Ancholik-NER, EmoNoBa) support the motivation of a research gap and are not load-bearing for the correctness of BIDWESH. No step reduces to itself by construction, so no circular step is entered.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The central claim rests on the source labels being correct and on translation preserving hate semantics. These are domain assumptions stated in the paper but not empirically validated.

assumptions (3)
  • domain assumption Hate speech semantics are preserved when translating standard Bangla into Barishal, Noakhali, and Chittagong.
    Section 3.2 assumes translation keeps the offensive content and label valid; no independent verification.
  • domain assumption BD-SHS annotations are accurate for the source sentences.
    The dataset is used as ground truth; no error analysis is reported.
  • domain assumption The three chosen dialects represent the relevant Bangla regional variation for hate speech.
    Section 3.2 justifies selection by 'notable differences' but provides no coverage or empirical support.

how reviews work

0 comments
Cite this review

Pith. "Pith review of BIDWESH: A Bangla Regional Based Hate Speech Detection Dataset." pith.science (2026). https://pith.science/paper/KXN2KLLE

@misc{pith2026250716183,
  author       = {Pith},
  title        = {Pith review of: BIDWESH: A Bangla Regional Based Hate Speech Detection Dataset},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/KXN2KLLE}},
  note         = {Machine review of arXiv:2507.16183}
}
read the original abstract

Hate speech on digital platforms has become a growing concern globally, especially in linguistically diverse countries like Bangladesh, where regional dialects play a major role in everyday communication. Despite progress in hate speech detection for standard Bangla, Existing datasets and systems fail to address the informal and culturally rich expressions found in dialects such as Barishal, Noakhali, and Chittagong. This oversight results in limited detection capability and biased moderation, leaving large sections of harmful content unaccounted for. To address this gap, this study introduces BIDWESH, the first multi-dialectal Bangla hate speech dataset, constructed by translating and annotating 9,183 instances from the BD-SHS corpus into three major regional dialects. Each entry was manually verified and labeled for hate presence, type (slander, gender, religion, call to violence), and target group (individual, male, female, group), ensuring linguistic and contextual accuracy. The resulting dataset provides a linguistically rich, balanced, and inclusive resource for advancing hate speech detection in Bangla. BIDWESH lays the groundwork for the development of dialect-sensitive NLP tools and contributes significantly to equitable and context-aware content moderation in low-resource language settings.

Figures

Figures reproduced from arXiv: 2507.16183 by the authors.

Figure 1
Figure 1. Examples of hate and non-hate sentences across Bangla Language [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. A Systematic Pipeline for the development of BIDWESH Dataset [PITH_FULL_IMAGE:figures/full_fig_p007_2.png] view at source ↗
Figure 3
Figure 3. Regional Distribution of Dialectal Instances [PITH_FULL_IMAGE:figures/full_fig_p011_3.png] view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: Type Classification 5.4 Target Classification The target classification identifies seven distinct categories, representing the diverse targeting mechanisms employed in hate speech across the three regional dialects [PITH_FULL_IMAGE:figures/full_fig_p012_4.png]
Figure 5
Figure 5. Figure 5: Target Classification 6 Dataset Characteristics The BIDWESH dataset exhibits several key characteristics that distinguish it as a valuable resource for hate speech detection research in regional Bangla languages: • Regional Diversity: Equal representation from Barishal…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

30 extracted references · 29 canonical work pages

  1. [1]

    Automated hate speech detection and the problem of offensive language

    Thomas Davidson, Dana Warmsley, Michael Macy, and Ingmar Weber. Automated hate speech detection and the problem of offensive language. In Proceedings of the international AAAI conference on web and social media, volume 11, pages 512–515, 2017

  2. [2]

    Deep learning for hate speech detection in tweets

    Pinkesh Badjatiya, Shashank Gupta, Manish Gupta, and Vasudeva Varma. Deep learning for hate speech detection in tweets. In Proceedings of the 26th international conference on World Wide Web companion, pages 759–760, 2017

  3. [3]

    Vashantor: a large-scale multilingual benchmark dataset for automated translation of bangla regional dialects to bangla language

    Fatema Tuj Johora Faria, Mukaffi Bin Moin, Ahmed Al Wase, Mehidi Ahmmed, Md Rabius Sani, and Tashreef Muhammad. Vashantor: a large-scale multilingual benchmark dataset for automated translation of bangla regional dialects to bangla language. arXiv preprint arXiv:2311.11142, 2023

  4. [4]

    Ancholik-ner: A benchmark dataset for bangla regional named entity recognition

    Bidyarthi Paul, Faika Fairuj Preotee, Shuvashis Sarker, Shamim Rahim Refat, Shifat Islam, Tashreef Muhammad, Mohammad Ashraful Hoque, and Shahriar Manzoor. Ancholik-ner: A benchmark dataset for bangla regional named entity recognition. arXiv preprint arXiv:2502.11198, 2025

  5. [5]

    Anubhuti: A comprehensive corpus for sentiment analysis in bangla regional languages

    Swastika Kundu, Autoshi Ibrahim, Mithila Rahman, and Tanvir Ahmed. Anubhuti: A comprehensive corpus for sentiment analysis in bangla regional languages. arXiv preprint arXiv:2506.21686, 2025

  6. [6]

    BD-SHS: A Benchmark Dataset for Learning to Detect Online Bangla Hate Speech in Different Social Contexts

    Nauros Romim, Mosahed Ahmed, Md Saiful Islam, Arnab Sen Sharma, Hriteshwar Talukder, and Moham- mad Ruhul Amin. Bd-shs: A benchmark dataset for learning to detect online bangla hate speech in different social contexts. arXiv preprint arXiv:2206.00372, 2022

  7. [7]

    Banglahatebert: Bert for abusive language detection in bengali

    Md Saroar Jahan, Mainul Haque, Nabil Arhab, and Mourad Oussalah. Banglahatebert: Bert for abusive language detection in bengali. In Proceedings of the second international workshop on resources and techniques for user information in abusive language analysis, pages 8–15, 2022

  8. [8]

    Interpretable multi labeled bengali toxic comments classification using deep learning

    Tanveer Ahmed Belal, GM Shahariar, and Md Hasanul Kabir. Interpretable multi labeled bengali toxic comments classification using deep learning. In 2023 International Conference on Electrical, Computer and Communication Engineering (ECCE), pages 1–6. IEEE, 2023

Show all 30 references
  1. [9]

    Banth: A multi-label hate speech detection dataset for transliterated bangla

    Fabiha Haider, Fariha Tanjim Shifat, Md Farhan Ishmam, Deeparghya Dutta Barua, Md Sakib Ul Rahman Sourove, Md Fahim, and Md Farhad Alam. Banth: A multi-label hate speech detection dataset for transliterated bangla. arXiv preprint arXiv:2410.13281, 2024. 14 A PREPRINT - S EPTEM...

  2. [10]

    Analyzing emotions in bangla social media comments using machine learning and lime

    Bidyarthi Paul, SM Rahman, Dipta Biswas, Md Ziaul Hasan, and Md Zahid Hossain. Analyzing emotions in bangla social media comments using machine learning and lime. arXiv preprint arXiv:2506.10154, 2025

  3. [11]

    Improving bangla regional dialect detection using bert, llms, and xai

    Bidyarthi Paul, Faika Fairuj Preotee, Shuvashis Sarker, and Tashreef Muhammad. Improving bangla regional dialect detection using bert, llms, and xai. In 2024 IEEE International Conference on Computing, Applications and Systems (COMPAS), pages 1–6. IEEE, 2024

  4. [12]

    L-boost: Identifying offensive texts from social media post in bengali

    Muhammad Firoz Mridha, Md Anwar Hussen Wadud, Md Abdul Hamid, Muhammad Mostafa Monowar, Mohammad Abdullah-Al-Wadud, and Atif Alamri. L-boost: Identifying offensive texts from social media post in bengali. Ieee Access, 9:164681–164699, 2021

  5. [13]

    Bengali hate speech detection from social media using ensemble machine learning approach

    Sadia Tarin, Farzina Akther, Pranta Paul, and Tanvinur Rahman Siam. Bengali hate speech detection from social media using ensemble machine learning approach. 2025

  6. [14]

    G-bert: an efficient method for identifying hate speech in bengali texts on social media

    Ashfia Jannat Keya, Md Mohsin Kabir, Nusrat Jahan Shammey, Muhammad Firoz Mridha, Md Rashedul Islam, and Yutaka Watanobe. G-bert: an efficient method for identifying hate speech in bengali texts on social media. IEEE Access, 11:79697–79709, 2023

  7. [15]

    Bengali multi-class text classification via enhanced contrastive learning techniques

    Farhana Hossain Swarnali, Jannatim MaishaL, Muhammad Azmain Mahtab, M Saymon Islam Iftikar, and Faisal Muhammad Shah. Bengali multi-class text classification via enhanced contrastive learning techniques. In 2024 27th International Conference on Computer and Information Technol...

  8. [16]

    Hate speech detection: a comparison of mono and multilingual trans- former model with cross-language evaluation

    Koyel Ghosh and Apurbalal Senapati. Hate speech detection: a comparison of mono and multilingual trans- former model with cross-language evaluation. In Proceedings of the 36th Pacific Asia Conference on Language, Information and Computation, pages 853–865, 2022

  9. [17]

    Bengalihatecb: A hybrid deep learning model to identify bengali hate speech detection from online platform

    Sagor Kumar Saha, Afrina Akter Mim, Sanzida Akter, Md Mehraz Hosen, Arman Habib Shihab, and Md Hu- maion Kabir Mehedi. Bengalihatecb: A hybrid deep learning model to identify bengali hate speech detection from online platform. In 2024 6th International Conference on Electrical...

  10. [18]

    Deephateexplainer: Explainable hate speech detection in under-resourced bengali language

    Md Rezaul Karim, Sumon Kanti Dey, Tanhim Islam, Sagor Sarker, Mehadi Hasan Menon, Kabir Hossain, Md Azam Hossain, and Stefan Decker. Deephateexplainer: Explainable hate speech detection in under-resourced bengali language. In 2021 IEEE 8th international conference on data scie...

  11. [19]

    Multimodal hate speech detection from bengali memes and texts

    Md Rezaul Karim, Sumon Kanti Dey, Tanhim Islam, Md Shajalal, and Bharathi Raja Chakravarthi. Multimodal hate speech detection from bengali memes and texts. In International Conference on Speech and Language Technologies for Low-resource Languages, pages 293–308. Springer, 2022

  12. [20]

    Stop the hate, spread the hope: An ensemble model for hope speech detection in english and dravidian languages

    Deepawali Sharma, Vedika Gupta, Vivek Kumar Singh, and Bharathi Raja Chakravarthi. Stop the hate, spread the hope: An ensemble model for hope speech detection in english and dravidian languages. ACM Transactions on Asian and Low-Resource Language Information Processing, 2025

  13. [21]

    Enhancing hate speech detection in the digital age: A novel model fusion approach leveraging a comprehensive dataset

    Waqas Sharif, Saima Abdullah, Saman Iftikhar, Daniah Al-Madani, and Shahzad Mumtaz. Enhancing hate speech detection in the digital age: A novel model fusion approach leveraging a comprehensive dataset. IEEE Access, 12:27225–27236, 2024

  14. [22]

    Hate or non-hate: Translation based hate speech identification in code-mixed hinglish data set

    Shankar Biradar, Sunil Saumya, and Arun Chauhan. Hate or non-hate: Translation based hate speech identification in code-mixed hinglish data set. In 2021 IEEE international conference on big data (Big Data), pages 2470–2475. IEEE, 2021

  15. [23]

    byteSizedLLM@NLU of Devanagari script languages 2025: Hate speech detection and target identification using customized attention BiLSTM and XLM-RoBERTa base embeddings

    Rohith Gowtham Kodali, Durga Prasad Manukonda, and Daniel Iglesias. byteSizedLLM@NLU of Devanagari script languages 2025: Hate speech detection and target identification using customized attention BiLSTM and XLM-RoBERTa base embeddings. In Kengatharaiyer Sarveswaran, Ashwini V...

  16. [24]

    Herdphobia: A dataset for hate speech against fulani in nigeria

    Saminu Mohammad Aliyu, Gregory Maksha Wajiga, Muhammad Murtala, Shamsuddeen Hassan Muhammad, Idris Abdulmumin, and Ibrahim Said Ahmad. Herdphobia: A dataset for hate speech against fulani in nigeria. arXiv preprint arXiv:2211.15262, 2022

  17. [25]

    From languages to geographies: Towards evaluating cultural bias in hate speech datasets

    Manuel Tonneau, Diyi Liu, Samuel Fraiberger, Ralph Schroeder, Scott A Hale, and Paul Röttger. From languages to geographies: Towards evaluating cultural bias in hate speech datasets. arXiv preprint arXiv:2404.17874, 2024

  18. [26]

    Hate speech detection in low-resourced indian languages: An analysis of transformer-based monolingual and multilingual models with cross-lingual experiments

    Koyel Ghosh and Apurbalal Senapati. Hate speech detection in low-resourced indian languages: An analysis of transformer-based monolingual and multilingual models with cross-lingual experiments. Natural Language Processing, 31(2):393–414, 2025. 15 A PREPRINT - S EPTEMBER 10, 2025

  19. [27]

    Generalizable multilingual hate speech detection on low resource Indian languages using fair selection in federated learning

    Akshay Singh and Rahul Thakur. Generalizable multilingual hate speech detection on low resource Indian languages using fair selection in federated learning. In Kevin Duh, Helena Gomez, and Steven Bethard, editors, Proceedings of the 2024 Conference of the North American Chapte...

  20. [28]

    Investigating the predominance of large language models in low-resource bangla language over transformer models for hate speech detection: A comparative analysis

    Fatema Tuj Johora Faria, Laith H Baniata, and Sangwoo Kang. Investigating the predominance of large language models in low-resource bangla language over transformer models for hate speech detection: A comparative analysis. Mathematics, 12(23):3687, 2024

  21. [29]

    Hate speech and offensive language detection in bengali

    Mithun Das, Somnath Banerjee, Punyajoy Saha, and Animesh Mukherjee. Hate speech and offensive language detection in bengali. arXiv preprint arXiv:2210.03479, 2022

  22. [30]

    Evaluation of hate speech detection using large language models and geographical contextualization

    Anwar Hossain Zahid, Monoshi Kumar Roy, and Swarna Das. Evaluation of hate speech detection using large language models and geographical contextualization. arXiv preprint arXiv:2502.19612, 2025. 16

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.