{"id":"b894f84c-8441-4ca8-94cc-85fa4a4dd4ea","arxiv_id":"2411.18577","paper_version":1,"verdict":"REJECT","confidence":"LOW","novelty_score":3.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"HingBERT and Hing-FastText are reported to beat vanilla BERT and FastText on Hinglish hate speech detection, but the paper provides no reproducible evidence.","lead":"This paper compares English BERT and FastText with Hindi-English code-mixed versions (HingBERT, Hing-FastText) for hate speech detection on two datasets. It claims the code-mixed models win, but the writeup lacks the details needed to verify that result.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The empirical advantage of Hing models over BERT is unverifiable because the manuscript reports no preprocessing pipeline or classifier hyperparameters; an uncontrolled comparison could make the Hing advantage a tuning artifact.","rationale":"The central claim is entirely empirical: code-mixed pre-training improves hate speech detection. The paper provides no numbers, no configuration details, and an empty preprocessing section; the only evidence is two unlabeled figures. The single most load-bearing assumption is therefore that all models went through the same preprocessing and classifier tuning, because if the authors experimented with different settings per model, the ranking can reflect engineering choices rather than pretraining data. This is not a manufactured objection: Section 4.2 is literally blank, Section 4.4 lists several classifiers without selection criteria, and Section 4.3 says 'mostly max pooling' rather than a fixed pooling scheme. The absence of error bars or repeated runs makes it impossible to distinguish real gains from noise. The reader's verdict of REJECT is appropriate, not because the paper is necessarily wrong, but because it fails to provide the evidence needed to support the comparison. The proposed test - a single controlled rerun with a shared protocol - would directly settle whether the concern lands. If the Hing advantage persists under identical settings and with significance testing, the claim would be supported; if not, the rejection stands. Hence I agree with the reader and leave the verdict unchanged.","tokens_in":4769,"tokens_out":6107,"duration_ms":51846,"concrete_test":"Obtain the missing experimental configuration (Sections 4.2 and 4.4) and re-run the full comparison with one shared preprocessing pipeline, one fixed pooling choice, and one identical classifier hyperparameter grid for BERT, mBERT, HingBERT, and both FastText variants; if the reported Hing advantage over standard models disappears or falls within noise, the central claim is an artifact of uncontrolled settings rather than code-mixed pre-training.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is an empirical comparison (Abstract: 'HingBERT models... outperform BERT models'; 'Hing-FastText performs better than standard English FastText and vanilla BERT models'). The load-bearing assumption is that every compared model was evaluated under the same, hermetic protocol. This is not established anywhere in the text. Section 4.2 ('Preprocessing') is empty, and Section 4.4 lists candidate classifiers ('Random Forest, Linear Regression, SVM and KNN') without stating which were used, their hyperparameters, or any validation procedure. Section 5 reports only figures ('testing scores are shown in Figures 2 and 3') with no numeric F1, precision, recall, or confidence intervals. Because the models themselves have different tokenizers and hidden sizes, small variations in pooling (Section 4.3 says 'mostly max pooling is used') or in SVM penalty/kernel settings can reverse rankings. Without a single fixed preprocessing and training configuration applied identically to all models, the reported Hing advantage cannot be attributed to code-mixed pre-training rather than to tuning or implementation differences. This is a missing-support problem, not an internal logical contradiction; the claim might survive a rigorous re-run, but the paper as submitted does not demonstrate it.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript compares BERT-family transformers and FastText embeddings, including code-mixed Hing models developed by L3Cube, for hate speech detection on the HASOC 2021 Hindi-English dataset and a smaller in-house HATE dataset. The central claim is that HingBERT variants and Hing-FastText outperform standard English BERT and FastText models on this task. The paper describes the models, gives a brief literature survey, outlines a methodology that includes embedding generation and classification, and concludes with the comparative advantage of code-mixed models.","tokens_in":5008,"tokens_out":2820,"duration_ms":25902,"significance":"If the empirical claims are correct, the paper would provide useful evidence that code-mixed pre-training (L3Cube-HingCorpus) improves hate speech detection for Hindi-English text, a socially important task. The use of the external HASOC 2021 benchmark avoids an obviously circular evaluation, and the range of compared models is appropriate. However, the manuscript as submitted does not report the actual numerical results, does not describe the preprocessing or classifier configuration, and leaves the central comparison unverifiable. The significance is therefore conditional on the authors supplying the missing experimental evidence.","major_comments":[{"comment":"The 'Preprocessing' subsection is empty. Since the paper's central claim is an empirical comparison, the preprocessing pipeline is load-bearing: different normalization, tokenization, handling of Devanagari script, emoji, URL, and case-folding choices can affect transformer and FastText models differently. The manuscript must state exactly what preprocessing was applied and confirm that it was identical for all compared models.","section":"Section 4.2"},{"comment":"The classification protocol is underspecified. Section 4.4 lists Random Forest, Linear Regression, SVM, and KNN as candidate algorithms but does not state which were actually used, what hyperparameters were chosen, how train/validation/test splits were formed, or whether any hyperparameter tuning was performed. Section 5 adds that 'mostly max pooling is used' without defining when pooling differs. Without a fixed, fully specified protocol applied identically to all models, the reported Hing advantage could be a tuning artifact rather than an effect of code-mixed embeddings.","section":"Section 4.4 and Section 5"},{"comment":"Results are presented only as aggregate figures ('testing scores are shown in Figures 2 and 3') with no numeric F1, precision, recall, or accuracy values for any model or dataset. This makes the abstract's claims—that HingBERT 'outperform BERT models' and that Hing-FastText 'performs better than standard English FastText and vanilla BERT models'—impossible to check or reproduce. The authors should include tables of exact numbers, with confidence intervals or standard deviations if multiple runs were performed.","section":"Section 5"},{"comment":"The HATE dataset is not sourced, described, or released. The tables give category counts for only 'a portion' of the dataset, and the total size, collection method, annotation procedure, and license are absent. Without provenance and full statistics, the results cannot be reproduced, and it is unclear whether the dataset is public or whether the evaluation set overlaps with the pre-training corpus. These details are necessary to rule out data leakage and to allow independent verification.","section":"Section 4.1, Tables 1 and 2"}],"minor_comments":[{"comment":"The text states that BERT-base has 345 million parameters; BERT-base actually has about 110 million parameters, while BERT-large has about 340 million. This should be corrected together with the associated citation.","section":"Section 2.1"},{"comment":"The manuscript refers to 'Fig. 1. Flowchart' and to 'Figures 2 and 3' as comparison charts, but none of these figures appear in the text. The figures must be included, or the references removed.","section":"Figures"},{"comment":"The abstract and conclusion assert that Hing models outperform BERT and vanilla FastText, but no numerical evidence is given anywhere in the paper. The claims should be backed by the reported numbers or softened to what the results actually support.","section":"Abstract and Conclusion"},{"comment":"The description of HingCorpus says it contains '52.93 million sentences and 1.04 billion tags'; it is unclear whether 'tags' is a typo for 'tokens', and it is not specified which subset (Roman, Devanagari, or both) was used for each Hing model in this study.","section":"Section 2.4"},{"comment":"Several in-text citations appear mismatched with the reference list: for example, reference [10] is described as a character-level GRU hate speech paper, but the reference list entry [10] is the Sabty et al. Arabic-English NER paper. The entire bibliography should be checked against the citations.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"I recommend major revision rather than reject because the central claim is plausible and the deficiencies—empty preprocessing section, unspecified classifiers, missing numeric results, and unsourced dataset—are in principle fixable by adding the missing experimental details. However, if the authors cannot supply those details, including exact numbers and a fully specified protocol, the paper would not be publishable in its current form. The authors' affiliation with L3Cube and the prominence of self-citations are not themselves a flaw, but they make it especially important that the evaluation protocol be transparent and that the comparison be demonstrably hermetic."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nHere's the short version: this paper wants to show that code-mixed HingBERT beats English BERT for Hinglish hate speech detection, and that result may well be true—but the manuscript as written doesn't let you check it. Section 4.2 (Preprocessing) is empty. Section 4.4 lists Random Forest, Linear Regression, SVM and KNN as candidate classifiers without saying which were used or with what hyperparameters. Results exist only as figures, with no numeric F1, precision, recall, or confidence intervals. The HATE dataset is never sourced. BERT-base is described as having 345 million parameters; the real number is 110 million. This is not a paper you can evaluate; it's a slide deck with an abstract.\n\nWhat's good: the question is a fair one, and the comparison to HASOC 2021 is a sensible external benchmark. The Hing-FastText versus vanilla FastText comparison is a small extension beyond earlier L3Cube work, and the authors cite the key prior papers (Nayak and Joshi 2021, Farooqi et al. 2021) that already compared mBERT and Indic-BERT on the same task. So the framing is honest about where the result comes from.\n\nThe soft spot is the missing support, not the research direction. The stress-test note is right: since HingBERT and BERT use different tokenizers and hidden sizes, small choices in pooling or SVM settings could reverse the ranking. Without a fixed preprocessing and training protocol applied to all models, the reported advantage could be a tuning artifact. The authors' own lab built the Hing models, which isn't circular because the benchmark is external, but it makes full reporting more important, not less.\n\nProportionately: this is a reject-for-incompleteness, not a reject-for-wrongness. The claim could survive a rigorous re-run.\n\nRecommendation: desk reject in current form. If the authors resubmit with the preprocessing details, the actual classifier and hyperparameters, and per-model numbers, it becomes a legitimate workshop paper. Not worth referee time as is.","headline":"Plausible HingBERT-over-BERT result, but the manuscript hides every detail that would let you check it—empty preprocessing section, no hyperparameters, no numeric results, and a wrong BERT parameter count.","tokens_in":5577,"tokens_out":3054,"would_cite":false,"duration_ms":25961,"reading_group":"no","serious_thinker":"unclear","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that code-mixed Hindi-English embeddings (HingBERT and Hing-FastText) trained on L3Cube-HingCorpus outperform monolingual BERT and English FastText for hate speech detection on HASOC 2021 and an in-house HATE dataset.","keywords":["code-mixed text","hate speech detection","HingBERT","Hing-FastText","Hinglish","word embeddings","HASOC","BERT"],"falsifier":"Run HingBERT, Hing-FastText, BERT, and English FastText through one identical preprocessing and hyperparameter search on a held-out Hinglish hate speech test set that is verified disjoint from L3Cube-HingCorpus; if the code-mixed models no longer beat the English-only ones on F1, the paper's central claim is disproved.","tokens_in":4577,"feed_emoji":"🛡️","tokens_out":7663,"duration_ms":58926,"temperature":0.7,"pith_summary":"Code-mixed Hindi-English (Hinglish) is common in Indian social media, but most NLP models are trained on monolingual English and struggle with it. This paper tries to establish that embeddings pre-trained on a large code-mixed Hindi-English corpus are better for hate speech detection than standard English-only embeddings. Using the HASOC 2021 hate speech dataset and an in-house HATE dataset, the authors report that HingBERT and Hing-FastText achieve higher F1 scores, recall, and accuracy than BERT and English FastText. If the claim holds, the practical lesson is that code-mixed pre-training, not architecture alone, is what makes hate speech models work in multilingual communities.","feed_headline":"HingBERT beats BERT at finding hate speech in Hinglish","feed_subtitle":"Code-mixed pre-training on Hindi-English text lifts F1, recall, and accuracy on HASOC 2021 and a private HATE set.","key_machinery":"The central object is the code-mixed embedding. HingBERT is a transformer-based language model pre-trained on the L3Cube-HingCorpus with contextual, bidirectional representations; Hing-FastText is a subword-level embedding model trained on the same corpus in Roman and Devanagari scripts. Both carry vocabulary and syntax from actual Hinglish text, which is what the paper credits for the measured hate speech detection advantage over models trained only on English.","core_discovery":"The central claim is that code-mixed representations are the decisive ingredient for hate speech identification in Hinglish: HingBERT and Hing-FastText, both trained on the L3Cube-HingCorpus (52.93 million sentences of Roman and Devanagari Hindi-English text), outperform standard BERT, multilingual BERT, and vanilla FastText on two hate speech test sets, HASOC 2021 and the authors' HATE dataset. The paper reports the code-mixed models winning on F1, recall, and accuracy, and attributes the gain to their ability to handle mixed-code context, subword information, and the real-world vocabulary of Hinglish social media.","pith_inferences":["The paper's own evidence would be more convincing if the preprocessing and classifier hyperparameters were disclosed, since Section 4.2 is empty and Section 4.4 lists no settings; a controlled replication with identical settings would test whether the reported advantage survives.","The same code-mixed pre-training recipe could transfer to other language pairs, such as Spanish-English or Arabic-French, where social media text mixes languages and monolingual models underperform.","The reported results also suggest that code-mixed embeddings may improve other Hinglish social media tasks, such as sentiment analysis, misinformation detection, or identifying in-group slurs that shift meaning across languages.","A natural next experiment is an ablation that trains HingBERT on only the Roman subset or only the Devanagari subset of L3Cube-HingCorpus to see which portion drives the hate speech gain."],"forward_implications":["If correct, hate speech detection systems for Hinglish should start from code-mixed pre-trained models like HingBERT rather than English BERT.","The reported results make Hing-FastText a lightweight alternative that beats vanilla BERT, which matters for deployment where transformer compute is limited.","The findings imply that the value of a multilingual corpus is not just language coverage but alignment with the code-mixed register of the target task.","The HASOC 2021 and HATE scores serve as baselines for future work on code-mixed hate speech detection.","The advantage of HingBERT over multilingual BERT suggests that a dedicated code-mixed pre-training corpus is more useful than a broad multilingual Wikipedia-trained model."],"supporting_citations":[{"why":"Supplies the HASOC 2021 code-mixed hate speech data and a transformer-based baseline (F1 73.07) that the paper compares against.","marker":"[1]"},{"why":"Gives a second code-mixed Hinglish hate speech baseline using transformer ensembles (macro F1 0.7253), which motivates the model comparison.","marker":"[2]"},{"why":"Compares pretrained embeddings for hate speech in Indian code-mixed text, the direct precedent for evaluating BERT and FastText on this task.","marker":"[5]"}],"fun_headline_variants":["Code-mixed embeddings win for hate speech detection","HingBERT beats BERT on Hinglish hate speech","Code-mixed embeddings boost hate speech identification","Why code-mixed embeddings matter for hate speech","Hinglish hate speech? Code-mixed embeddings help"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The central claim presupposes a fair and fully specified comparison—identical preprocessing and classifier tuning across all models, with no overlap between the L3Cube-HingCorpus pre-training data and the HASOC or HATE test sets—yet the paper does not report the preprocessing details or classifier settings.","fun_headline_variants_meta":{"raw":{"variants":["Code-mixed embeddings win for hate speech detection","HingBERT beats BERT on Hinglish hate speech","Code-mixed embeddings boost hate speech identification","Why code-mixed embeddings matter for hate speech","Hinglish hate speech? Code-mixed embeddings help"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00075,"raw_usage":{"total_tokens":3294,"prompt_tokens":851,"completion_tokens":2443,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":467,"completion_tokens_details":{"reasoning_tokens":2368}},"tokens_in":467,"tokens_out":2443,"duration_ms":15109,"temperature":1.0,"reasoning_tokens":2368,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T11:03:32.192403+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run HingBERT, Hing-FastText, BERT, and English FastText through one identical preprocessing and hyperparameter search on a held-out Hinglish hate speech test set that is verified disjoint from L3Cube-HingCorpus; if the code-mixed models no longer beat the English-only ones on F1, the paper's central claim is disproved.","supporting_citations":[{"cited_title":"Contextual Hate Speech Detection in Code Mixed Text using Transformer Based Approaches","cited_arxiv_id":null,"evidence_quote":"Supplies the HASOC 2021 code-mixed hate speech data and a transformer-based baseline (F1 73.07) that the paper compares against."},{"cited_title":"Leveraging Transformers for Hate Speech Detection in Conversational Code-Mixed Tweets","cited_arxiv_id":null,"evidence_quote":"Gives a second code-mixed Hinglish hate speech baseline using transformer ensembles (macro F1 0.7253), which motivates the model comparison."},{"cited_title":"Comparison of Pretrained Embeddings to Identify Hate Speech in Indian Code-Mixed Text,","cited_arxiv_id":null,"evidence_quote":"Compares pretrained embeddings for hate speech in Indian code-mixed text, the direct precedent for evaluating BERT and FastText on this task."}],"review_version":1}