{"id":"f17a33c6-d9c4-4c3e-b1a3-52ca42076d20","arxiv_id":"2505.12160","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":3.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A BERTurk model fine-tuned on TREMO reaches about 92.6% accuracy on six Turkish emotion classes, and its predictions on 'Sessiz Istila' tweets suggest anger and surprise dominate the discourse.","lead":"This paper fine-tunes BERTurk, a Turkish language model, on the TREMO emotion dataset and reports 92.62% accuracy across six emotion categories, then applies the model to about 47,000 Turkish tweets about 'Sessiz Istila' (Silent Invasion). The analysis finds anger and surprise dominate the discourse, with surprise more common in 2021 shifting to anger in 2022.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The headline anger/surprise findings rest on unvalidated model labels; the paper's own admitted anger/fear confusion directly threatens the central claim.","rationale":"The reader's weakest assumption — that TREMO-trained model outputs on out-of-domain political tweets are treated as actual emotions without human validation — is exactly the load-bearing concern I identify. The paper's own Limitations section strengthens this concern by admitting poor anger/fear discrimination, while anger is the headline emotion in the target corpus. The concrete test of human annotation on a target-domain sample would settle whether the transfer holds. I also note the internal inconsistencies in the reported accuracy and the missing reproducibility artifacts, but those are secondary to the domain-transfer gap. Since the reader already issued a CONDITIONAL verdict that requires target-domain validation and reproducible artifacts, my analysis does not change that verdict; it reinforces it. No basis emerges for outright rejection, because the underlying fine-tuning approach is standard and the TREMO accuracy, if properly verified, is plausible. Equally, no basis emerges for acceptance without the requested validation.","tokens_in":11950,"tokens_out":3622,"duration_ms":38783,"concrete_test":"Draw a stratified random sample of roughly 1,000 'sessiz istila' tweets spanning the collection period, including the May–June 2022 peak. Have two or three native Turkish annotators independently label each tweet with the same six emotion categories or 'ambiguous', measure inter-annotator agreement, and compare the ERM's thresholded predictions against the human majority labels. Report per-class precision and recall on this target-domain sample, and recompute the anger and surprise percentages using human labels. If anger precision on the target domain is materially below the TREMO test-set precision (roughly 0.90–0.95) or the human-labeled anger share differs from 43.6% by more than about 10 percentage points, the Section 4.3 temporal findings should be treated as model artifacts rather than discourse facts.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central social-science claim — that 'sessiz istila' discourse is dominated by anger (43.6%) and surprise (33.4%) — rests entirely on labels produced by a TREMO-fine-tuned BERTurk applied to out-of-domain political tweets, with no human-annotated target-domain evaluation reported anywhere in the manuscript. This is not a peripheral issue: Section 5.1 explicitly states that the model 'struggled to differentiate between anger and fear due to semantic overlaps in linguistic expressions.' Anger is the headline category, while fear is the fourth-largest category at 11.2%; if even a modest fraction of the 17,839 anger labels are actually fear, the anger-dominance narrative weakens substantially. The Section 3.5 confidence threshold of 0.6 and the exclusion of ambiguous '-1' predictions also mean Table 4's denominator (40,880) and percentages are not reconciled with the 47,024 collected tweets, so the reported distribution is not robustly defined. The 92.62% TREMO accuracy itself is not independently checkable because no code, data, seeds, or baselines are provided, and the confusion-matrix total of 3,605 does not match the stated 90/10 split on the balanced 18,018-instance dataset; but the decisive gap for the paper's contribution is the missing validation of the model on the target discourse.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a BERTurk-based emotion recognition model fine-tuned on the balanced TREMO dataset for Turkish, claiming 92.62% accuracy across six Ekman emotion categories. The model is then applied to 47,024 Turkish X posts containing the keyword \"sessiz istila\" from June 2021 to December 2022, yielding a reported distribution with anger at 43.6% and surprise at 33.4%, and a shift from surprise-dominated discourse in 2021 to anger-dominated discourse in 2022. The paper presents this as a contribution to low-resource Turkish emotion recognition and to computational social science, with implications for monitoring social sentiment in crisis- and policy-related contexts.","tokens_in":12158,"tokens_out":4209,"duration_ms":39260,"significance":"If the central claims were fully supported, the paper would provide a useful demonstration of fine-tuning a transformer model for Turkish emotion classification and a substantive case study of emotional dynamics in anti-refugee discourse. The strengths are the relevant task, the use of a language-specific pretrained model (BERTurk), and the effort to balance the TREMO dataset. However, several load-bearing issues currently undermine the claims: the reported accuracy is internally inconsistent with the described train/test split, the target-domain emotion counts are not validated against human annotation, and the paper's own limitation statement acknowledges confusion between anger and fear, which is the headline emotion. The paper also does not provide code, data, seeds, or baselines, so the reported figures are not independently checkable. As presented, the contribution is not yet established.","major_comments":[{"comment":"The confusion matrix total of 3,605 predictions is inconsistent with the stated 90/10 split on the balanced 18,018-instance TREMO dataset, which should yield approximately 1,802 test predictions. Moreover, the reported diagonal sum of 3,338 exceeds the entire expected test set, so the computed accuracy of 92.62% cannot be correct as described. The Discussion section also states 92.47% accuracy instead. Please provide the actual test-set size and confusion-matrix dimensions, correct the accuracy calculation, and report all metrics consistently (or make the evaluation code and split indices public so the discrepancy can be resolved).","section":"§3.3, §4.2, Eq. (2)"},{"comment":"The emotion percentages for the target corpus are based on counts that sum to 40,880, whereas the collected dataset is described as 47,024 tweets. Predictions below the 0.6 confidence threshold are labeled as ambiguous (-1) and are excluded from the reported counts, but the number of excluded tweets is never given, and no analysis shows that exclusion is not systematically correlated with particular emotions or time periods. The reported anger (43.6%) and surprise (33.4%) distribution is therefore defined over an uncharacterized subset and cannot be taken as the emotion distribution of the collected discourse without reporting the excluded counts and a sensitivity analysis.","section":"§3.5, §4.3, Table 4"},{"comment":"The headline finding that anger dominates the sessiz istila corpus (43.6%) is directly threatened by the model's admitted difficulty in differentiating anger from fear, as stated in Section 5.1. Fear is the fourth-largest category at 11.2%, so even a modest error rate between these two classes could materially change the reported dominance. The paper provides no target-domain validation: there is no comparison of model labels with human annotations on a sample of sessiz istila tweets, no inter-annotator agreement, and no error analysis on political or Xenophobic discourse. Without such validation, the temporal anger/surprise findings remain possible artifacts of applying a TREMO-trained model to out-of-domain text. Please add a target-domain evaluation with human-annotated samples and report confusion and agreement metrics for those samples.","section":"§4.3, §5.1, Table 4"}],"minor_comments":[{"comment":"The paper refers to \"Sessis Istila\" in the paragraph on Erbaysal Filibeli & Öneren Özbek; this appears to be a typo for \"Sessiz Istila\" and should be corrected.","section":"§2"},{"comment":"The text states that \"fear was recorded 5,581 times\" while Table 4 lists the fear count as 4,581; these numbers must be reconciled.","section":"§4.3"},{"comment":"The data collection is described as occurring on January 26, 2022, yet the corpus extends to December 31, 2022; clarify whether the collection was retrospective through the academic API or whether multiple collection rounds were performed.","section":"§3.1"},{"comment":"The normalization rules in Table 1 are expressed as R code, while the emoji conversion is described as a Python program; the paper would benefit from a unified description of the preprocessing pipeline, including the language/toolchain used at each step.","section":"§3.2"},{"comment":"The confusion matrix in Figure 6 would be much clearer with explicit emotion labels on both axes; currently the reader has to infer the class ordering from the text.","section":"§4.2, Figure 6"},{"comment":"The reference list uses inconsistent name spellings (e.g., \"Tocoglu\" vs. \"Toçoglu\") and inconsistent entry formatting; please unify the style and check that all cited works, including those with DOIs, are listed with complete information.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper addresses an important and timely topic, and the approach is sensible in principle, but the reported numbers contain internal inconsistencies that make the current claims unreliable. The most decisive gap is the absence of any target-domain validation, which is essential for the paper's social-science contribution. I believe these issues are addressable within the scope of a revision, provided the authors can supply corrected metrics and a target-domain evaluation. I would also encourage the editor to request that the authors clearly identify the version of the manuscript being submitted, since the arXiv posting is from May 2025 while the document is dated September 2025."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: the real content here is the application, not the method. The authors fine-tune BERTurk on TREMO and run the resulting classifier over tweets containing \"sessiz istila.\" That is a standard pipeline, and the only genuinely new empirical material is the temporal emotion curve for anti-refugee discourse between June 2021 and December 2022. As a descriptive case study it could be useful to people working on Turkish NLP and on anti-refugee sentiment.\n\nThe paper does several things right. The tweet normalization and emoji translation are described in unusual detail, the monthly and yearly breakdowns are tied to concrete events (the Niğde truck incident, lira devaluation, release of the short film), and the authors engage with the relevant Turkish emotion-analysis literature. The balanced TREMO setup is straightforward and defensible.\n\nThe soft spots are real. First, the load-bearing claim—43.6% anger and 33.4% surprise in Sessiz Istila discourse—rests entirely on labels produced by a TREMO-trained model applied to out-of-domain political tweets. There is no human-annotated validation on the target corpus, no agreement score, and no error analysis on political tweets. The paper itself admits in Section 5.1 that the model struggled to separate anger from fear. Anger is the headline category; fear is 11.2%. Without a target-domain check, the anger-dominance narrative is not established. That is the main problem.\n\nSecond, the numbers do not add up in places. The confusion matrix sums to 3,605 predictions, which is not 10% of the 18,018-instance balanced dataset (that would be roughly 1,802). The results section reports 92.62% accuracy, while the discussion says 92.47%. Table 4's denominator of 40,880 is never reconciled with the 47,024 collected tweets, because the thresholded \"-1\" predictions are excluded without a clear accounting. And no code, data, seeds, or baselines are provided, so the \"state of the art\" claim is uncheckable. None of these is fatal alone; together they make the reported numbers hard to trust.\n\nThis is not a methodological advance and should not be published as-is. But it is a serious descriptive effort on a timely topic. With a small target-domain validation, corrected arithmetic, and released artifacts, it could become a useful data point for Turkish emotion analysis and for social scientists tracking anti-refugee sentiment. I would send it to peer review, with a clear request for major revision.","headline":"A standard BERTurk fine-tune with a potentially useful descriptive case study, but the central emotion distribution rests on unvalidated out-of-domain labels and the reported numbers do not add up.","tokens_in":12766,"tokens_out":4160,"would_cite":false,"duration_ms":40348,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Fine-tuning BERTurk on TREMO yields 92.62% Turkish emotion accuracy, and applying it to 'sessiz istila' tweets shows anger and surprise dominating anti-refugee discourse.","keywords":["emotion recognition","Turkish NLP","BERTurk","TREMO","social media analysis","anti-refugee discourse","sessiz istila","sentiment analysis"],"falsifier":"Have Turkish-speaking annotators label a random sample of the 47,024 'sessiz istila' tweets into the same six emotion categories, then compare their labels with the model's confident predictions; if agreement on anger and surprise is near chance, the temporal trends are artifacts of the model rather than public sentiment.","tokens_in":11662,"feed_emoji":"😠","tokens_out":9551,"duration_ms":81632,"temperature":0.7,"pith_summary":"This paper claims that a Turkish-language emotion classifier built by fine-tuning BERTurk on the balanced TREMO dataset can recognize six basic emotions—happiness, fear, anger, sadness, disgust, and surprise—with 92.62% accuracy on a held-out test set. The paper then applies that classifier to 47,024 Turkish X posts containing the phrase 'sessiz istila' ('silent invasion') collected from June 2021 to December 2022. The model labels anger as the dominant emotion (43.6% of confident predictions) and surprise second (33.4%), with surprise leading in 2021 and anger taking over in 2022, peaking after the release of the short film that popularized the term. If the transfer from the training dataset to real political discourse holds, the result matters because it turns a low-resource language's social-media text into measurable emotional signals, with direct uses in monitoring anti-refugee sentiment, polarization, and public reaction to political events.","feed_headline":"92.62% accuracy: Turkish emotion model maps anti-refugee anger","feed_subtitle":"Fine-tuned on TREMO, the model finds anger and surprise dominate 'sessiz istila' tweets from 2021-2022.","key_machinery":"The mechanism is a fine-tuning pipeline: BERTurk, a transformer pretrained on Turkish text, is trained for three epochs on the TREMO emotion dataset, which has been balanced to 3,003 labeled sentences per emotion across six categories. Tweets are normalized by replacing retweets, URLs, mentions, hashtags, and emojis with fixed Turkish tokens, and the classifier's predictions are kept only when the highest softmax probability is at least 0.6, with lower-confidence predictions marked ambiguous. The resulting model is then run over the 47,024-tweet 'sessiz istila' corpus, and the monthly and yearly counts of each emotion are compared.","core_discovery":"The paper's central claim is that fine-tuning BERTurk on a balanced version of the TREMO dataset produces a reliable Turkish emotion-recognition model for six basic emotion categories. On the held-out test set the model reaches 92.62% overall accuracy, with F1 scores per emotion ranging from 0.9091 to 0.9505 and with happiness and disgust easiest to recognize, sadness and anger hardest. Applied to the 'sessiz istila' corpus, the model reports anger in 43.6% and surprise in 33.4% of confident predictions, and it tracks a year-by-year shift: surprise dominates 2021, anger dominates 2022, with the highest tweet volumes and strongest anger appearing in May and June 2022. The paper presents this as evidence that a localized transformer model can capture emotional dynamics in Turkish anti-refugee discourse.","pith_inferences":["Editorial inference: the surprise-to-anger shift could be read as a two-stage public response—initial shock, then sustained moral outrage—but the paper does not test that causal story.","Editorial inference: because the model discards predictions below a 0.6 confidence threshold, the reported anger and surprise percentages apply only to confident predictions; if the ambiguous tweets differ, the true emotion proportions in the full 47,024-post corpus could shift.","Editorial inference: the same fine-tuning recipe could generalize to other morphologically rich low-resource languages that have a small emotion-labeled dataset, though the paper does not claim this."],"forward_implications":["A Turkish emotion classifier at this accuracy level can support real-time monitoring of social-media sentiment for marketing, public relations, and crisis management.","The emotion timeline—surprise in 2021, anger in 2022, peaking with the film's release—shows that counting emotions by month is a practical way to track how anti-refugee discourse intensifies around events.","Balancing TREMO to 3,003 sentences per emotion before fine-tuning is a straightforward recipe for improving emotion classification in low-resource languages with skewed label distributions.","The difficulty the model has separating anger from fear pinpoints where future Turkish emotion models need better contextual representations."],"supporting_citations":[{"why":"Supplies the Turkish emotion dataset (TREMO) with six emotion labels that the model is fine-tuned on.","marker":"[35]"},{"why":"Provides BERTurk, the Turkish pretrained transformer that is fine-tuned into the emotion classifier.","marker":"[30]"},{"why":"Defines the BERT fine-tuning approach and the three-epoch training strategy used here.","marker":"[12]"},{"why":"Defines the six basic emotion categories that structure the label set.","marker":"[14]"},{"why":"The data-collection tool used to obtain the 47,024 Turkish 'sessiz istila' tweets.","marker":"[7]"},{"why":"Supplies the tweet normalization rules that prepare raw posts before classification.","marker":"[28]"}],"fun_headline_variants":["BERTurk tuned on TREMO hits 92.62% on Turkish emotions","Turkish emotion model: anger and surprise dominate 'sessiz istila' tweets","Fine-tuned BERTurk maps anti-refugee anger in Turkish tweets","From TREMO to 'sessiz istila': BERTurk emotion model at 92.62%","92.62% Turkish ERM: anger and surprise shift from 2021 to 2022"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper assumes that the emotions the model learned from the TREMO training sentences are the same emotions expressed in real Turkish tweets about refugees, without human-checked labels on those tweets to confirm it.","fun_headline_variants_meta":{"raw":{"variants":["BERTurk tuned on TREMO hits 92.62% on Turkish emotions","Turkish emotion model: anger and surprise dominate 'sessiz istila' tweets","Fine-tuned BERTurk maps anti-refugee anger in Turkish tweets","From TREMO to 'sessiz istila': BERTurk emotion model at 92.62%","92.62% Turkish ERM: anger and surprise shift from 2021 to 2022"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000594,"raw_usage":{"total_tokens":2780,"prompt_tokens":942,"completion_tokens":1838,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":558,"completion_tokens_details":{"reasoning_tokens":1724}},"tokens_in":558,"tokens_out":1838,"duration_ms":10938,"temperature":1.0,"reasoning_tokens":1724,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T20:39:57.190462+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have Turkish-speaking annotators label a random sample of the 47,024 'sessiz istila' tweets into the same six emotion categories, then compare their labels with the model's confident predictions; if agreement on anger and surprise is near chance, the temporal trends are artifacts of the model rather than public sentiment.","supporting_citations":[{"cited_title":"A., and Alpkocak, A","cited_arxiv_id":null,"evidence_quote":"Supplies the Turkish emotion dataset (TREMO) with six emotion labels that the model is fine-tuned on."},{"cited_title":"BERTurk - BERT models for Turkish (1.0.0).Zenodo., 2020","cited_arxiv_id":null,"evidence_quote":"Provides BERTurk, the Turkish pretrained transformer that is fine-tuned into the emotion classifier."},{"cited_title":"Basic Emotions","cited_arxiv_id":null,"evidence_quote":"Defines the six basic emotion categories that structure the label set."},{"cited_title":"academictwitteR: an R package to access the Twitter Academic Research Product Track v2 API endpoint.Journal of Open Source Software., 6(62), page 3272, 2021","cited_arxiv_id":null,"evidence_quote":"The data-collection tool used to obtain the 47,024 Turkish 'sessiz istila' tweets."},{"cited_title":"Multilingual evaluation of pre-processing for BERT-based sentiment analysis of tweets.Expert Systems with Applications., page 181, 2021","cited_arxiv_id":null,"evidence_quote":"Supplies the tweet normalization rules that prepare raw posts before classification."}],"review_version":1}