{"id":"30dfae6a-c26b-4b11-a2f9-7170998a5b31","arxiv_id":"2412.17533","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A new 24,768-sentence Polish dataset labels erotic and related content, and Polish-specific language models outperform multilingual models at detecting it.","lead":"This paper introduces forePLay, a Polish-language dataset of over 24,000 sentences labeled for erotic, ambiguous, violent, socially unacceptable, and neutral content. It then benchmarks Polish and multilingual AI models on detecting these categories, finding Polish-specific models generally do better.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Final labels are not a stable ground truth: the outlier annotator Fem1, whose pairwise kappa is 0.14–0.18, is retained in the majority vote, so the benchmark scores in Tables 5–6 may measure label noise.","rationale":"The reader's weakest assumption—that aggregated labels are reliable enough to serve as ground truth—is the same concern I identify as most load-bearing. The paper's own numbers make this concrete: raw Krippendorff alpha is 0.387, Fem1's pairwise kappas are 0.14–0.18, and Fem1 is nevertheless retained in the majority vote. Since every item is annotated by only three people, one unreliable annotator can systematically bias the final labels whenever the other two disagree. This is not a peripheral issue: the benchmark tables, the Polish-vs-multilingual superiority claim, and the multiclass degradation pattern all depend on those labels. I considered the reproducibility problem (only 3,704 of 24,768 sentences are released, with neutral and rare classes omitted) as an alternative concern, but that affects usability and verification, whereas label instability would invalidate the empirical conclusions even if all artifacts were public. The proposed test—recomputing labels without Fem1 and rerunning the evaluations—directly settles whether the central benchmark conclusions are an artifact of the outlier. Because the reader already conditioned acceptance on robustness to label noise, my analysis does not change the verdict; it sharpens the specific mechanism and the required check.","tokens_in":17603,"tokens_out":6971,"duration_ms":69285,"concrete_test":"Release the raw per-annotator labels for the full 24,768 sentences and recompute the aggregated labels with Fem1 excluded. Measure the flip rate, i.e., the fraction of final labels that change when Fem1's votes are removed from the majority vote. Then rerun the RoBERTa and PLLuM-SFT evaluations from Tables 5–6 on the revised labels. If the flip rate exceeds about 5% or any macro-F1 in the Core/Extended/Full settings moves by more than about 0.05, the benchmark conclusions are not robust to the outlier annotator's inclusion.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central empirical claim—that specialized Polish models outperform multilingual alternatives on forePLay—rests on majority-vote labels that Section 4.2 shows are not stable. Krippendorff's alpha is 0.387 overall, and reaches 0.716 only after dropping Fem1, whose pairwise kappas against the other annotators are 0.14–0.18, with near-zero agreement on the ambiguous class (0.015–0.053 in Table 4). Despite this, Fem1's votes are included in the majority aggregation, and only the 830 fully split triples go to a superannotator. Because each sentence is labeled by exactly three annotators, Fem1's unreliable vote is decisive whenever the other two disagree, and her over-use of the ambiguous label can create an ambiguous majority or force a three-way tie. The final labels for the ambiguous class (5.43% of the data) and for the rare violence/unacceptable classes are therefore partly determined by an annotator whose judgments are nearly uncorrelated with the rest of the team. The paper's Limitations section concedes that 'significant variation in individual annotator interpretations—particularly evident in the use of the ambiguous category—suggests potential instability in ground truth labels.' If the labels are unstable, the macro-F1 scores in Tables 5 and 6, and the Polish-vs-multilingual ranking drawn from them, are measurements of a noisy target. The observed degradation from binary to multiclass performance could be an artifact of the ill-defined ambiguous class rather than a real property of the models. A related internal tension supports this concern: Appendix B states that each sentence was presented in isolation without context, while the ambiguous label is defined as context-related, making the label hard to apply consistently.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces forePLay, a Polish-language dataset of 24,768 sentences annotated for erotic content detection, with a five-class taxonomy (erotic, ambiguous, violence-related, socially unacceptable, neutral). The authors document the data collection from online fiction repositories and literary works, the annotation process with six annotators, and the aggregation via majority vote with a superannotator for three-way ties. They report inter-annotator agreement (Krippendorff's alpha 0.387 overall, 0.716 after excluding one outlier annotator, Fem1) and benchmark a wide range of Polish-specific and multilingual models on binary, three-class, four-class, and five-class classification tasks. The central claims are that forePLay is the first Polish manually annotated dataset of erotic content and that specialized Polish language models outperform multilingual alternatives on this benchmark.","tokens_in":17813,"tokens_out":4213,"duration_ms":39107,"significance":"If the dataset labels are reliable, forePLay is a valuable resource: it addresses a genuine gap for a morphologically complex, non-English language, provides a multidimensional taxonomy rather than a simple binary, includes LGBTQ+ representation, and offers transparent documentation of the annotation process and limitations. The paper also provides a broad empirical comparison including Polish encoder models, Polish LLMs, and multilingual LLMs, with detailed error analyses. The authors are unusually candid about annotation difficulties and label instability, which is a strength. However, the central empirical claim rests on ground-truth labels whose stability is questionable, and the reported aggregate agreement is low; this tempers the significance of the benchmark comparisons until the label reliability issue is addressed.","major_comments":[{"comment":"The final labels are determined by majority vote over three annotators, but the statistical evidence in Section 4.2 and Table 4 shows that one annotator (Fem1) had pairwise Cohen's kappa values of only 0.14–0.18 against all other annotators, with near-zero agreement on the ambiguous class (0.015–0.053). Because each sentence is labeled by exactly three annotators, Fem1's unreliable vote is decisive whenever the other two annotators disagree, and her documented over-use of the ambiguous label can create an ambiguous majority or force a three-way tie. The overall Krippendorff's alpha of 0.387 is low, and it rises to 0.716 only when Fem1 is excluded. The paper does not quantify how many final labels are determined by Fem1's vote, nor does it report the distribution of label changes if Fem1's annotations were removed from the aggregation. This is a load-bearing issue for the ground-truth labels used in all subsequent experiments, and it is acknowledged in the Limitations section as 'potential instability in ground truth labels' without being mitigated. I request an analysis of the stability of the final labels under alternative aggregation rules (e.g., excluding Fem1, or using soft labels), and if the instability is substantial, the experiments should be re-run on a cleaner label set.","section":"Section 4 and Section 4.2, Table 4"},{"comment":"The central empirical claim—that specialized Polish language models achieve superior performance compared to multilingual alternatives—is supported by macro-F1 scores computed against the majority-vote labels just described. If those labels are not a stable ground truth, the scores in Tables 5 and 6, and the ranking drawn from them, may be measuring noise rather than detection ability. This concern is especially acute for the ambiguous class, which has the lowest annotator agreement (Table 4) and which is exactly the class whose addition causes the sharp performance drop from the Basic to the Core configuration. The paper argues that the degradation reflects the difficulty of finer-grained distinctions, but it could equally be an artifact of an ill-defined label. To support the central claim, the authors should report at least one sensitivity analysis: for example, macro-F1 on the subset of sentences with full annotator agreement, or re-run the main comparisons using labels aggregated without Fem1. Without such evidence, the superiority claim is not yet established.","section":"Section 6, Tables 5 and 6"},{"comment":"The released dataset (Release 1.0) contains only 3,704 erotic and ambiguous sentences, which is 15% of the full dataset; neutral sentences and the rare violence/unacceptable classes are excluded. However, the experiments in Tables 5 and 6 use the full 24,768-sentence dataset, including the rare classes and the neutral majority. As a result, the benchmark results cannot be reproduced or independently verified from the public release. This is a significant limitation for a resource paper whose main contribution is a dataset: the authors should either release the full label distribution (even without the text, to address copyright and ethical concerns) or provide a clear protocol for reconstructing the full dataset, and should state explicitly that the public release is only a subset and that the reported results pertain to the full, non-public dataset.","section":"Section 9 and Section 3"}],"minor_comments":[{"comment":"The row for PLLuM-Mistral-12B (SFT) in the 1-shot condition contains only three numeric entries instead of four, which appears to be a formatting error: the table should be checked and corrected.","section":"Appendix D, Table 11"},{"comment":"The term 'Type I error percentage' is defined as FP/(FP+FN), which is not a conventional Type I error rate but rather the false positive proportion among all errors; this definition should be stated more prominently in the text and the figure caption should be adjusted accordingly.","section":"Section 7, Figure 1"},{"comment":"The prompt template does not include the label definitions, and the Limitations section correctly notes this omission; however, the main text describing the LLM evaluation (Section 5.3 and 5.4) should mention this design choice and its potential impact on the reported LLM scores, rather than relegating it only to the Limitations section.","section":"Appendix C, Figure 2"},{"comment":"There are inconsistent decimal formats in the table, such as '0.58' alongside '0.580' and '0.42' alongside '0.420'; these should be unified for readability.","section":"Table 6"},{"comment":"The sentence 'for datasets labeled as Extended and Core, which are marked by pronounced class imbalance' contains a typo ('asExtended') and is also somewhat imprecise, since Core has a modest class imbalance compared to Extended and Full; consider rewording.","section":"Section 6"},{"comment":"Several references use nonstandard author formatting, such as 'cjadams' and 'inversion' in the Jigsaw corpus entry; these should be converted to the journal's citation style.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper is a serious resource contribution and the authors are transparent about their annotation challenges, which is commendable. However, the central empirical claim hinges on the reliability of the majority-vote labels, and the paper's own agreement statistics reveal a significant risk that the labels are unstable. I believe the paper is within scope for a major revision: the authors could address this by (a) quantifying the impact of Fem1 on the final labels, (b) re-running key experiments with a cleaner aggregation or with soft labels, and (c) releasing full label distributions or a reproducibility protocol. I would not reject the paper outright, because the resource itself is potentially valuable and the limitations are honestly stated. I would also encourage the editor to weigh whether the 'first Polish dataset' claim is adequately verified, though this is not a blocking issue for me."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Read this if you work on non-English content moderation: it is a real dataset paper with honest limitations, but don't trust the headline benchmark numbers until the label-noise question is settled. The genuinely new thing is forePLay, the first manually annotated Polish erotic-content dataset: 24,768 sentences, five labels (erotic, ambiguous, violence-related, socially unacceptable, neutral), drawn from amateur fiction repositories and published literature with deliberate LGBTQ+ representation. That fills an obvious gap, and the paper does several things right: the taxonomy is explicit, the annotation process is described in detail, and the discussion of human label variation is unusually candid. The Limitations section openly says the ambiguous category creates potential instability in ground truth.\n\nThe soft spots are concentrated in one place: the ground truth. Krippendorff's alpha is 0.387 overall; it only reaches 0.716 after excluding annotator Fem1, whose pairwise kappas are 0.14–0.18 and whose agreement on the ambiguous class is near zero (0.015–0.053). Fem1's votes were still included in the majority vote. With only three annotators per sentence, this is not harmless: whenever the other two disagree, Fem1 decides. The ambiguous class, and the rare violence/unacceptable classes, are partly defined by the person whose judgments are nearly uncorrelated with everyone else. The benchmark macro-F1 scores in Tables 5 and 6 are therefore at risk of measuring label noise, not model skill. The paper's own Appendix B makes this worse: sentences were presented in isolation, yet the ambiguous label is defined context-dependently.\n\nOther issues are minor or acknowledged: the prompt omitted label definitions (they admit this), the Polish-vs-multilingual comparison mixes model sizes and prompting setups, and the public release excludes neutral and the rare classes, so the central evaluation cannot be fully reproduced from the GitHub files.\n\nNone of this kills the paper. The dataset is usable, the documentation is transparent, and the authors clearly know the limits. What I'd want before treating the benchmark as a result: release full labels (or at least per-annotator labels), and run a robustness check dropping Fem1 and re-reporting the tables. If the Polish-vs-multilingual ranking survives that, the claim is solid. As it stands, I'd send it to peer review with major revision conditions rather than desk reject it. The resource deserves to exist; the empirical claim needs to be re-grounded.","headline":"A real Polish erotic-content dataset with honest documentation, but the benchmark claims rest on labels that may be too unstable to trust until the outlier annotator is removed and the results re-checked.","tokens_in":18469,"tokens_out":2076,"would_cite":true,"duration_ms":20337,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper introduces forePLay, the first Polish manually annotated dataset of erotic content—24,768 sentences labeled across five categories—and reports that Polish-specific models beat multilingual alternatives on every label…","keywords":["forePLay","Polish language","erotic content detection","annotated dataset","content moderation","inter-annotator agreement","language-specific models","morphologically complex languages"],"falsifier":"Re-annotate a random sample of roughly 500 forePLay sentences with a fresh annotator pool using the same guidelines, then measure agreement between the new labels and the published majority-vote labels; if agreement for the ambiguous category falls near chance, the ground truth that all model comparisons rest on is not stable. Alternatively, retrain the same models on per-annotator labels rather than majority votes—if the ranking of Polish-specific versus multilingual models flips, the paper's central comparison is an artifact of label aggregation.","tokens_in":17335,"feed_emoji":"🔞","tokens_out":7181,"duration_ms":58661,"temperature":0.7,"pith_summary":"The paper is trying to establish that Polish-language erotic content can be detected reliably by models trained on a new resource, forePLay, the first manually annotated Polish dataset of erotic discourse. The dataset uses a five-way taxonomy—erotic, ambiguous, violence-related, socially unacceptable, and neutral—that the authors argue captures the context-dependence and moral dimensions that binary schemes miss. Their evaluations show that specialized Polish models (HerBERT, Polish RoBERTa, PLLuM, Bielik) outperform multilingual and general-purpose systems on this benchmark, which they take as evidence that content moderation for morphologically complex languages needs language-specific resources. The paper also documents substantial disagreement among annotators, so part of its contribution is a candid account of how unstable the ground truth is for this task.","feed_headline":"Polish-trained models win on Polish erotic-text detection","feed_subtitle":"forePLay labels 24,768 sentences in five categories; Polish-specific models beat multilingual rivals on every task.","key_machinery":"The load-bearing object is the annotation scheme: a five-way exclusive taxonomy (erotic, ambiguous, violence-related, socially unacceptable, neutral) with a fixed priority order for overlapping categories and a deliberately separated ambiguous class for context-dependent erotic connotations. This scheme produces the final labels through majority voting with a superannotator resolving total disagreements, and every model score in the paper is a measurement of that aggregated ground truth.","core_discovery":"The central discovery is forePLay itself: a dataset of 24,768 Polish sentences drawn from online fiction repositories and published literature, annotated by six raters with majority vote and a superannotator breaking three-way ties. The five exclusive labels follow a fixed priority order (socially unacceptable outranks violence-related, which outranks erotic), and the separate 'ambiguous' label is designed for sentences whose erotic reading depends on context. On this benchmark the paper reports that Polish-specialized encoder models reach macro-F1 scores of 0.929–0.944 in binary classification, with Polish RoBERTa leading HerBERT, and that fine-tuned PLLuM-Mistral-12B reaches 0.946. Polish-specific models consistently beat multilingual baselines such as GPT-4o, Llama 3.1, Mixtral, and Command-R, although all models lose accuracy as the number of classes grows, dropping to roughly 0.66 on the five-class task.","pith_inferences":["The paper's own agreement statistics suggest the majority-vote labels may be too noisy to serve as a stable ground truth; a natural test is to re-annotate a random sample with a fresh annotator pool and measure agreement with the published labels.","If the ranking of Polish-specific versus multilingual models were computed on per-annotator labels or soft labels instead of majority votes, the reported advantage might change, since the aggregated labels are where the noise is concentrated.","Because 69% of the corpus comes from amateur online fiction, the models' edge on this benchmark may not transfer to other Polish registers such as chat, social media, or professional prose, which the evaluation does not directly test.","A direct extension would be to train the same models on disagreement-weighted objectives and test whether the Polish-model advantage survives when the target is an individual reader's judgment rather than an aggregated label."],"forward_implications":["Content moderation for Polish should use Polish-specific models rather than English-centric multilingual tools, since the paper's comparisons show a consistent advantage for the specialized models on every label configuration.","The taxonomy is a reusable template for erotic-content detection in other morphologically complex languages, provided each language receives its own manually annotated dataset.","The strong performance drop with more classes (from roughly 0.94 in binary to 0.66 in five-class macro-F1) implies that fine-grained moderation needs either more training data for rare categories or a different evaluation strategy such as learning with disagreements.","The public release of 3,704 erotic and ambiguous sentences gives researchers a reproducible benchmark, although the rare violence-related and socially unacceptable classes are excluded from the release for ethical reasons."],"supporting_citations":[{"why":"Supplies HerBERT, one of the two specialized Polish encoder models whose scores anchor the comparison.","marker":"(Mroczkowski et al., 2021)"},{"why":"Supplies Polish RoBERTa, the encoder model the paper finds superior among the specialized transformer baselines.","marker":"(Dadas, 2023)"},{"why":"Supplies Bielik, the Polish-aligned LLM family used in the few-shot and fine-tuned evaluations.","marker":"(Ociepa et al., 2024)"},{"why":"Supplies Mixtral, a strong multilingual open-source baseline the Polish models must beat.","marker":"(Jiang et al., 2024)"},{"why":"Supplies Llama 3.1, the general-purpose instruction-tuned baseline with high refusal rates in the comparison.","marker":"(Dubey et al., 2024)"},{"why":"Supplies GPT-4o, the commercial generalist baseline that performs best among non-Polish models.","marker":"(Hurst et al., 2024)"},{"why":"Provides the human-label-variation framing the paper uses to interpret annotator disagreement and motivate future soft-label work.","marker":"(Plank, 2022)"},{"why":"Offers the learning-with-disagreements approach the paper identifies as the natural next step for the dataset.","marker":"(Uma et al., 2021)"}],"fun_headline_variants":["Polish models beat multilingual rivals on erotic text","forePLay: 24k Polish sentences push erotic detection forward","Polish-specific encoders top erotic discourse benchmark","New forePLay dataset sharpens Polish erotic content detection"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the aggregated majority-vote labels, with a superannotator breaking ties, are reliable enough to serve as ground truth—but the paper's own numbers show a Krippendorff's alpha of only 0.387 before removing the most divergent annotator, so if the labels are unstable every reported model score is a measurement of that instability.","fun_headline_variants_meta":{"raw":{"variants":["Polish models beat multilingual rivals on erotic text","forePLay: 24k Polish sentences push erotic detection forward","Polish-specific encoders top erotic discourse benchmark","New forePLay dataset sharpens Polish erotic content detection"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000154,"raw_usage":{"total_tokens":1162,"prompt_tokens":847,"completion_tokens":315,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":463,"completion_tokens_details":{"reasoning_tokens":252}},"tokens_in":463,"tokens_out":315,"duration_ms":3288,"temperature":1.0,"reasoning_tokens":252,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T05:27:28.258191+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-annotate a random sample of roughly 500 forePLay sentences with a fresh annotator pool using the same guidelines, then measure agreement between the new labels and the published majority-vote labels; if agreement for the ambiguous category falls near chance, the ground truth that all model comparisons rest on is not stable. Alternatively, retrain the same models on per-annotator labels rather than majority votes—if the ranking of Polish-specific versus multilingual models flips, the paper's central comparison is an artifact of label aggregation.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies HerBERT, one of the two specialized Polish encoder models whose scores anchor the comparison."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies Polish RoBERTa, the encoder model the paper finds superior among the specialized transformer baselines."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies Bielik, the Polish-aligned LLM family used in the few-shot and fine-tuned evaluations."}],"review_version":1}