{"id":"8105cd0b-bf93-44f6-8e80-841550e19029","arxiv_id":"2412.09203","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"CleanComedy delivers cleaner bilingual joke datasets, but its fine-tuned models produce less funny jokes than GPT-4o or unfiltered human jokes.","lead":"CleanComedy is a new toxicity-filtered dataset of 44,481 English and 40,926 Russian jokes, with human humor scores for 2,000 of them, plus a two-stage fine-tuning recipe for Llama 3.1 8B. The resource is useful for training and evaluating humor models that need to avoid offensive content.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"English CleanComedy evaluation scores rest on a nearly non-native annotator pool, so the English toxicity-halving claim and Gold humor scores lack target-population validity.","rationale":"The paper's central empirical payoff is the toxicity-halving comparison in Table 2. For English, every number in that table is a function of the volunteer pool described in Section 5, which contained almost no native English speakers and which the authors say significantly influenced the English humor scores. The same raters supplied the binary toxicity judgments used to compute 26.09% versus 11.41%. If native English speakers perceive toxicity or humor differently, the half-toxicity claim and the utility of CleanComedy Gold as an English humor benchmark are both in question. This matters because the strongest claim explicitly promises an English and Russian evaluation resource; losing English validity removes half the contribution. The reader flagged annotator skew only in passing and concentrated on threshold proxies; the stress-test makes the annotator population the primary concern. A re-rating study with native speakers would settle it. Until then, CONDITIONAL/UNCHANGED is the right verdict: the dataset artifact remains useful, but the English evaluation evidence is unverified.","tokens_in":746,"tokens_out":1212,"duration_ms":56113,"concrete_test":"Use the released per-joke annotations to identify the exact 100 English jokes in the Clean and Unfiltered groups. Recruit a new panel of at least 30 self-reported native English speakers, plus a comparison panel of non-native speakers, and have them answer the same binary vulnerability/inappropriateness question for the same 200 jokes. Compute each group's toxicity percentage for each panel. If the native panel does not show approximately a halving, e.g. the ratio is not around 0.5 or the difference is within sampling noise, the headline toxicity-halving claim for English fails. Also compare mean humor scores on the Gold subset to see if English humor scores shift systematically.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing support for the central claim is Table 2's toxicity comparison, and for English that comparison is produced by a panel that, by the paper's own admission (Section 5, Figure 10), contained almost no native English speakers and 'significantly influenced' the humor scores. Since the same annotator pool rated toxicity via the binary 'Do you find this text vulnerable or inappropriate?' question, the reported English drop from 26.09% to 11.41%, and the English CleanComedy Gold scores, are measurements of non-native perception rather than of the English-speaking audience the dataset is meant to serve. The claim that CleanComedy roughly halves toxicity is therefore not established for the target population. This is not an internal inconsistency, but an external-validity gap that is directly acknowledged in the paper. The Russian results are less affected, but the English half of the central claim is load-bearing and currently unsupported. Additionally, the 100-sample proportions in Table 2 are reported without confidence intervals or significance tests, so even within this pool the halving could reflect sampling noise.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces CleanComedy, a toxicity-filtered and deduplicated joke dataset in English and Russian, with human humor scores for 1,000 jokes per language (CleanComedy Gold). The curation pipeline combines exact and semantic deduplication, automated toxicity classifiers, zero-shot labeling, and BERTopic cluster removal. The authors then fine-tune Llama-3.1-8B on the clean data and apply an alignment stage that uses the Gold humor scores as soft labels. They evaluate the dataset and the models by sampling 100 jokes from each of six groups (SFT, SFT+Aligned, Instruct, GPT-4o, Unfiltered Dataset, Clean Dataset) and collecting human ratings of humor and toxicity. The central claim is that the clean dataset has about half the toxicity of the unfiltered data, and that the two-stage training yields ethically aligned humor generation. The paper releases the dataset and individual annotations on GitHub.","tokens_in":11279,"tokens_out":6122,"duration_ms":58737,"significance":"If the claims hold, CleanComedy is a useful resource: it provides a reproducible filtering pipeline, a public dataset in two languages, and per-joke human scores that can support personalization research. The explicit release of individual annotations rather than only aggregates is a strength. The paper is also commendably transparent about annotator demographics. However, the central quantitative claims are weakened by the annotator pool for English and by the absence of significance testing; the alignment conclusion is not supported by the reported table. These issues are fixable within the manuscript's scope, so the contribution is potentially valid but needs revision.","major_comments":[{"comment":"The English half of the central toxicity-halving claim is not established for the target population. The paper states that 'there was almost no native speakers of English among our volunteers (Figure 10), which has significantly influenced the humor scores obtained for the English jokes.' Because the same annotators answered the binary 'vulnerable or inappropriate?' question, the English toxicity drop from 26.09% to 11.41% and the Clean Dataset humor score of 2.96 are measurements of non-native perception, not of the English-speaking audience the dataset is intended to serve. The authors should either re-annotate a subsample with native English speakers, or explicitly restrict the claim to the actual annotator population and provide evidence that toxicity and humor judgments are stable across populations. This is an external-validity gap, not an internal inconsistency, but it is load-bearing for the paper's main quantitative contribution.","section":"Section 5, Table 2 and Figure 10"},{"comment":"The toxicity percentages are not defined or tested. The text reports a percentage but does not state how individual 'yes/no' answers were aggregated into a joke-level or group-level toxicity label (e.g., any annotator, majority, or all annotators). In addition, no confidence intervals or significance tests are provided. Under a simple two-proportion z-test, the Russian drop from 20.93% to 11.37% is not significant at the 0.05 level (z ≈ 1.85, p ≈ 0.06), so the claimed 'half' reduction is not statistically supported for Russian. The English drop is significant under the same naive test, but the lack of a defined aggregation rule and the absence of annotation-clustering corrections make all point estimates difficult to interpret. The authors should report the aggregation rule, confidence intervals, and significance tests that account for annotator random effects.","section":"Section 5, Table 2"},{"comment":"The claim that alignment improves generated humor is not supported by the reported data. The alignment stage trains the model to predict average humor scores from CleanComedy Gold, using the same five-point scale and annotation protocol as the evaluation in Section 5. Yet Table 2 shows that the SFT+Aligned model does not outperform the SFT-only model on humor scores in either language (English: 2.02 vs 2.11; Russian: 1.74 vs 1.68), and no significance tests are given. The conclusion that the results 'underscore the importance of alignment techniques in improving the quality and relevance of generated humor' is therefore overstated. The authors should either provide statistical evidence of improvement or temper the conclusion to reflect that alignment mainly reduced Russian toxicity (4.01% to 3.3%) in their sample.","section":"Section 4.2, Eq. (1)-(2); Section 5 and 6"}],"minor_comments":[{"comment":"The text contains two unresolved cross-references to 'Table ?? in the Appendix'; the intended tables appear to be Tables 5 and 6 in the appendix, and the references should be fixed.","section":"Section 3.1"},{"comment":"The paper does not report inter-annotator agreement for the Gold annotations or for the evaluation (e.g., Krippendorff's alpha or ICC). Such measures would help readers interpret the reliability of the humor and toxicity scores.","section":"Section 3.2 and Section 5"},{"comment":"The definition of the toxicity percentage is ambiguous: it is not stated whether a joke is counted as toxic if at least one annotator marks it as vulnerable/inappropriate or only if a majority does. This should be clarified in the text or a footnote.","section":"Section 5"},{"comment":"The procedural difference for the English LLaMA 3.1 8B (Instruct) generation (temperature 0.9 and semantic deduplication) is mentioned in a footnote but not discussed in terms of comparability. The authors should report how many duplicates were removed and justify the asymmetry, or use a common protocol for both languages.","section":"Section 5"},{"comment":"The phrase 'unbiased toxicity classifier' is imprecise, since Detoxify is trained on a particular definition of toxicity and the authors later acknowledge classifier limitations. Consider rewording to avoid overclaiming objectivity.","section":"Section 3.1"},{"comment":"The CleanComedy English row lists a '2-scale score' while the Gold row lists a '5-scale score'; the '2-scale' appears to refer to the binary toxicity annotation, but this is confusing and should be labeled more explicitly.","section":"Table 1"}],"recommendation":"major_revision","confidential_remarks":"The paper is transparent about its annotator demographics, which is a strength, but the English evaluation pool is a serious external-validity issue for the central claim. The absence of significance tests, especially for Russian, is also a concern. The alignment conclusion overreaches the data. All three issues are fixable with additional analysis or careful rewording, so I recommend major revision rather than rejection. The dataset release and individual annotations are valuable contributions."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a real dataset paper, not a breakthrough. The useful thing is CleanComedy, a bilingual English/Russian joke corpus that goes through a documented toxicity/deduplication pipeline, with 1,000 jokes per language rated by five annotators and the individual ratings released. That is genuinely reusable for humor generation and evaluation, and publishing per-annotator data is a nice touch. The models the authors train don't work particularly well, and they are honest about that.\n\nWhat the paper does well: the curation is transparent. They list the filters (Detoxify, ruBERTConv, SBERT dedup, zero-shot label removal, BERTopic cluster pruning), show examples, and give counts. The clean dataset does show roughly half the toxicity of the unfiltered pool in both languages, in their numbers. The Russian side is more credible than the English side.\n\nSoft spots, in order of importance. First, the English evaluation is built on a nearly non-native annotator pool; the paper admits this. That means the English toxicity-halving number (26.09 to 11.41) and the English humor scores measure non-native perception, not the English-speaking audience the dataset targets. The central claim is only established for Russian; for English it is a gap the authors themselves flag. Second, there are no confidence intervals or significance tests on proportions from 100 samples, so the halving could be noise within the panel. Third, the alignment stage uses the same annotation population and protocol for training soft labels as for evaluation, so the claim that alignment learns funniness is partly circular—and Table 2 shows the aligned model is not funnier than the SFT model, so the alignment claims are not supported by the paper's own results. Fourth, the filtering thresholds are post hoc and culturally specific; the dataset is clean by a particular set of proxies, not by a universal standard.\n\nWho is this for: people working on computational humor who want a moderately clean bilingual resource and per-annotator scores. It is a useful contribution, but the evaluation claims need to be read with the annotator caveat in mind. The paper deserves a serious referee, not a desk reject; a reviewer should ask for significance testing and either a native-English re-annotation or a re-scoped claim.","headline":"A genuinely useful bilingual humor dataset, honest about its limits, but the English evaluation rests on non-native annotators and the alignment claims are weaker than the paper suggests.","tokens_in":11772,"tokens_out":2276,"would_cite":true,"duration_ms":22251,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A multi-stage filtering pipeline cuts joke toxicity by about half and yields reusable English and Russian humor datasets with human scores, while two-stage fine-tuning produces cleaner yet less funny LLM humor.","keywords":["computational humor","humor generation","toxicity filtering","dataset curation","human evaluation","LLM alignment","English and Russian jokes"],"falsifier":"Re-annotate the CleanComedy Gold sets with a demographically diverse panel of native English and Russian speakers, then compare their funniness and offensiveness ratings with the pipeline's automated filter decisions; a weak correlation, or a toxicity gap much smaller than half under a different toxicity classifier, would contradict the central curation claim.","tokens_in":10863,"feed_emoji":"😂","tokens_out":8571,"duration_ms":75067,"temperature":0.7,"pith_summary":"This paper introduces CleanComedy, a toxicity-filtered corpus of English and Russian jokes built by merging existing joke collections and passing them through a multi-stage cleaning pipeline. The paper reports that the pipeline cuts the toxicity percentage in the final datasets to about half of the unfiltered source data, with 44,481 English and 40,926 Russian jokes surviving, and adds CleanComedy Gold, 1,000 jokes per language each rated by five annotators on a 1–5 humor scale. Using these datasets, the authors fine-tune an 8-billion-parameter language model with LoRA and a soft-label alignment stage that turns average human scores into training targets. In a human evaluation of 600 jokes, the curated human jokes and GPT-4o score highest on funniness, while the fine-tuned models produce cleaner but less funny output. The contribution, if correct, is a reusable lower-toxicity humor resource plus a reproducible training recipe for safer joke generation.","feed_headline":"Filtered joke dataset cuts toxicity by half in two languages","feed_subtitle":"A new joke dataset with human scores lets small LLMs generate cleaner humor, though big models still win on funniness.","key_machinery":"The load-bearing mechanism is a cascaded filtering pipeline followed by a two-stage training procedure. The filter chain keeps only entries between 50 and 150 characters, removes toxic texts with Detoxify for English and ruBERTConv for Russian (scores above 0.1 are dropped), deduplicates semantically with Sentence-BERT cosine thresholds of 0.7 (English) and 0.9 (Russian), labels jokes with a zero-shot DeBERTa-v3 classifier to remove political, racist, and insulting content, and finally deletes whole BERTopic clusters on sensitive topics such as religion, funerals, bathrooms, officers, pregnancy, nations, disabilities, and divorce. On the modeling side, the paper fine-tunes an 8-billion-parameter language model (the base, non-instruction version of Llama 3.1) with LoRA at rank 4, first by supervised fine-tuning on the clean jokes, then by an alignment stage where average human scores are linearly mapped to soft labels and trained with binary cross-entropy, an approach inspired by DPO and SimPO but using soft labels instead of chosen-rejected pairs.","core_discovery":"The paper's central claim is that automated curation can produce a humor dataset that is roughly half as toxic as its raw sources, and that training on that curated data shifts a language model toward cleaner humor generation without matching the funniness of large general models or human-written jokes. The reported toxicity percentages for the filtered datasets are 11.41% (English) and 11.37% (Russian), compared with 26.09% and 20.93% for unfiltered samples; the human funniness scores place clean human jokes at 2.96 (English) and 2.84 (Russian), GPT-4o at 3.02 and 2.38, and the aligned fine-tuned model at 2.02 and 1.74. The authors take these numbers as evidence that the filtering pipeline works and that two-stage fine-tuning with soft labels can steer generation toward friendlier humor, while acknowledging that generative humor remains an open problem.","pith_inferences":["Nothing in the paper tests whether the same filter cascade transfers to other languages or joke formats; a natural extension is to run the pipeline on a third language and check whether the toxicity-halving result persists.","The reported drop in funniness for filtered jokes suggests a possible toxicity–funniness tradeoff that the paper does not isolate; an experiment that varies the toxicity threshold and measures both metrics would make that tradeoff explicit.","Because the annotator pool skews young (20–30) and includes few native English speakers, the human scores likely reflect a specific demographic's humor; re-annotation with a broader panel would show how much the headline numbers depend on that panel.","The soft-label alignment loss could be adapted to other subjective attributes besides humor, such as politeness or helpfulness, where scalar human ratings are easier to collect than pairwise preference data."],"forward_implications":["CleanComedy provides roughly 44,000 English and 41,000 Russian jokes with reduced toxicity and duplication, plus per-joke human scores for 1,000 jokes per language, as reusable training and evaluation material.","The two-stage LoRA plus soft-label alignment recipe offers a lightweight way to make an 8B model generate cleaner humor, which could transfer to other constrained creative-text tasks.","Toxicity filtering roughly halves the share of offensive jokes compared with the unfiltered source collections, making the pipeline a candidate template for other dataset curation efforts.","Because the filtered fine-tuned models score lower on funniness than GPT-4o and human jokes, the paper implies that safety filtering alone does not produce funnier jokes, so humor quality needs separate optimization.","Publishing individual annotator scores rather than only averages enables downstream research on humor personalization and on how demographic factors shape funniness ratings."],"supporting_citations":[{"why":"Provides the Detoxify classifier used to remove English toxic jokes.","marker":"Hanu and Unitary team, 2020"},{"why":"Supplies the ruBERTConv Russian toxicity classifier and its detoxification motivation.","marker":"Kuratov and Arkhipov, 2019; Dementieva et al., 2021"},{"why":"Provides Sentence-BERT embeddings used for semantic duplicate detection.","marker":"Reimers and Gurevych, 2019, 2020"},{"why":"Supplies the zero-shot DeBERTa-v3 model that labels and filters political, racist, and insulting jokes.","marker":"Laurer et al., 2024"},{"why":"Provides BERTopic, the topic model used to identify and delete sensitive joke clusters.","marker":"Grootendorst, 2022"},{"why":"Supplies LoRA, the parameter-efficient fine-tuning method used in both training stages.","marker":"Hu et al., 2021"},{"why":"Inspires the alignment loss used after supervised fine-tuning.","marker":"Rafailov et al., 2024"},{"why":"Inspires the reference-free soft-label variant of the alignment loss.","marker":"Meng et al., 2024"}],"fun_headline_variants":["CleanComedy halves joke toxicity in English and Russian","Filtered joke data slashes toxicity for cleaner AI humor","CleanComedy cuts toxicity by half, but funniness lags","Toxicity-filtered jokes yield cleaner but less funny AI humor","CleanComedy: half the toxicity, but humor still lacks punch"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The pipeline assumes that its automated proxies—toxicity classifier scores, SBERT similarity cutoffs, zero-shot content labels, and deleted BERTopic clusters—are a reliable stand-in for what a broad audience finds offensive; if those proxies are biased, the 'clean' dataset and the reported toxicity reduction inherit that bias.","fun_headline_variants_meta":{"raw":{"variants":["CleanComedy halves joke toxicity in English and Russian","Filtered joke data slashes toxicity for cleaner AI humor","CleanComedy cuts toxicity by half, but funniness lags","Toxicity-filtered jokes yield cleaner but less funny AI humor","CleanComedy: half the toxicity, but humor still lacks punch"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001259,"raw_usage":{"total_tokens":5100,"prompt_tokens":833,"completion_tokens":4267,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":449,"completion_tokens_details":{"reasoning_tokens":4180}},"tokens_in":449,"tokens_out":4267,"duration_ms":29369,"temperature":1.0,"reasoning_tokens":4180,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T17:12:25.952133+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-annotate the CleanComedy Gold sets with a demographically diverse panel of native English and Russian speakers, then compare their funniness and offensiveness ratings with the pipeline's automated filter decisions; a weak correlation, or a toxicity gap much smaller than half under a different toxicity classifier, would contradict the central curation claim.","supporting_citations":[],"review_version":1}