{"id":"5958d613-d2d5-42f8-ae4b-9a1e7984f185","arxiv_id":"2504.15392","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Co-created expert-LLM definitions of sexism rarely beat LLM-generated definitions, but expert-written definitions perform substantially worse in zero-shot sexism detection.","lead":"Researchers of sexism interacted with an LLM to write definitions of sexism, then tested those definitions in zero-shot detection across five benchmarks. The study found that LLM-written definitions beat expert-written ones on average, while a few experts improved detection by co-editing definitions with the model.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Co-creation result rests on near-identical prompts; highlighted per-expert gains are not distinguished from repeated-run noise.","rationale":"The reader's CONDITIONAL verdict is appropriate, and my stress-test does not move it. The large average gap between expert-written definitions (F1 = .532) and LLM-generated/co-created definitions (.765/.762) is likely real and is not plausibly explained by length, because Table 9 shows within-type length-F1 correlations that are mostly negative. The more serious problem is the co-creation sub-claim. Since most co-created definitions are semantic duplicates of LLM-generated definitions, the per-expert 'improvements' cited in Section 4.5, particularly for the LLM-inexperienced Expert 6, compare nearly identical prompts. Without reported repeated-run variance or significance tests, those differences cannot be distinguished from noise. This does not invalidate the released interaction dataset, the qualitative taxonomies, or the framework contribution, but it means the abstract's specific claim that 'some experts do improve ... also experts who are inexperienced in using LLMs' is not yet supported. The paper should either report the repeated-run analysis, restrict co-creation claims to experts whose definitions actually differ, or soften the conclusion to say co-created definitions performed on par with LLM-generated ones overall. None of this changes the CONDITIONAL verdict, but it identifies the precise evidence needed for acceptance.","tokens_in":26156,"tokens_out":5461,"duration_ms":49747,"concrete_test":"Re-run the same 27 definition prompts through GPT4o at temperature=0 for at least 10 independent repetitions (or, minimally, compare the two runs the authors state they completed), computing macro-F1 for every expert-definition-dataset cell. For each expert, compute delta = F1(co-created) - F1(LLM-generated). If the deltas for the highlighted experts (especially Expert 6) are no larger than the maximum delta observed between two repetitions of the identical prompt, then the per-expert co-creation improvements are statistically indistinguishable from repeated-run noise. As a stricter check, restrict the co-creation analysis to the two experts (2 and 4) whose co-created and LLM-generated definitions are materially different per Table 8; if neither shows a positive delta, there is no evidence that co-creation adds value beyond selecting an LLM-generated definition.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The distinctive part of the central claim is that co-created definitions improve performance for some experts, including an LLM-inexperienced one. That claim is undercut by the paper's own Appendix E. For seven of nine experts (1, 3, 5, 6, 9, 10, 11), the co-created definition is identical to or a lightly cleaned-up version of the LLM-generated definition: Table 8 reports TF-IDF cosine similarity between co-created and LLM-generated at or above .95 for all of these experts except Expert 5, with Experts 3 and 11 at 1.00. The paper even labels these cases 'robustness tests.' Yet Section 4.5 then treats per-expert differences between co-created and LLM-generated F1 as evidence of a co-creation effect, specifically for Expert 6, whose co-created and LLM-generated definitions have similarity .97 (TF-IDF) and .87 (SBERT). When the compared prompts are nearly identical, any F1 difference is within the noise of repeated prompting; the authors report no confidence intervals, significance tests, or repeated-run variance for the main temperature=0 GPT4o experiment, despite stating they ran it twice. The aggregate difference is also tiny (co-created mean F1 = .762 vs LLM-generated .765), so the 'some experts improve' conclusion is not established. The length confound raised by the reader is less decisive: Appendix E.1, Table 9 shows length correlates negatively with F1 within each definition type, so length alone cannot explain the large expert-written versus LLM-generated gap. The load-bearing weakness is the non-independence of the co-created and LLM-generated conditions for most experts.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies how sexism researchers interact with an LLM (GPT-3.5) to produce definitions of sexism, and whether those definitions improve zero-shot sexism detection when placed in a prompt template for GPT-4o. Nine experts each produced three definitions: an expert-written definition, a preferred LLM-generated definition, and a co-created definition. These 27 definitions were used to classify 2,500 texts from five sexism benchmarks, yielding 67,500 classification decisions. The main reported results are that expert-written definitions perform much worse (mean macro-F1 = .532) than LLM-generated definitions (.765) and co-created definitions (.762), and that some experts, including an LLM-inexperienced one, improve performance with co-created definitions. The paper also presents qualitative taxonomies of expert interaction strategies and definition similarity analyses, and it releases the interaction framework, definitions, and code.","tokens_in":26465,"tokens_out":4084,"duration_ms":38533,"significance":"The paper is methodologically interesting and fills a real gap: it connects open-ended expert–LLM interaction, definition co-creation, and downstream benchmark evaluation in one pipeline. Its strengths include a clearly described and reproducible experimental protocol, a released corpus of definitions and interactions, multiple benchmarks, and robustness checks with a second model and a different temperature. The large gap between expert-written and LLM-generated definitions is a robust and potentially useful finding for prompt design. However, the more distinctive claim about co-created definitions is not currently established. The manuscript's own Appendix E shows that most co-created definitions are identical or nearly identical to LLM-generated definitions, and Section 4.4.1 reports no confidence intervals or significance tests. The aggregate co-created vs. LLM-generated difference (.762 vs. .765) is within any reasonable estimate of sampling noise, and the per-expert comparisons promoted in the abstract and Section 4.5 are therefore likely to reflect repeated-run variation rather than a co-creation effect.","major_comments":[{"comment":"The claim that \"some experts do improve classification performance with their co-created definitions\" is not supported by the evidence as presented. For seven of nine experts (1, 3, 5, 6, 9, 10, 11), the co-created definition is identical to or a lightly cleaned version of the LLM-generated definition, with TF-IDF cosine similarity at or above .95 (Experts 3 and 11 have similarity 1.00). The paper itself labels these cases \"robustness tests\" in Appendix E. Under these conditions, any F1 difference between the co-created and LLM-generated conditions is repeated-prompt noise, not a co-creation effect. Section 4.5 nevertheless treats Expert 6's difference as evidence of co-creation success despite a TF-IDF similarity of .97 between that expert's co-created and LLM-generated definitions. The analysis should be restricted to the experts whose co-created definitions genuinely differ from the LLM-generated ones (notably Experts 2 and 4), or the claim should be reformulated as a comparison of editing behavior rather than an independent contribution of expert knowledge.","section":"Section 4.5, Table 8, Appendix E"},{"comment":"No confidence intervals, significance tests, or repeated-run variance are reported for the main temperature=0 GPT-4o results, yet the central comparisons are close. The aggregate means are LLM-generated F1 = .765 and co-created F1 = .762, a difference of three thousandths that is almost certainly within sampling noise. The manuscript notes in Appendix C.2 that the full experiment was run twice, so run-level variance is available and should be reported. Without a proper uncertainty quantification, the abstract's statement that \"some experts do improve\" is indistinguishable from a noise-driven claim. I would expect per-expert bootstrap confidence intervals, a paired significance test, or an explicit effect size with uncertainty for the co-created vs. LLM-generated difference.","section":"Section 4.4.1 and Appendix C.2"},{"comment":"The interpretation of the expert-written vs. LLM-generated gap is incomplete because definition length and content are confounded. Expert-written definitions average 34.44 tokens versus 119.89 for LLM-generated and 110.55 for co-created definitions. Appendix E.1 shows negative within-type correlations between length and F1, which weakens a simple \"longer is better\" explanation, but the cross-type comparison still conflates length, lexical style, and content. The paper would be stronger if it included a control analysis (e.g., length-matched definitions, or a comparison restricted to experts who wrote substantially longer definitions) or explicitly acknowledged that the expert-written gap may be driven by informativeness and style rather than by expert knowledge per se.","section":"Section 4.4.1, Appendix E.1"}],"minor_comments":[{"comment":"The sentence \"the LLM-generated definition performs slightly lower (F1 = .699) than the co-written definition (F1 = .584) or participant-written definitions (F1 = .703)\" is internally inconsistent: .584 is the lowest of the three values, not lower than .699. The numbers or the wording should be corrected.","section":"Appendix H.2"},{"comment":"The text reports \"M F1 = .760 vs M F1 = .68 for the temperature=0 run,\" but Section 4.4.1 reports a mean F1 of .765 for the temperature=0 condition. This inconsistency should be reconciled.","section":"Appendix G.2"},{"comment":"The last row entry \"Similarity (TF-IDF) - LLM-generated\" on HateCheck shows \"33\" rather than a decimal value; it should presumably be \".33\".","section":"Table 9"},{"comment":"The text uses European decimal notation \"2.500 texts\" where an English-language paper should use \"2,500 texts\".","section":"Section 3.4"},{"comment":"The terminology for the three definition types is inconsistent: \"expert-written,\" \"participant definitions,\" and \"hybrid\" are all used in different places to refer to the same conditions. Please standardize the labels.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The paper's core subject is timely and the released artifacts are valuable, but the headline co-creation finding rests on near-identical prompts and absent uncertainty quantification. The authors have the data to fix this—they report running the main experiment twice and can restrict the analysis to genuinely distinct co-created definitions—so I see this as a major-revision issue rather than a rejection. I would also encourage the editor to ask for the per-expert raw F1 tables and run-level results to be included in the main text or a clearly labeled appendix."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The short version: read this for the pipeline, not for the headline claim. The four-part design (survey, free-form interaction, definition co-creation, zero-shot evaluation) is new, and the released repository of expert interactions and 27 definitions is a real resource for computational social science. The qualitative taxonomy of expert prompting strategies is also thoughtful and well-presented. Credit where due: this is a serious, clearly written empirical study with an honest limitations section.\n\nThe soft spot is the co-creation result. The abstract and Section 4.5 claim that some experts, including an LLM-inexperienced one, improve zero-shot sexism detection with co-created definitions. But Appendix E shows that for seven of nine experts the co-created definition is identical to or a lightly cleaned-up version of the LLM-generated one (TF-IDF cosine similarity .95–1.00). The paper itself calls these cases robustness tests. Comparing nearly identical prompts and attributing per-expert F1 gaps to a co-creation effect is not justified, especially with no confidence intervals or significance tests. The aggregate difference is also tiny (co-created .762 vs LLM-generated .765), so the average direction does not support the claim either.\n\nThe length confound is real but less decisive than the similarity issue. Expert-written definitions average 34 tokens versus 119 for LLM-generated and 110 for co-created, so the expert-written deficit may partly be a length effect. However, the appendix's within-type negative correlations between length and F1 complicate a pure length story. Still, the paper should have reported a no-definition baseline and ideally controlled for length.\n\nWhat holds up: the stark difference between expert-written and LLM-generated definitions (F1 .532 vs .765) is robust in direction and large in magnitude, even if its interpretation is muddied. The method for collecting and evaluating definitions is the contribution, and the authors are transparent enough to present the similarity data that undermines their own stronger claim.\n\nThis paper deserves a serious referee, but the referee should push for revision: either drop the per-expert co-creation claim or restrict it to the two experts (2 and 4) whose co-created definitions genuinely differ from the LLM-generated ones, and report variance across the repeated runs. The resource contribution and the qualitative analysis are publishable as they stand; the quantitative co-creation story needs to be re-cast.","headline":"A genuinely useful pipeline and resource paper whose central co-creation claim is undercut by its own appendix: for most experts the co-created definition is nearly identical to the LLM-generated one.","tokens_in":27022,"tokens_out":2081,"would_cite":true,"duration_ms":19968,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"LLM-generated definitions outperform expert-written ones in zero-shot sexism detection, with co-creation helping only some experts.","keywords":["sexism detection","zero-shot classification","large language models","hybrid intelligence","co-created definitions","expert knowledge","prompt engineering","benchmark evaluation"],"falsifier":"A direct test: take the nine expert-written definitions, expand each to the average length of the LLM-generated definitions without adding new substantive criteria, and re-run GPT-4o on the same 2,500 texts; if $F_1$ jumps to the LLM-generated level, the advantage is length. A second test: compare only the experts whose co-created and LLM-generated definitions are semantically identical and check whether any performance difference remains.","tokens_in":25936,"feed_emoji":"🤖","tokens_out":6681,"duration_ms":53090,"temperature":0.7,"pith_summary":"This paper asks whether sexism researchers can improve zero-shot sexism detection by writing, selecting, or co-creating the definition of sexism that an LLM is prompted with. Nine sexism researchers each produced three definitions: their own, the best one the LLM generated, and a version they edited together with the LLM. Prompting GPT-4o with these 27 definitions on 2,500 texts from five sexism benchmarks shows that expert-written definitions do markedly worse on average than LLM-generated ones. At the same time, some experts, including at least one with little LLM experience, get better results from the co-created definition than from the model's alone. The central claim is that the definition's provenance and content substantively change zero-shot detection performance, and that expert knowledge does not automatically transfer into a better prompt.","feed_headline":"LLM-written definitions beat experts' own in sexism detection","feed_subtitle":"Model-generated definitions scored far higher across five benchmarks; co-creation helped only some experts.","key_machinery":"The load-bearing object is the definition-as-prompt: each expert contributes three textual definitions of sexism (expert-written, LLM-generated, co-created), and each is inserted into a fixed prompt template that asks GPT-4o to label a text 'sexist,' 'non-sexist,' or 'can't say.' The pipeline—survey, interactive assessment of the model's knowledge, definition co-creation, and zero-shot classification—ties the qualitative interaction strategies to measured $F_1$ scores, so the paper can attribute performance differences to how the definition was produced.","core_discovery":"On the paper's own terms, the core discovery is that in zero-shot sexism classification, the provenance of the prompt definition changes performance: LLM-generated definitions (mean macro $F_1 = .765$) and co-created definitions (mean $F_1 = .762$) outperform expert-written definitions (mean $F_1 = .532$) across five benchmarks. The pattern is dataset-dependent, with near-tied results on CallMeSexist and a large gap on RedditGuest. A few experts nevertheless obtain their best results from co-created definitions, including Expert 6, who reported low confidence with LLMs. The authors read this as partial evidence for hybrid intelligence: expert and model can sometimes jointly produce a better detection prompt than either alone.","pith_inferences":["Because expert-written definitions average 34 tokens versus roughly 110–120 for the other two types, a plausible reading is that length, not authorship, drives the gap; this is an inference, not the paper's claim.","Since most co-created definitions are near-identical to LLM-generated ones (cosine similarities up to 1.00), the observed co-creation advantage may be an expert-selection effect—which model output the expert chose to keep—rather than an editing effect.","The strong dataset dependence suggests that prompt definitions are a high-variance intervention on rare-class benchmarks like RedditGuest; practitioners should probably evaluate several definition variants before relying on one.","Extending the same pipeline to other contested constructs would show whether the expert-written deficit is specific to sexism or a general feature of short expert definitions in zero-shot prompting."],"forward_implications":["Definition provenance changes zero-shot sexism detection: model-generated definitions are better prompts on average than short expert definitions.","Expert knowledge can still enter a prompt usefully, but only through co-creation, and success depends on the individual expert, not on self-reported LLM confidence.","Prompt-sensitivity varies by dataset: on CallMeSexist all definition types perform similarly, while on RedditGuest expert-written definitions fall far behind.","The co-created definition can lift performance above the majority-class baseline on some datasets and experts, so hybrid definitions are not uniformly useless.","The effect is tied to the model and temperature: at higher temperature or with LLaMa the definition-type differences shrink, so the finding is not model-independent."],"supporting_citations":[{"why":"Supplies the prompt template used for zero-shot classification and the prior framing of LLM-generated counterfactually augmented data for harmful language detection.","marker":"Sen et al. (2023)"},{"why":"Provides the CallMeSexist dataset, one of the five sexism benchmarks used for evaluation.","marker":"Samory et al. (2021)"},{"why":"Provides the EDOS dataset, one of the five sexism benchmarks used for evaluation.","marker":"Kirk et al. (2023)"},{"why":"Provides the Reddit Misogyny dataset, the benchmark with the most skewed class distribution.","marker":"Guest et al. (2021)"},{"why":"Provides the EXIST dataset, one of the five sexism benchmarks used for evaluation.","marker":"Rodriguez-Sanchez et al. (2021)"},{"why":"Provides the HateCheck subset targeting women, one of the five sexism benchmarks used for evaluation.","marker":"Röttger et al. (2021)"},{"why":"Prior work comparing expert-written and GPT-generated definitions, which this paper extends by adding co-created definitions.","marker":"Peskine et al. (2023)"},{"why":"Provides the zero-shot definition-prompting methodology and temperature setting that the modeling experiments build on.","marker":"Korre et al. (2025)"},{"why":"Defines hybrid intelligence, the conceptual frame that motivates the co-creation experiment.","marker":"Dellermann et al. (2019)"}],"fun_headline_variants":["LLM definitions beat expert prompts in sexism detection","Co-created definitions rival pure LLM prompts for sexism detection","Expert-LLM co-creation boosts sexism detection for some experts","Model-written definitions beat expert-written ones in sexism tasks"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The claim that definition provenance changes detection rests on the assumption that the measured $F_1$ differences come from the definition content, not from confounds such as definition length (expert-written definitions average 34 tokens versus 110–120 for the others) or from the high similarity between co-created and LLM-generated definitions.","fun_headline_variants_meta":{"raw":{"variants":["LLM definitions beat expert prompts in sexism detection","Co-created definitions rival pure LLM prompts for sexism detection","Expert-LLM co-creation boosts sexism detection for some experts","Model-written definitions beat expert-written ones in sexism tasks"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000248,"raw_usage":{"total_tokens":1522,"prompt_tokens":897,"completion_tokens":625,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":513,"completion_tokens_details":{"reasoning_tokens":554}},"tokens_in":513,"tokens_out":625,"duration_ms":5769,"temperature":1.0,"reasoning_tokens":554,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T11:27:37.560542+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A direct test: take the nine expert-written definitions, expand each to the average length of the LLM-generated definitions without adding new substantive criteria, and re-run GPT-4o on the same 2,500 texts; if $F_1$ jumps to the LLM-generated level, the advantage is length. A second test: compare only the experts whose co-created and LLM-generated definitions are semantically identical and check whether any performance difference remains.","supporting_citations":[{"cited_title":"SemEval-2023 Task 10: Explainable Detection of Online Sexism","cited_arxiv_id":"2303.04222","evidence_quote":"Provides the EDOS dataset, one of the five sexism benchmarks used for evaluation."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the EXIST dataset, one of the five sexism benchmarks used for evaluation."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the zero-shot definition-prompting methodology and temperature setting that the modeling experiments build on."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines hybrid intelligence, the conceptual frame that motivates the co-creation experiment."}],"review_version":1}