{"id":"6936c40d-5a0d-464e-919e-cd5b55ec2757","arxiv_id":"2501.11433","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"LLM assistance increases the quantity of meme ideas and lowers perceived effort, but does not raise average meme quality; fully AI-generated memes score highest on average, while human-made memes dominate the funniest top performers.","lead":"The paper ran a crowdsourced experiment where people made memes alone, with an LLM assistant, or with the AI working fully alone, then had other users rate the results on humor, creativity, and shareability. It finds that AI help generated more ideas with less effort but did not make human memes better on average, while fully AI-made memes scored higher on average, even though the funniest individual memes were mostly human-made.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The AI-only meme condition is not demonstrably comparable: the paper reports generating 20 captions per image/topic but never states how the 150 rated AI memes were selected, so the headline AI-superiority result could be a selection artifact.","rationale":"The study has genuine strengths: a between-subjects design, a plausible productivity/effort result, and a null human-AI quality difference that is internally consistent with prior co-writing findings. However, the paper's most striking empirical claim, that AI-only memes outperform both human-involved conditions, depends on the comparability of the AI-only stimulus set. The text itself creates an arithmetic inconsistency: 15 combinations x 20 captions is 300, yet the study reports 150 AI images, with no selection rule stated. The human conditions are explicitly randomly sampled, so the absence of an equivalent statement for the AI condition is not a stylistic omission; it changes the interpretation. If the AI images were selected by the researchers for quality or caption length, the average-superiority result is not evidence about LLM meme generation but about the selected subset. The per-topic result reinforces this fragility, since the significant pairwise differences are driven primarily by work memes. Secondary issues noted by the reader, such as the abstract's participant-count error and the lack of inter-rater reliability, are real but less central. My read leaves the verdict at conditional acceptance because the problem is fixable by disclosure and reanalysis; the productivity/effort findings and the null human-AI quality difference would likely survive even if the AI-only comparison were rerun with a clearly random sample.","tokens_in":15684,"tokens_out":4203,"duration_ms":45049,"concrete_test":"Examine the supplementary material or study scripts to recover the exact AI-only selection rule. If the rule is not explicit, ask the authors to state whether the 150 were all generated captions, a random sample, or a curated subset. Then rerun the meme-rating analysis on a random sample of all 300 AI-generated captions (or on all 300) and recompute the pairwise comparisons against the human conditions. As a second check, rerun the analysis on the \"food\" and \"sports\" topics separately; if the AI-only advantage disappears or reverses in those subsets, the headline claim needs to be qualified to \"work\" memes or withdrawn.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing assumption is the construction of the AI-only condition. In Section 3.3 the authors say they instructed the model to \"generate 20 meme captions\" for each image/topic combination, with five images and three topics giving 15 combinations. That would produce 300 captions, but Section 3.2 says this yielded \"additional 150 images.\" The missing step is how 150 were chosen from 300. For the human conditions, the 150 memes are explicitly a random sample of ten per image/topic after exclusions. No equivalent statement exists for the AI condition; it could be a random subset, the first ten generated, or a researcher-selected subset based on perceived quality or caption length. If any non-random curation occurred, the finding that AI-only memes scored higher on average is an artifact of selection rather than a property of AI generation. The concern is not merely procedural: the AI-only comparison drives the paper's headline claim, and the paper's own per-topic analysis shows the significant pairwise effects are concentrated in \"work\" memes, so the unqualified abstract claim is already more fragile than the data. The selection ambiguity is therefore load-bearing for the paper's central conclusion.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a two-part empirical study of LLM-assisted meme creation. In the first phase, human participants produced meme captions either alone or with a GPT-4o chat assistant, and the authors measured idea counts, NASA-TLX workload, and subjective ownership. In the second phase, crowdsourced raters evaluated 450 memes (150 human-only, 150 human-AI collaborative, 150 AI-only) on humor, creativity, and shareability. The central reported findings are that LLM assistance increases idea quantity and reduces perceived effort, that human-AI collaboration does not significantly improve rated meme quality relative to human-only creation, and that AI-only memes receive the highest average ratings overall, with the significant differences concentrated in work-related memes.","tokens_in":15824,"tokens_out":6005,"duration_ms":66123,"significance":"If the results hold, the study is a useful empirical contribution to human-AI co-creativity in a culturally nuanced domain, with a comparatively large stimulus set and honest reporting of per-topic analyses. The authors also make interaction logs and generated captions available as supplementary material and apply standard nonparametric/parametric tests. However, several methodological details and claim-calibration issues need to be resolved before the central conclusions can be fully endorsed.","major_comments":[{"comment":"The construction of the AI-only condition is underspecified and potentially load-bearing. Section 3.3 says the model was instructed to generate 20 meme captions for each of the 15 image-topic combinations, which would produce 300 captions, yet Section 3.2 states that this step yielded 150 images (i.e., 10 per combination). The paper never states whether the 150 rated captions were the first 10 generated, a random subsample, or a researcher-selected subset. If any non-random curation occurred, the finding that AI-only memes scored higher on average would be an artifact of selection rather than a property of AI generation. The authors should specify the exact selection procedure, or state explicitly if all 20 captions per combination were used and 10 were dropped for a documented reason.","section":"§3.2 and §3.3"},{"comment":"The unit of statistical analysis is not reported. With 98 raters each rating 50 images, there are 4,900 ratings of 450 memes; the paper does not state whether meme-level means were computed before the ANOVA/Kruskal-Wallis tests or whether individual ratings were treated as independent observations. If individual ratings were used, the independence assumption is violated because each meme contributes multiple ratings and each rater contributes 50 ratings. The authors should report how ratings were aggregated, the number of ratings per meme, and an inter-rater reliability statistic (e.g., ICC) for humor, creativity, and shareability.","section":"§4.2 and Table 1"},{"comment":"The abstract and conclusion claim that AI-only memes performed better than both human-only and human-AI collaborative memes in all areas on average, but the paper's own analysis qualifies this in two ways. First, the pairwise comparison for shareability between the collaborative and AI-only conditions was not significant. Second, the omnibus differences are not significant for the sports and food topics; the authors state that the significant differences 'seem to stem primarily from the memes about the topic of work.' The claims should be qualified to report the topic-specific pattern and the non-significant shareability comparison.","section":"Abstract, §4.2, §7"},{"comment":"The null quality result for the human-AI condition may be a weak test of the paper's own definition of co-creativity. Section 6.2 reports that less than half of participants interacted with the LLM multiple times and only six participants had more than eight interactions. Since the introduction defines co-creativity in terms of iterative, dialog-based refinement, many participants may not have engaged in the iterative process being studied. The authors should either restrict the 'collaboration does not improve quality' conclusion to the observed level of engagement or report a post-hoc analysis of participants who used the LLM more extensively.","section":"§6.2 and §5.2"}],"minor_comments":[{"comment":"The abstract says 'three groups of 50 participants each,' but Section 3.6 reports 124 recruited and 98 completing the task, and the AI-only condition had no human participants. This wording should be corrected.","section":"Abstract"},{"comment":"The significance legend reads '*: p < 0.05, **: p < 0.01, **: p < 0.001'; the third symbol should be '***'.","section":"Figure 8"},{"comment":"The header contains typographical errors: 'ANOV A' should be 'ANOVA' and 'Kruska-Wallis' should be 'Kruskal-Wallis'.","section":"Table 1"},{"comment":"There are several typos: 'meed' should be 'meet' in the introduction, 'signiciantly' should be 'significantly', 'simiarly' should be 'similarly', and 'participated provided' should be 'participants provided'.","section":"Throughout"},{"comment":"The analysis of top-performing memes is purely descriptive; the text should label it as such and avoid causal phrasing such as 'humans can be wittier still' without a statistical comparison.","section":"§5.3"},{"comment":"The introduction refers to 'two user studies,' but the design is one study with two phases; this terminology should be made consistent.","section":"§1"}],"recommendation":"major_revision","confidential_remarks":"The most serious risk is the underspecified selection of the 150 AI-only memes from the 300 generated captions. If the authors can report the exact selection procedure or verify from logs that it was random, the central claim may be salvageable; otherwise the headline AI-superiority result is not interpretable. The other major comments concern reporting and claim calibration rather than irreparable flaws."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The core result here is believable: LLM help increases the number of ideas and cuts perceived effort without raising average quality, and the paper honestly shows the AI-only advantage is mostly driven by 'work' memes. That is a useful, if not shocking, addition to the co-creativity literature, and it aligns with prior co-writing studies like Wan et al. The three-way comparison for memes is new, and the study is competently run: between-subjects, adequate sample, appropriate nonparametric tests, and a per-topic breakdown that the authors did not bury. I give them credit for that.\n\nThe soft spots are real but mostly fixable. The biggest one is the AI-only condition. They say they prompted the model to generate 20 captions per image/topic combination, which with five images and three topics gives 300 captions, yet they used 150 images. The paper never states how those 150 were selected. If it was a random sample, fine, but if anyone picked the funnier ones, the headline that AI beats humans on average is an artifact. The paper also lacks an equivalent to the favorite-selection step that humans performed, so the AI-only condition is not just a different creator but also a different pipeline. They need to clarify this before the AI-superiority claim can be taken at face value.\n\nTwo smaller issues: the abstract says 'three groups of 50 participants each,' but only 98 humans completed the first phase, split across two conditions, and the AI condition was not a group of 50 people. That is just wrong and should be corrected. Also, no inter-rater reliability is reported for the subjective ratings, which is a minor omission for a crowdsourced humor study.\n\nNone of this sinks the central finding about human-AI collaboration: more ideas, less effort, no quality gain. That result does not depend on the AI-only comparison and is consistent with the literature. The AI-only finding is the flashy part but also the most fragile piece.\n\nI would send this to peer review. The study deserves a serious referee, but the authors need to clarify the AI-only selection procedure, fix the abstract, and report reliability stats. A conditionally accept with those changes is reasonable.","headline":"Solid HCI experiment on LLM-assisted meme creation; the productivity-without-quality result is credible, but the AI-superiority headline rests on an underspecified selection procedure and the abstract misstates the design.","tokens_in":16428,"tokens_out":1971,"would_cite":true,"duration_ms":23620,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"LLM assistance raises meme output and cuts effort but not quality; AI-only memes rate highest on average.","keywords":["Human-AI collaboration","LLM","co-creativity","memes","humor","crowdsourcing","creativity evaluation","GPT-4o"],"falsifier":"Generate memes from every caption the LLM produces for each image-topic pair, or from a pre-registered random sample, and have the same crowd rate them; if the AI-only average no longer exceeds human and human-AI memes, the reported advantage is a selection artifact rather than a property of AI generation. Separately, rerunning the analysis on non-work topics alone would test whether the headline effect survives outside the topic that drove it.","tokens_in":15424,"feed_emoji":"😂","tokens_out":5067,"duration_ms":49834,"temperature":0.7,"pith_summary":"This paper tests whether a large language model can act as a genuine co-creative partner in a humor-rich, culturally specific task: writing captions for internet memes. In a between-subjects study, participants produced memes alone, with a chat-based LLM assistant, or not at all, with the LLM generating the third set of memes autonomously, and a separate crowd then rated humor, creativity, and shareability. The authors report that LLM assistance significantly increased the number of ideas and reduced perceived effort, but it did not improve rated quality of human-involved memes. Fully AI-generated memes scored higher on average than both human-only and human-AI memes on all three scales, although that advantage mostly came from the 'work' topic, and top-rated memes tell a different story: human-made memes were funniest, while human-AI collaborations led in creativity and shareability. The paper's point is that productivity gains from AI co-creation do not automatically translate into better creative output, and that broad average appeal and deep human resonance are different things.","feed_headline":"AI memes beat human and human-AI teams on average","feed_subtitle":"User study finds AI-only memes rate highest on average, yet top-rated humor still comes from people.","key_machinery":"The load-bearing mechanism is a three-condition between-subjects meme-creation workflow followed by crowdsourced rating. In the human-AI condition, participants brainstormed captions in a chat interface with a multimodal LLM (GPT-4o) that returned up to three caption ideas per reply; ideas extracted from the chat were added to the participant's idea list, and the participant chose favorites and composed final images. The AI-only condition was produced by prompting the same model to generate captions for each image-topic pair. Ratings of humor, creativity, and shareability by a separate crowd of raters, analyzed with ANOVA/Kruskal-Wallis and Mann-Whitney-U tests, carry the conclusion that AI assistance changes productivity and appeal without changing average quality of human-involved output.","core_discovery":"The paper's central empirical discovery is a split between process and product in human-AI co-creation of humor. On the process side, people who could chat with an LLM produced significantly more caption ideas while reporting no more workload, and actually less effort, than people working alone; they also felt somewhat less ownership of the ideas. On the product side, crowdsourced ratings showed no significant quality difference between human-only and human-AI memes on humor, creativity, or shareability, whereas memes generated entirely by the LLM were rated higher on average than both human-involved conditions on all three dimensions. The authors qualify this by showing that the overall effect is driven mainly by memes about work and by the top-meme analysis, where humans took most of the funniest slots and human-AI teams took the most creative and shareable ones.","pith_inferences":["If the AI-only captions used in the evaluation were curated from the 20 generated per image-topic pair, the average superiority of AI memes could be an artifact of selection; a replication that rates all generated captions would settle this.","Because most participants used the chat sparingly, a more guided or iterative co-creative interface might produce quality gains that the current open-ended chat did not; this is a design implication the paper leaves open.","The broad-audience result likely depends on the evaluator pool; rating the same memes with smaller, culturally homogeneous audiences could reverse the AI advantage, consistent with the paper's cultural-context caveat."],"forward_implications":["LLM assistants can be used to expand the idea space in humor tasks without adding perceived workload, making ideation cheaper and faster.","Simply handing users a chat assistant is not enough to improve average meme quality; the collaboration needs more structure or iteration to beat unaided humans.","Average ratings over a broad crowd favor AI-generated content, so evaluations of creative AI should separate average performance from top-performance and per-topic effects.","The work-topic result implies that domain or template content moderates the AI advantage, not just the creative process itself.","Top-meme findings suggest humans and AI contribute differently: humans for humor, human-AI pairs for creativity and shareability."],"supporting_citations":[{"why":"Supplies the closest prior result that LLM-assisted prewriting yields more ideas without better quality, which the paper's ideation finding aligns with.","marker":"[49]"},{"why":"Prior autonomous LLM meme generation that this study extends to human-AI co-creation.","marker":"[50]"},{"why":"Evidence that LLMs outperform humans on divergent-thinking fluency, motivating the productivity result.","marker":"[29]"},{"why":"Baseline showing LLM jokes can be rated comparably to human jokes, framing the AI-only humor result.","marker":"[23]"},{"why":"Defines meme humor and creativity as multimodal, supporting the three rating dimensions.","marker":"[48]"},{"why":"Provides the co-creative AI interaction framework the paper uses to define human-AI collaboration.","marker":"[24]"},{"why":"Example of AI as creative sketching partner, supporting the claim that AI assistance aids ideation.","marker":"[31]"}],"fun_headline_variants":["AI-only memes beat human and hybrid teams on average","Average memes: AI wins, but top humor still human","More ideas, less effort with AI, but meme quality doesn't rise","In meme co-creation, AI boosts quantity not quality","Top memes: humans funnier, AI teams more creative"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The comparison between AI-only and human-involved memes assumes that the 150 AI memes fairly represent what the model produces; the paper does not say whether all, random, or hand-picked captions from the model's 20 per image-topic were used, so the headline AI advantage rests on an unstated selection step.","fun_headline_variants_meta":{"raw":{"variants":["AI-only memes beat human and hybrid teams on average","Average memes: AI wins, but top humor still human","More ideas, less effort with AI, but meme quality doesn't rise","In meme co-creation, AI boosts quantity not quality","Top memes: humans funnier, AI teams more creative"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000185,"raw_usage":{"total_tokens":1358,"prompt_tokens":1019,"completion_tokens":339,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":635,"completion_tokens_details":{"reasoning_tokens":253}},"tokens_in":635,"tokens_out":339,"duration_ms":3764,"temperature":1.0,"reasoning_tokens":253,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T18:15:44.740142+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Generate memes from every caption the LLM produces for each image-topic pair, or from a pre-registered random sample, and have the same crowd rate them; if the AI-only average no longer exceeds human and human-AI memes, the reported advantage is a selection artifact rather than a property of AI generation. Separately, rerunning the analysis on non-work topics alone would test whether the headline effect survives outside the topic that drove it.","supporting_citations":[],"review_version":1}