{"id":"4171e5f3-edae-483b-a4b8-ceef3bbd372a","arxiv_id":"2506.05415","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Game length and guess distances predict GPT-3.5-labeled amusement in Wordle with 54.5% accuracy, just above the 50% baseline.","lead":"A small study tests whether software can predict which Wordle games amuse people, using about 80,000 Reddit comments labeled by GPT-3.5. Simple game features predict the amused labels with 54.5% accuracy, barely above chance, so the effect is real but weak.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Predictive accuracy is measured against GPT-3.5 labels whose human agreement is weak (κ 0.16–0.40); without evaluating the model on human amusement labels, the claim that user amusement is predictable is unsupported.","rationale":"The reader's weakest assumption is exactly the load-bearing condition for the central claim: the prediction target must be human amusement, not merely GPT-3.5's labels. The evidence in the paper shows that GPT-3.5's agreement with humans is weak (κ ≤ 0.398), and the few-shot prompt itself mismatches the construct by asking whether 'the comment is funny' rather than whether the commenter was amused. Since all model training and evaluation is on GPT labels, the 54.5% accuracy demonstrates only that game features are weakly predictive of GPT-3.5's labels, which the paper's own data do not validate as a proxy for human amusement. The additional observation that a single feature, number of guesses, nearly reproduces the full-model accuracy makes it especially important to rule out a trivial confound such as comment conventions about score reporting. The proposed human-label evaluation directly targets this weak point: if the fitted model cannot distinguish human-amused from human-not-amused comments, the abstract's claim about user amusement must be withdrawn or explicitly restricted to GPT-3.5-labeled amusement. Because the reader already made the label-proxy issue the basis for a conditional verdict, my assessment does not change that verdict; it sharpens the condition and provides a concrete way to test it.","tokens_in":5330,"tokens_out":9597,"duration_ms":111055,"concrete_test":"On a held-out random sample of at least 500 comments from the test set, have three annotators (not authors) rate whether each comment expresses amusement, using the same thresholding as the authors. Then evaluate the already-fitted logistic regression (trained on GPT-3.5 labels) on these human labels: report accuracy and AUC with 95% confidence intervals, and report human–human and GPT–human kappas. If the CI for AUC includes 0.5, the claim about human amusement is unsupported; if it excludes 0.5, the weak-signal claim has direct evidence.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim requires that the GPT-3.5 binary labels are a usable proxy for human amusement. The paper's own validation undermines this: Cohen's Kappa between GPT-3.5 and the five authors is 0.158–0.398 (Table 1), described in the Conclusions as 'likely very imperfect.' Moreover, the few-shot prompt asks GPT-3.5 to rate whether 'the comment is funny,' while the construct of interest is whether the commenter was amused by the game; these are not the same judgment. Every reported result—54.5% accuracy, coefficients in Table 2—is a prediction of these GPT labels. The fact that a model using only number of guesses is within 0.2% of the full model suggests the learned signal could be a convention about game length (e.g., comments reporting the score) rather than humor. Without an evaluation of the fitted model on held-out human amusement labels, the abstract's 'user amusement ... can be predicted computationally' is not established.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper explores whether objective properties of a Wordle game (number of guesses, reductions in the space of possible answers, Levenshtein and GloVe distances between consecutive guesses, and a regression-based measure of intrinsic word funniness) can predict whether a Reddit comment about that game expresses amusement. The authors scrape about 80,000 game-comment pairs from r/Wordle, obtain binary amusement labels from GPT-3.5 via few-shot prompting, validate those labels against five human annotators (reporting weak to moderate Cohen's Kappa), and train logistic regression and neural network classifiers on balanced subsampled data. The best model achieves 54.5% correct classification on a balanced test set of 15,000 examples, with a univariate model using only game length within 0.2% of the full model. The paper concludes that user amusement in Wordle is computationally predictable to a modest extent.","tokens_in":5565,"tokens_out":7576,"duration_ms":59117,"significance":"If the central claim were supported, the paper would provide a modest but interesting empirical result: objective game statistics carry a weak signal about perceived humor, complementing prior work on perceived move brilliance in chess. The authors are transparent about their limitations, report coefficient tables with standard errors and p-values, and compare against a univariate baseline, which are strengths. The feature set is grounded in psycholinguistic humor norms and previous Wordle-behavior studies. However, the current experiments only establish that GPT-3.5's amusement labels are predictable; the abstract's claim about 'user amusement' is not yet supported because the agreement between GPT-3.5 and human annotators is weak (Kappa 0.158–0.398), and no evaluation of the final model on human labels is provided. The negligible difference from the game-length-only baseline further tempers the substantive significance of the feature-level findings.","major_comments":[{"comment":"The central claim that 'user amusement at Wordle games can be predicted computationally' is not established by the reported experiments, because the prediction target is GPT-3.5's binary label, not human amusement. Table 1 reports Cohen's Kappa between GPT-3.5 and each of the five human annotators in the range 0.158–0.398 (weak agreement), and the Conclusions describe the rating as 'likely very imperfect.' The 54.5% accuracy and the coefficients in Table 2 therefore characterize the model's ability to predict GPT-3.5's judgments; they do not, by themselves, show that human amusement is predictable. I recommend either evaluating the fitted logistic regression on held-out human amusement labels (e.g., the author-annotated comments, supplemented if necessary), or explicitly limiting the paper's claims to prediction of GPT-annotated amusement.","section":"Abstract and Conclusions"},{"comment":"The full model's improvement over a single-feature baseline is negligible in practical terms: a univariate logistic regression using only num_possible_guesses_length achieves a correct classification rate that is only 0.2% lower than the full model. With a test set of 15,000 balanced examples, a chi-squared test can detect such a small difference as statistically significant, but the incremental predictive value of the remaining features is essentially zero. Since the paper's discussion interprets multiple features (Levenshtein distance, GloVe distance, intrinsic word humor) as contributing to amusement, the authors should report the univariate model's accuracy explicitly and provide effect sizes (e.g., accuracy difference, AUC) for the full versus univariate model. They should also consider whether the game-length signal is an artifact of comments that simply report the score rather than expressing amusement.","section":"Model and Performance / Results"},{"comment":"The operational definition used to obtain labels is inconsistent with the paper's construct. The system prompt defines humor as 'the extent that the commenter is amused by the Wordle game,' but the instruction to GPT-3.5 asks whether 'the comment is funny' (0 or 1). These are different judgments: a comment can be funny without the commenter being amused by the game (e.g., a witty complaint), and a commenter can be amused while writing a non-funny comment. Unless the authors provide evidence that these two phrasings are empirically interchangeable, the reported labels are ambiguous with respect to the construct of interest. The prompt should be aligned with the intended construct, or the paper should acknowledge and justify the conflation.","section":"Appendix (few-shot prompt)"},{"comment":"The dataset construction introduces avoidable label noise: the authors state that the scraped replies 'often, though not always, are reactions to the Wordle game in the original post.' If a substantial fraction of replies are not responses to the game posted in the original post, their amusement cannot be predicted from game features, which dilutes the measured signal. The paper should report the proportion of replies that are direct top-level reactions to the game post, and either restrict the analysis to those or analyze the sensitivity of the results to this filtering.","section":"Dataset"}],"minor_comments":[{"comment":"The phrase 'verify that GPT-3.5's labels roughly correspond to human labels' overstates the agreement reported in Table 1 (Kappa 0.158–0.398); suggest 'weakly correspond' or similar.","section":"Abstract"},{"comment":"The reported 'R2 of 0.37022' should be formatted as 'R^2 = 0.370' and the decimal precision is excessive; also, the RMSE of 7.67 over a range of 72.3 is a useful effect-size indicator and should be stated as such.","section":"Feature: Intrinsically Funny Words"},{"comment":"The claim that 'the significant p-values are generally very small, and would be robust to a Bonferroni correction' is not accurate for all significant entries in Table 2: the coefficient for 'num possible guesses reduction mean' has p = 0.0396, which exceeds the Bonferroni threshold of approximately 0.00385 for 13 predictors, and 'levenshtein distance max' (p = 0.0577) is not significant at the 0.05 level. Please specify which p-values survive the correction.","section":"Results"},{"comment":"The sentence 'Performance on the test set, for all the settings described above, was 54%±0.5%, using all the settings above' is ambiguous: it is unclear whether this is the best, average, or representative performance across settings. Please report the selection procedure and the variability across the three architectures and regularization options.","section":"Model and Performance"},{"comment":"The reference 'CMLOEGCMLUIN. 2012. Relative frequencies of English phonemes' appears to contain a garbled author name; please verify the source.","section":"Reference list"},{"comment":"The interpretive claim that a larger Levenshtein distance 'might imply the user is having fun with different guesses' is speculative and not supported by the data; consider softening or removing such motivational language from the feature descriptions.","section":"Feature: Luck or Skill"}],"recommendation":"major_revision","confidential_remarks":"The paper is an honest, exploratory study on an interesting question, and the reporting of uncertainty is commendable. The main gap is construct validity: the fitted model predicts GPT-3.5's labels, not human amusement, and the authors' own human-agreement data are weak. This is fixable either by validating on human labels or by carefully reframing the claims. I would not reject the paper, provided the authors address this gap and the univariate-baseline issue in a revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear [Colleague],\n\nHere's the short version: this paper shows that a handful of hand-crafted game features (number of guesses, reduction in possible answers, word-funniness norms, etc.) can predict whether GPT-3.5 calls a Reddit comment about a Wordle game 'amusing,' with 54.5% accuracy on a balanced test set of 15,000 games, against a 50% baseline. The effect is statistically significant and the authors are appropriately modest in the conclusions. The problem is the abstract's last sentence goes a step further: 'user amusement is predictable.' That isn't what they measured.\n\nWhat's new: nobody has tried to predict amusement at Wordle games from game statistics before. The dataset scraping and the feature engineering are straightforward but sensible, and the statistical reporting is unusually careful for a short paper—p-values, standard errors, a Bonferroni mention, and a balanced test set. The funniness regression using Westbury & Hollis features is also a reasonable transfer.\n\nThe soft spots. The target label is GPT-3.5's binary 'amusing/not amusing' judgment. The agreement between GPT-3.5 and the five human annotators is weak (Cohen's Kappa 0.16–0.40, Table 1). The authors acknowledge this in the Conclusions, but the abstract and the final claim don't carry the same caveat. The stress-test note about the prompt is mostly wrong: the system prompt actually defines humor as 'the extent that the commenter is amused by the Wordle game.' So the construct is not as mismatched as the note claims. However, the weak Kappa still means we don't know how well GPT's labels track human amusement, and every result in the paper is about those labels.\n\nSecond, the full model is only 0.2% better than a univariate model using just the number of guesses. That tells you the signal is dominated by game length, not by the clever features. The authors report this honestly, but it undercuts the 'creativity infused through humor' framing.\n\nThird, no code or data is released, and the few-shot examples are referenced but not shown in the text (they're in an appendix that isn't included here). That makes it hard to reproduce or evaluate the prompt quality.\n\nWho is this for? People working on computational humor or LLM-based annotation. It's a useful data point, but not a smoking gun. It deserves a serious referee—the methodology is sound enough, and the label-validity issue is exactly what peer review should hammer on. I'd recommend sending it to review, with the expectation that the authors need to either validate against human labels or rewrite the claims to be about GPT-3.5's judgments.","headline":"Weak but real signal for predicting GPT-3.5-labeled amusement in Wordle; the human-label validation is too thin to support the 'user amusement' framing.","tokens_in":6089,"tokens_out":2846,"would_cite":false,"duration_ms":24969,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that whether a Wordle game amuses Reddit users can be predicted weakly but above chance from game statistics alone, with a logistic regression reaching 54.5% correct classification against a 50% baseline.","keywords":["Wordle","amusement prediction","humor detection","few-shot GPT-3.5 classification","logistic regression","game features","computational creativity","Reddit comments"],"falsifier":"Collect direct amusement ratings from human annotators on a few thousand Wordle games and retrain the same logistic regression on those labels; if accuracy is not significantly above 50%, or if the features that matter change substantially, the claim that game features predict user amusement would be unsupported.","tokens_in":5152,"feed_emoji":"😂","tokens_out":7443,"duration_ms":60082,"temperature":0.7,"pith_summary":"The paper claims that user amusement at Wordle games can be predicted from game statistics, to a modest but real extent. The authors scrape about 80,000 Reddit reactions to Wordle games, label them as amused or not with a few-shot GPT-3.5 classifier, and verify that those labels coincide weakly with human annotations. They then train a logistic regression on aggregate game features—number of guesses, reductions in the number of possible answers, Levenshtein and GloVe distances between guesses, and predicted intrinsic funniness of guessed words—and reach 54.5% correct classification on a balanced test set of 15,000 games, against a 50% baseline. The authors read this as evidence that a measurable aspect of the humor and creativity infused into Wordle games lives in the game itself, not only in the comment threads.","feed_headline":"Game stats predict Wordle amusement at 54.5 percent","feed_subtitle":"Shorter games, bigger answer-space drops, and funnier words carry a small but real signal of player amusement.","key_machinery":"The central mechanism is a feature set computed from each Wordle game transcript. For every guess the authors compute the reduction in the number of words still consistent with the color feedback, the Levenshtein and GloVe distances to the previous guess, and a predicted 'intrinsic funniness' score for the guessed word; the funniness score comes from a linear regression trained on 4,858 words with human humor ratings using features such as category-defining vectors and valence-arousal-concreteness estimates. These aggregate statistics are fed into a logistic regression classifier whose coefficients reveal which game properties shift the predicted probability of an amused reaction, and the label side is a few-shot GPT-3.5 model that turns each Reddit comment into a binary amused/not-amused judgment. This pipeline—game transcript to features to probability—is what carries the claim.","core_discovery":"On its own terms, the core discovery is that amusement at Wordle games is predictable to a very modest extent from the games: a logistic regression with only aggregate game-level features classifies GPT-labeled amused versus not-amused reactions correctly 54.5% of the time on a balanced 15,000-game test set, where guessing the majority class would give 50%. The features that matter are game length, the reduction in the number of words still possible after guesses, edit distance between consecutive guesses, and how funny the guessed words are on their own. In particular, shorter games, larger final reductions in the candidate set, and a larger distance between the last two guesses predict more amusement, while average edit distance predicts less, and GloVe semantic distance has no measurable effect. The authors acknowledge the amusement labels are likely very imperfect, since GPT-3.5's agreement with human annotators is weak, and hypothesize that better labels would improve predictive performance.","pith_inferences":["Editorial extension: because the labels are GPT-3.5's, the 54.5% figure may partly measure how predictable GPT-3.5's humor judgments are; a direct human-labeled replication is needed to confirm the signal attaches to human amusement.","Editorial extension: if the signal is real, puzzle generators could optimize for amusement by selecting Wordle answers that produce short games with dramatic collapses of the candidate set, turning humor into a tunable game-design objective.","Editorial extension: the small gap over baseline suggests most of the amusement signal lives in the comment text and social context rather than the game transcript, so combining game features with text features is a natural next step.","Editorial extension: the same feature family could be tested on other constrained word games (Quordle, Semantle, or custom Wordle variants) to see whether the pattern generalizes beyond one game's answer list."],"forward_implications":["Shorter games and large late reductions in the number of possible answers both predict amusement, suggesting that visibly lucky or skilled solves are a main emotional trigger.","A one-standard-deviation increase in last-guess Levenshtein distance raises predicted amused probability by about 4%, and a one-standard-deviation increase in the final answer-space reduction raises it by about 2%.","Game length alone is within 0.2 percentage points of the full model's accuracy, so the number of guesses is a near-sufficient summary of the signal in this feature set.","Average Levenshtein distance predicts less amusement while last-guess distance predicts more, so the final dramatic jump to the answer matters more than overall guess style.","The authors hypothesize that higher-quality amusement labels from humans would make the game features more predictive than they appear here."],"supporting_citations":[{"why":"Supplies the 4,997-word humor norms used to train the intrinsic-funniness predictor.","marker":"(Engelthaler and Hills 2018)"},{"why":"Supplies the nineteen word features and category-defining vectors shown to predict word funniness.","marker":"(Westbury and Hollis 2019)"},{"why":"Supplies the GloVe embeddings used to compute semantic distance between guesses.","marker":"(Pennington, Socher, and Manning 2014)"},{"why":"Defines the edit distance used to measure orthographic distance between consecutive guesses.","marker":"(Levenshtein 1966)"},{"why":"Provides the antecedent approach of predicting human perception of game moves from game-tree features, which this paper adapts to amusement.","marker":"(Zaidi and Guerzhoy 2024)"},{"why":"Motivates using semantic and orthographic distance between guesses as gameplay-bias features.","marker":"(Liang et al. 2024)"},{"why":"Supplies the extrapolated valence, arousal, dominance, and concreteness estimates used in the funniness model.","marker":"(Hollis, Westbury, and Lefsrud 2017)"}],"fun_headline_variants":["Game stats predict Wordle amusement at 54.5%","Shorter Wordles, bigger answer drops predict fun","AI labels help spot amusing Wordles, 54.5% accurate","Weak but real signal in Wordle amusement found"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that GPT-3.5's binary amused/not-amused judgments on Reddit comments are a trustworthy stand-in for what human players actually find amusing, even though agreement between GPT-3.5 and the human annotators is weak (Cohen's kappa about 0.16–0.40).","fun_headline_variants_meta":{"raw":{"variants":["Game stats predict Wordle amusement at 54.5%","Shorter Wordles, bigger answer drops predict fun","AI labels help spot amusing Wordles, 54.5% accurate","Weak but real signal in Wordle amusement found"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001035,"raw_usage":{"total_tokens":4318,"prompt_tokens":868,"completion_tokens":3450,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":484,"completion_tokens_details":{"reasoning_tokens":3382}},"tokens_in":484,"tokens_out":3450,"duration_ms":24582,"temperature":1.0,"reasoning_tokens":3382,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T10:41:39.136657+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Collect direct amusement ratings from human annotators on a few thousand Wordle games and retrain the same logistic regression on those labels; if accuracy is not significantly above 50%, or if the features that matter change substantially, the claim that game features predict user amusement would be unsupported.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the 4,997-word humor norms used to train the intrinsic-funniness predictor."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the nineteen word features and category-defining vectors shown to predict word funniness."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the GloVe embeddings used to compute semantic distance between guesses."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the edit distance used to measure orthographic distance between consecutive guesses."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the antecedent approach of predicting human perception of game moves from game-tree features, which this paper adapts to amusement."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Motivates using semantic and orthographic distance between guesses as gameplay-bias features."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the extrapolated valence, arousal, dominance, and concreteness estimates used in the funniness model."}],"review_version":1}