{"id":"4e496fd5-6647-470b-8a1c-292b601d541f","arxiv_id":"2507.00769","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Small reward models trained on LitBench reach 78% agreement with upvote-derived human preferences in creative writing, beating all zero-shot LLM judges tested.","lead":"LitBench is a new benchmark of 2,480 paired Reddit stories, labeled by upvote preferences, for measuring how well AI systems judge creative writing. Small trained reward models reach 78% agreement and beat the best off-the-shelf judge, Claude-3.7-Sonnet, at 73%.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central accuracy numbers are agreements with a filtered upvote proxy; the 2,480 test labels are never directly validated against human pairwise judgments, so 'human-labeled' and 78% agreement overstate what is shown.","rationale":"I read the paper in good faith: it is a substantial, well-constructed resource with plausible findings, and the reader's conditional verdict is appropriate. The most load-bearing concern is exactly the one the reader identified: upvote counts are the sole source of preference labels for the test set, and the paper never directly validates those labels against independent human pairwise judgments on the test set itself. The online human study in Section 5 is genuine supporting evidence, but it only validates a trained reward model's rankings on 64 newly generated stories, not the benchmark labels. If the labels are a poor proxy for human literary preference, then the 73% and 78% accuracy numbers are not measurements against human taste, and the phrase 'human-labeled' overstates the data. This is a correctness risk, not merely a framing issue, because the central scientific claim is that LitBench enables reliable automated evaluation of creative writing. The concern is not fatal: the ablation in Figure 7 demonstrates internal consistency of the curation pipeline, and the resource may still be useful as a proxy benchmark if the proxy is disclosed clearly. The concrete experiment I propose—direct pairwise human rating of a stratified sample of LitBench test pairs—would settle the question. Other concerns, such as missing error bars or the unfinished citation, are real but secondary. I agree with the reader's weakest-assumption analysis and recommend keeping the conditional verdict, pending the label-validation test.","tokens_in":11397,"tokens_out":3012,"duration_ms":40291,"concrete_test":"Sample 150–200 LitBench test pairs stratified by upvote margin and length difference; recruit at least three independent human raters per pair to judge which story is better creative writing, using the evaluation dimensions from the appendix prompts (e.g., imagery, tension, originality, emotional impact). Compute majority-human agreement with the upvote-derived labels and inter-annotator agreement (e.g., Fleiss' kappa). If majority-human agreement with upvote labels is near 57–60%, the 'human-labeled' claim and headline accuracies need re-scoping; if it reaches roughly 75% or higher, the proxy is substantially validated. Report the same human agreement on the paired stories when presented in permuted order.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim—that LitBench is a benchmark of human-labeled story comparisons and that a trained reward model reaches 78% human agreement—rests on equating Reddit upvotes with human preference for creative writing quality. Section 3.2 constructs labels purely from upvote counts (minimum 10 upvotes, ≥25% upvote differential, and the higher-upvote story published later), but the test set is never validated against direct human pairwise ratings. Section 5's online study (46 crowdworkers, 64 GPT-generated stories, 10–13 annotators per pair) tests only whether annotators prefer the RM's ranking on newly generated stories; it does not establish that the 2,480 upvote-derived test labels match what human readers think. The Limitations (Section 7) explicitly concede that upvotes may encode exposure, demographic skew, altruism, and other factors rather than literary quality. The headline numbers (73% for Claude-3.7-Sonnet, 78% for the BT-RM) are therefore accuracies against an unvalidated proxy. The curation ablations in Figure 7 show the filters improve agreement with the proxy, but they cannot by themselves show the proxy is human preference. A direct label-validation experiment is missing and is load-bearing for the benchmark's core interpretation.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"LitBench is a new benchmark and training dataset for evaluating automated judges of creative writing. The authors collect story pairs from Reddit's r/WritingPrompts, constructing labels from upvote counts with several curation filters (engagement threshold, upvote differential, temporal ordering, and length-balancing). The test set contains 2,480 pairs and the training set 43,827 pairs. The paper benchmarks zero-shot LLM judges, trains Bradley-Terry and generative reward models, and conducts an online human study on 64 newly generated stories. Headline results are that a fine-tuned Llama-8B Bradley-Terry reward model reaches 78% agreement with the labels, surpassing the best zero-shot judge (Claude-3.7-Sonnet at 73%), and that chain-of-thought distillation hurts generative reward model accuracy. The authors release the dataset and models.","tokens_in":11745,"tokens_out":3020,"duration_ms":36631,"significance":"If the benchmark labels are a valid proxy for human creative-writing preferences, LitBench would be a valuable resource: it is the first standardized pairwise creative-writing evaluation set of its scale, it ships with a training corpus, and the public release of code and models supports reproducibility. The paper also makes two useful empirical observations: small fine-tuned reward models can beat much larger zero-shot judges, and chain-of-thought supervision harms generative reward models in this domain. The ablations on curation filters (Figure 7) are a good-faith attempt to justify design choices. However, the central interpretation of the accuracy numbers depends on an assumption that is explicitly acknowledged but not directly tested: that upvotes, after filtering, encode human preference for writing quality. Because the test labels are never validated against independent human pairwise judgments, the headline 'human agreement' figures are currently agreement with a proxy.","major_comments":[{"comment":"The test set is called 'human-labeled' in the abstract and throughout, but the labels are derived solely from upvote counts (minimum 10 upvotes, ≥25% upvote differential, higher-upvote story published later). Section 7 explicitly concedes that upvotes may encode exposure, demographic skew, altruism, and other factors rather than literary quality. The online human study in Section 5 validates a trained reward model on 64 newly generated stories, not on the 2,480 test pairs. Consequently, the 73% and 78% accuracy figures in Figure 4 are agreements with the upvote proxy, not with independent human judgments. This is load-bearing for the benchmark's core interpretation. The authors should directly validate a random sample of test pairs with human pairwise ratings (reporting agreement, e.g., accuracy or Cohen's kappa) or, failing that, reword all 'human-labeled' and 'human agreement' claims to say 'upvote-preference proxy.'","section":"§3.2, §5, §7"},{"comment":"The paper reports accuracy and preference percentages without any uncertainty quantification. With 2,480 test pairs, a difference between 78% and 73% may be significant, but the reader cannot tell; with only 64 generated pairs in the human study, the 57% versus 41% preference result needs a confidence interval or a significance test (e.g., a binomial test per pair or a mixed-effects model). Without this, the claims that trained models 'outperform' zero-shot judges and that the human study 'confirms' alignment are not statistically grounded. Please add error bars or confidence intervals to the key accuracy figures and report the outcome of a significance test for the human study.","section":"Figure 4, Table A.1, Figure 8"},{"comment":"The claim that the test set enables 'true zero-shot evaluation' is undermined for models with training cutoffs after January 2023. The test stories are all posted after January 2, 2023, and several evaluated models, such as GPT-4.1 and Claude-3.7-Sonnet, are likely to have been trained on data from that period. The paper should either restrict the zero-shot claim to models with confirmed earlier cutoffs, or include a leakage analysis (e.g., checking whether judge accuracy varies by story date or by model training cutoff).","section":"§3.3, §4"}],"minor_comments":[{"comment":"In the sentence 'In domains where human where ground truth are usually collected from human raters, LLM judges are often used,' the phrase 'where human where ground truth' appears to be a typo; it should read 'where ground truth labels are usually collected from human raters.'","section":"§1"},{"comment":"The paper repeatedly refers to 'ROGUE' metrics; the standard acronym is ROUGE. Please correct this throughout, including the related work discussion.","section":"§2.1"},{"comment":"'This entire process is performed independently for both our benchmark and training dataset' is redundant with the preceding paragraph; consider removing or condensing. Also, 'reduce the affect of noise' should be 'reduce the effect of noise.'","section":"§3.2"},{"comment":"The text says the best LLM judge (Claude-3.7-Sonnet) 'performs at chance' in the human study, but the details of how Claude-3.7-Sonnet was evaluated on the generated stories are not given. Please specify the evaluation protocol and the exact chance-level accuracy, or remove the comparison if it is not central.","section":"§5 and Figure 8"},{"comment":"There is a missing reference in the motivation paragraph: 'prompts enable (a) introduction of criteria ... (CITE).' Please fill in the citation.","section":"Appendix A.2"},{"comment":"The paper contains minor typographical errors, including 'GemRM-CoT' for 'GenRM-CoT', 'Across across' in Section 6, and inconsistent hyphenation of 'Bradley-Terry'/'Bradley Terry.' A careful proofreading pass is recommended.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The manuscript addresses a real gap and the resource could be useful, but the gap between the upvote-derived labels and the 'human agreement' framing is too large to ignore. I would ask the authors to add a direct label-validation experiment or to substantially reword the claims. I also recommend confirming that the statistical comparisons hold up once error bars are added. The paper is otherwise within the journal's scope."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"LitBench is a useful new resource: a pairwise creative-writing benchmark (2,480 pairs) and a 43k-pair training corpus, with explicit length and temporal debiasing, plus an ablation showing those curation steps matter. The empirical finding that distilled chain-of-thought hurts generative reward models for this task is genuinely new and worth knowing. The paper also does the right thing by releasing ids, code, and a MIT-licensed training set, so the resource is reproducible.\n\nThe caveat is real, and the authors state it themselves in the Limitations: the labels come from Reddit upvotes, not from direct human pairwise judgments on the test set. The 78% and 73% figures are agreements with a filtered upvote proxy, not with independent human raters on those 2,480 pairs. The online human study goes some way toward showing that a reward model trained on this proxy selects better stories when judged fresh, but it does not validate the test-set labels. The phrase \"human-labeled\" in the abstract overstates what was measured.\n\nThe missing pieces are fixable: a small direct-annotation study on a sample of test pairs would let the authors say how well the proxy matches explicit pairwise preferences, and error bars or significance tests would make the model comparisons credible. The curation ablations are a good start but they only show the filters improve agreement with the proxy, not that the proxy is human taste. Minor polish issues (an unfinished citation, a mis-cited SHP reference, a typo in the intro) are easy to clean up.\n\nAll of this is in proportion: this is a solid benchmark paper, not a flawed one. The central resource is likely useful to anyone working on LLM evaluation or reward modeling for open-ended generation, as long as the accuracy claims are read as proxy agreement. The finding on CoT is a real signal for practitioners.\n\nI would send this to peer review. It is important enough and the methodology is careful enough to deserve referee time, with a request for the label-validation experiment and tightening of the wording.","headline":"Useful benchmark, honest limitations, but the headline accuracy numbers measure agreement with a filtered upvote proxy, not independent human taste.","tokens_in":12187,"tokens_out":2130,"would_cite":true,"duration_ms":24100,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper introduces LitBench, the first standardized benchmark and paired dataset for creative-writing verification, and claims that a small trained reward model reaches 78% human agreement, beating all zero-shot LLM judges tested.","keywords":["creative writing evaluation","LLM-as-a-judge","reward models","Bradley-Terry preference model","generative reward models","Reddit upvotes as human preference","chain-of-thought reasoning","r/WritingPrompts"],"falsifier":"Take a random sample of LitBench test pairs, have a panel of expert readers rate each pair directly, and compare those ratings to the upvote-derived 'chosen' labels. If human pairwise agreement with the upvote labels is near chance, or if the 78%-accurate verifier agrees with the upvote labels but not with the direct human ratings, then the benchmark measures agreement with a voting proxy rather than with human literary taste.","tokens_in":11167,"feed_emoji":"📝","tokens_out":13559,"duration_ms":128458,"temperature":0.7,"pith_summary":"The paper tries to make creative-writing evaluation measurable, the way math and code verification already are. It builds LitBench, a benchmark of 2,480 paired story comparisons from r/WritingPrompts whose preferred story is inferred from upvotes after heavy filtering, plus a 43,827-pair training corpus. It claims that a small fine-tuned Bradley-Terry reward model reaches 78% agreement with these human-preference labels, beating every zero-shot LLM judge tested, including the strongest, Claude-3.7-Sonnet, at 73%. The paper also reports that chain-of-thought reasoning hurts generative reward models on this task, and that an online study on newly generated stories confirms the trained verifier's rankings transfer. This matters because reliable automated judgment is the missing ingredient for steering creative-writing generation the way verifiers accelerated math and code.","feed_headline":"78% beats 73%: trained verifiers outrank LLM story judges","feed_subtitle":"Reddit upvote preferences, filtered and debiased, help small models judge new stories better than bigger chatbots.","key_machinery":"The central object is LitBench's label-construction pipeline, which converts noisy Reddit upvote counts into pairwise 'chosen versus rejected' story labels by removing low-engagement stories, dropping pairs with small upvote gaps, requiring that the chosen story be published later, and balancing the length distribution so preference is not confounded with length. This pipeline produces the training signal for a Bradley-Terry discriminative reward model (a linear head on a Llama-8B backbone) that scores each story independently and is trained so the chosen story scores higher; the benchmark then measures how well any judge, trained or zero-shot, reproduces those labels.","core_discovery":"The central claim is that creative writing has convergent human preferences, those preferences can be captured at scale from Reddit upvotes, and a verifier trained on such preferences can be evaluated and used. LitBench supplies the evaluation: 2,480 pairwise comparisons built from stories posted after January 2023, filtered for engagement, paired only when the higher-upvoted story was published later, and pruned on length so that 'chosen' stories are not simply the longer ones. On that test set, a Bradley-Terry reward model fine-tuned on LitBench's 43,827-pair training corpus reaches 78% accuracy, and the best zero-shot judge, Claude-3.7-Sonnet, reaches 73%. The paper also reports that adding chain-of-thought reasoning degrades generative reward models for this task (72% versus 78%), and that an online human study on 64 newly generated stories confirms the trained verifier ranks quality: humans chose the reward-model-preferred story 57% of the time versus 41% for the rejected story.","pith_inferences":["If the upvote-preference assumption holds, the same curation recipe could be applied to other writing communities, poetry, or serialized fiction to build additional verifier datasets without fresh human annotation.","Because the paper's own limitation notes that Reddit demographics skew male, educated, and middle-aged, the 78% figure may overstate alignment with other reader populations; measuring verifier accuracy on ratings stratified by reader demographics would test this.","The explanation-text statistics, which found plot discussion most predictive of correct verdicts, suggest that rubric-guided prompts might outperform free-form reasoning for LLM judges, a testable extension of the paper's prompt-optimization results.","The 41% disagreement with the rejected story in the online study implies pairwise upvote labels leave substantial room for richer supervision, such as rubric-based ratings or expert critiques, before automatic rewards can fully capture literary taste."],"forward_implications":["Small open-source verifiers (1B-8B) trained on LitBench can match or beat large proprietary LLM judges on creative-writing evaluation at a fraction of the cost.","Zero-shot LLM judges should not be treated as reliable for story quality: only frontier proprietary models in this comparison clear roughly 70% agreement, while smaller models hover near chance.","Chain-of-thought reasoning is not automatically helpful for judging narratives; in this domain it lowered generative-reward-model accuracy, so its use should be tested rather than assumed.","A verifier that scores well on LitBench generalizes to newly generated stories, so the benchmark can serve as a reward signal for steering creative-writing generators.","Because the test set is drawn from post-2023 stories, it provides a genuinely zero-shot evaluation for models with earlier training cutoffs."],"supporting_citations":[{"why":"Supplies the base set of r/WritingPrompts post IDs from which the test set's stories are searched and collected.","marker":"[Fan et al., 2018b]"},{"why":"Provides the temporal-pairing methodology that pairs stories so the higher-upvoted story was also published later, mitigating exposure bias.","marker":"[Ethayarajh et al., 2022b]"},{"why":"Supplies the 2048-token length cutoff and the observation that creative-writing reward models are hard to build, which the curation filters adopt.","marker":"[Chung et al., 2025]"},{"why":"Defines the paired-comparison objective used to train the best-performing discriminative reward model.","marker":"[Bradley and Terry, 1952]"},{"why":"Origin of the core assumption that Reddit upvotes encode human preference, which the entire label construction rests on.","marker":"[Ethayarajh et al., 2023]"},{"why":"Evidence that expert writers converge in rating creative products, the psychological basis for treating aggregated upvotes as a quality signal.","marker":"[Amabile, 1982]"},{"why":"Prior work on generative reward models that the paper's GenRM and GenRM-CoT experiments build on and compare against in the creative-writing setting.","marker":"[Mahan et al., 2024]"},{"why":"Defines chain-of-thought prompting, the technique whose degradation on story judgment the paper documents.","marker":"[Wei et al., 2022]"},{"why":"Documents non-quality motives for upvoting, the confound the paper cites against its upvote-preference assumption.","marker":"[Kassaeyan, 2016]"}],"fun_headline_variants":["Reddit upvotes beat chatbots at judging stories","Trained verifiers outrank zero-shot LLM story judges","Why chain-of-thought hurts generative story judges","LitBench: Reddit preferences train better story verifiers","Small verifiers beat big LLMs on story judgment"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that upvote counts on r/WritingPrompts, after the curation filters, indicate human preference for creative writing quality; the paper never directly validates the benchmark's labels against fresh human pairwise ratings, only a trained verifier built from them.","fun_headline_variants_meta":{"raw":{"variants":["Reddit upvotes beat chatbots at judging stories","Trained verifiers outrank zero-shot LLM story judges","Why chain-of-thought hurts generative story judges","LitBench: Reddit preferences train better story verifiers","Small verifiers beat big LLMs on story judgment"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000269,"raw_usage":{"total_tokens":1669,"prompt_tokens":1041,"completion_tokens":628,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":657,"completion_tokens_details":{"reasoning_tokens":551}},"tokens_in":657,"tokens_out":628,"duration_ms":7359,"temperature":1.0,"reasoning_tokens":551,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T21:07:04.806564+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a random sample of LitBench test pairs, have a panel of expert readers rate each pair directly, and compare those ratings to the upvote-derived 'chosen' labels. If human pairwise agreement with the upvote labels is near chance, or if the 78%-accurate verifier agrees with the upvote labels but not with the direct human ratings, then the benchmark measures agreement with a voting proxy rather than with human literary taste.","supporting_citations":[],"review_version":1}