{"id":"ec7a0fb4-05b8-4d47-bd39-13bf29d127a6","arxiv_id":"2608.03340","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Commonsense benchmark scores are task-dependent predictors of downstream performance: they transfer reliably only to a few tasks (False Beliefs, TRIP), and revised benchmarks do not improve criterion validity.","lead":"This paper tests whether scores on popular commonsense benchmarks (WinoGrande, HellaSwag, Social IQA, Physical IQA) predict how well 23 language models do on eight downstream reasoning tasks. It finds that benchmark rankings transfer only to a narrow subset of tasks such as False Beliefs and TRIP, and that reworked benchmark versions do not improve prediction.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Downstream criterion reliability: tiny filtered sets (CEI n=31, TimeDial n=60) may explain the narrow-transfer result; low item count attenuates correlations.","rationale":"The paper's central finding is a negative, task-dependent pattern: only two or three downstream tasks show consistent transfer. The most load-bearing assumption is that the eight downstream tasks are reliable and comparable criteria. The reader flagged this, and I agree. My concern sharpens it: the small item counts after manual filtering (CEI n=31, TimeDial n=60) mechanically inflate the variance of model scores, attenuating Spearman correlations toward zero. This directly threatens the comparison across tasks—the very evidence for \"task-dependent\" usefulness. The paper acknowledges small sets in Limitations but does not quantify the effect. A split-half reliability check or subsampling experiment would settle whether the pattern is real. Since this is a conditional concern that the current evidence does not resolve, the reader's CONDITIONAL verdict is appropriate; no change is needed.","tokens_in":13488,"tokens_out":3435,"duration_ms":42680,"concrete_test":"Compute split-half reliability for each downstream task by randomly partitioning items into two halves, computing each model's score on each half, and correlating the two score vectors across the 23 models (or bootstrap the per-model scores). Then apply attenuation correction: rho_corrected = rho_obs / sqrt(reliability_downstream) (assuming benchmark reliability high) and re-examine which tasks remain significantly correlated. Alternatively, subsample False Beliefs and TRIP to n≈31 (matching CEI) and check whether their correlations fall to CEI levels; if they do, the narrow-transfer result is an artifact of item count. Also report inter-annotator agreement for the manual filtering in Appendix B.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that CS benchmarks transfer only to False Beliefs and TRIP—rests on comparing Spearman correlations across eight downstream tasks, but those scores are estimated from very few items after manual filtering (Table 2: CEI n=31, TimeDial n=60, Indirect Requests n≈64; Appendix B). Score reliability scales with item count; classical attenuation will compress correlations for short tasks toward zero, so the observed \"narrow transfer\" may simply reflect which tasks have enough items to yield stable model rankings. False Beliefs (n≈192) and TRIP (n≈100) are among the larger sets, and they are exactly the tasks that show consistent associations. The authors do not report split-half reliability or a sensitivity analysis for the manual filtering, nor do they quantify the impact of small n on the correlation comparisons. Without this, the task-dependence of the headline result is confounded with task-score reliability.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper asks whether widely used commonsense benchmarks have criterion validity: do model scores on these benchmarks predict performance on downstream tasks that require implicit social, pragmatic, temporal, or physical reasoning? The authors evaluate 23 open-weight LLMs from six families on four benchmark/revision pairs (WinoGrande/WinoWhat, HellaSwag/GoldenSwag, Social IQA/filtered Social IQA, Physical IQA/Physical IQA-RUP), three non-commonsense controls, and eight downstream tasks grouped into social/emotional, pragmatic, and event/physical axes. Using Spearman correlations, partial correlations controlling for model size, family, and controls, bootstrap confidence intervals, and leave-one-family-out ridge regression, they report that reworked benchmarks largely preserve original model rankings and do not improve downstream prediction; that consistent transfer is limited to False Beliefs and TRIP, with metric-specific gains for Presupposition, Implicature, Indirect Requests, and TimeDial; and that axis-specific transfer is weak. The paper concludes that commonsense benchmarks have task-dependent and uneven downstream usefulness and should not be read as broad evidence of pragmatic or contextual reasoning competence.","tokens_in":13733,"tokens_out":4709,"duration_ms":53920,"significance":"If the result holds, it is practically important: benchmark scores need external validation before being used for model selection or capability claims, and dataset repair alone does not automatically create criterion validity. The study has clear strengths: uniform log-likelihood scoring across heterogeneous tasks, inclusion of non-commonsense controls, bootstrap-based uncertainty estimation, FDR correction, leave-one-family-out prediction, and public code/data. It engages directly with the current benchmarking-crisis literature. However, the evidential core is narrower than the abstract and first paragraphs suggest. The most robust positive associations are raw correlations to two of eight downstream tasks; partial correlations are uniformly non-significant after FDR correction; LOFOCV uses only six family folds; and some downstream tasks have very few items after manual filtering. These limitations are acknowledged in Section 6 but are not quantitatively addressed. The paper remains a clearly reported and valuable empirical contribution, but the headline claim about narrow transfer needs additional reliability analysis or a more cautious formulation.","major_comments":[{"comment":"The central negative result—that transfer is limited to False Beliefs and TRIP—is vulnerable to criterion-score attenuation. After manual filtering, CEI has 31 items and TimeDial 60; Indirect Requests has about 64 (Table 2; Appendix B). Spearman correlations across 23 models computed on such short tasks have low reliability, and classical attenuation compresses correlations toward zero. The two tasks with the most consistent associations, False Beliefs (≈192) and TRIP (≈100), are among the larger sets, so the observed task-dependence may reflect test length rather than construct validity. The paper reports no split-half reliability, no item-level bootstrap by task, and no sensitivity analysis around the manual filtering decisions. Please add reliability estimates for each downstream task (e.g., split-half or rank split, bootstrap CIs for per-task model scores) and either use larger unfil","section":"§3.2, Table 2, Appendix B"},{"comment":"Section 4.3 reports that after controlling for model size, family, and non-CS controls, no partial Spearman correlation survives FDR correction, while the positive LOFOCV gains for False Beliefs and TRIP are based on only six family folds and are presented descriptively. Given 23 models clustered in six families, the evidence for 'consistent cross-family predictive validity' is thinner than the wording in Sections 4.2–4.3 suggests. The manuscript flags this in Section 6, but the conclusion still leans on raw correlations. To make the positive claims load-bearing, please present family-level scatterplots or bootstrap distributions for the two key tasks, report the per-family LOFOCV predictions rather than only pooled R2, and explicitly state the effective sample size and its consequences. If this is infeasible, the abstract and conclusion should downgrade 'consistent cross-family' to 'sug","section":"§4.3, Figure 3"},{"comment":"The axis-specific analysis is underpowered and the axes themselves are constructed by the authors. Panel (b) omits the pragmatic axis because all CS benchmarks are considered potentially relevant, so the global comparison mixes different question types; with only a handful of tasks per axis the bootstrap CIs are wide. The conclusion that 'transfer does not follow domains' may be correct, but the current test cannot cleanly separate a true absence of domain specificity from low power or arbitrary axis assignment. A clearly justified axis assignment (or a preregistered alternative) plus an item-level task-by-task analysis would strengthen this claim. At minimum, the wording should reflect that the axis test is exploratory.","section":"§4.4, Figure 4"}],"minor_comments":[{"comment":"'common sense benchmarks' should be 'commonsense benchmarks' for consistency with the rest of the manuscript.","section":"Figure 1 caption"},{"comment":"The manual filtering description gives only final n, not the number of removed items per task or the inter-annotator agreement if multiple annotators were involved. Adding a table with original n, removed n, final n, and annotation agreement would help readers assess the reliability of the filtered criteria.","section":"§3.2 / Appendix B"},{"comment":"The sentence 'the top-three models overlap never exceeds an average Jaccard score of .25' is confusing. Appendix C reports mean top-3 Jaccard values averaged over benchmark–task pairs, not a maximum. Please rephrase to indicate that the mean Jaccard overlap is around .25 and clarify whether this is averaged over accuracy, macro-F1, or both.","section":"§4.2 / Appendix C"},{"comment":"The sign convention for Δρ is stated in text but the equation itself would benefit from an explicit 'positive values favor the original benchmark' annotation. Also, the text uses 'ρOG' and 'ρRW' without defining these acronyms in the equation block.","section":"Eq. (2)"}],"recommendation":"major_revision","confidential_remarks":"The paper is within the journal's scope and likely of interest to the benchmarking and evaluation community. The main concern is not novelty or framing but the reliability of the small downstream criteria and the tension between raw-correlation claims and null partial correlations. I would not reject; the revisions should focus on adding reliability analyses and softening the cross-family generalization language. The authors' self-citations are appropriate given their prior work on benchmark criticism, and the data/code availability statement is a positive feature."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe paper does something genuinely useful: first direct look at whether reworked commonsense benchmarks improve downstream criterion validity, with non-commonsense controls and leave-one-family-out cross-validation. The revisions-preserve-rankings finding is solid—WinoWhat is the only real exception, and that result doesn't depend on the noisy small tasks. The LOFOCV gains for False Beliefs and TRIP are the most interesting thing in the paper.\n\nBut the headline \"narrow transfer\" claim is fragile. CEI has 31 items, TimeDial 60, Indirect Requests around 64. At those sizes, Spearman correlations are attenuated and model rankings are just noisy. The stress-test note is correct: the tasks that show consistent transfer (False Beliefs, TRIP) are the larger ones. You can't separate task-dependence from test length. The authors flag the small sets in limitations but never quantify the impact—no split-half reliability, no sensitivity analysis of their manual filtering. That filtering (Appendix B) is subjective and could systematically shift which tasks correlate; it's a real confound.\n\nThe partial correlations also undercut the specificity story: after controlling for size, family, and controls, nothing survives FDR. So the raw correlations partly reflect broad capability. The authors admit this, but it means the central claim rests on uncontrolled correlations plus a handful of LOFOCV results from six family folds—descriptive at best.\n\nThat said, the paper is honest, well-executed within its limits, and ships code and data. Bootstrap CIs and FDR correction are appropriate. The revision-preserves-rankings result is not affected by the small-n issue. So the paper deserves a serious referee, but the revision should add reliability analyses (split-half on the filtered sets) and a sensitivity check that repeats the correlation analysis without the filtered items.\n\nWho gets value: anyone building or using commonsense benchmarks, and anyone reading leaderboard claims. It's a useful corrective, not a breakthrough. Recommendation: referee it, with the reliability concerns on the table.\n\nBest,\n\n[You]","headline":"A solid, honest meta-evaluation whose headline 'narrow transfer' result is partly confounded by tiny filtered downstream sets; worth refereeing with a demand for reliability analyses.","tokens_in":14173,"tokens_out":2011,"would_cite":true,"duration_ms":25504,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Standardized commonsense benchmarks predict downstream performance unevenly: rankings transfer most consistently to False Beliefs and TRIP, while reworked benchmarks do not improve criterion validity.","keywords":["commonsense reasoning","benchmark validity","criterion validity","LLM evaluation","model rankings","pragmatic inference","downstream transfer","benchmark revision"],"falsifier":"Run the same benchmark-to-downstream prediction with downstream tasks at their original, unfiltered sizes (keeping all CEI and TimeDial items) and with at least 1,000 items per task; if CEI and Sarcasm correlations turn positive and leave-one-family-out R2 rises above zero, the narrow-transfer result is an artifact of manual filtering. Alternatively, if the HellaSwag–TRIP partial correlation survives FDR correction when family and size are matched more tightly, the claim of no incremental validity beyond general capability fails.","tokens_in":13406,"feed_emoji":"📊","tokens_out":7073,"duration_ms":72908,"temperature":0.7,"pith_summary":"This paper asks whether a high score on standardized commonsense benchmarks tells you anything about how a model will behave on tasks that require implicit social, pragmatic, temporal, or physical reasoning. The authors evaluate 23 open-weight models on four established benchmarks, four reworked versions, three non-commonsense controls, and eight downstream tasks, then compare model rankings and test whether benchmark scores predict downstream performance for models from unseen families. The answer is that benchmark rankings transfer unevenly: they generalize most consistently to False Beliefs and TRIP, provide partial or metric-specific signals for a few other tasks, and fail for CEI and Sarcasm. Revised, 'cleaner' benchmarks keep the original rankings and do not improve downstream prediction. So a model's commonsense-benchmark rank is task-dependent evidence, not broad evidence of commonsense competence.","feed_headline":"Commonsense benchmark rankings transfer to just 2 of 8 tasks","feed_subtitle":"Across 23 models, WinoGrande and HellaSwag predict False Beliefs and TRIP, not broad pragmatic reasoning; cleaned versions don't help.","key_machinery":"The central machinery is a cross-model criterion-validity comparison: Spearman correlations between benchmark and downstream rankings, partial Spearman correlations that residualize log model size, model family, and non-commonsense control scores, and leave-one-family-out cross-validation with ridge regression. A uniform log-likelihood multiple-choice scorer with label-prior calibration makes rankings comparable across heterogeneous tasks. The design tests whether benchmark-induced ordering of models transfers to external tasks and whether it adds information beyond general capability.","core_discovery":"The paper's central claim is that standardized commonsense benchmarks have task-dependent and uneven downstream usefulness. Across the suite, model rankings induced by WinoGrande, HellaSwag, Social IQA, and Physical IQA (and their reworked variants) correlate most strongly and consistently with False Beliefs and TRIP; Presupposition shows a smaller but robust positive signal; Implicature, Indirect Requests, and TimeDial show metric-specific gains; CEI and Sarcasm remain unpredicted. Once model size, family, and non-commonsense control performance are accounted for, no extended partial correlation survives correction, and leave-one-family-out prediction improves only for the same narrow subse","pith_inferences":["If this narrow-transfer pattern holds beyond the eight tasks studied, leaderboard-based model selection for real applications should give way to task-specific evaluation; the paper's data suggest a wrong pick is common.","The strong transfer to False Beliefs and TRIP may reflect shared surface formats or shared simple-inference demands rather than a deep commonsense faculty; a testable extension is to rewrite those tasks in open-ended form and see whether transfer collapses.","The WinoWhat case, where the revision changes rankings but not downstream prediction, suggests that artifact filtering and construct validity can move independently; future revision studies should report both rather than internal quality alone."],"forward_implications":["A model that tops a commonsense leaderboard will often be six rank positions away from the top model on a downstream task; top-three overlap averages below 25%.","Cleaning or paraphrasing a benchmark is not enough: reworked benchmarks preserve rankings but do not make scores more predictive of downstream behavior.","Benchmark scores should be reported as evidence about that benchmark, not as a measure of broad commonsense, pragmatic, or social reasoning ability.","Accuracy and macro-F1 can disagree: some tasks, such as Implicature, Indirect Requests, and TimeDial, show gains only on one metric, so predictive-validity claims need to state the metric."],"supporting_citations":[{"why":"Supplies WinoGrande, the coreference benchmark whose rankings are tested against downstream tasks.","marker":"Sakaguchi et al., 2021"},{"why":"Supplies HellaSwag, the event-plausibility benchmark paired with GoldenSwag.","marker":"Zellers et al., 2019"},{"why":"Supplies Social IQA, the social-reasoning benchmark paired with the filtered variant.","marker":"Sap et al., 2019"},{"why":"Supplies Physical IQA, the physical-reasoning benchmark paired with Physical IQA-RUP.","marker":"Bisk et al., 2020"},{"why":"Supplies WinoWhat, the paraphrased WinoGrande revision whose ranking shift does not improve downstream prediction.","marker":"Gevers et al., 2025a"},{"why":"Supplies GoldenSwag, the filtered HellaSwag revision showing no downstream-validity gain.","marker":"Chizhov et al., 2025"},{"why":"Supplies the filtered Social IQA revision, preserving rankings without improving transfer.","marker":"Mousavi et al., 2026"},{"why":"Supplies Physical IQA-RUP, the paraphrased PIQA revision.","marker":"Wang and Zhao, 2024"},{"why":"Supplies the False Beliefs and Indirect Requests downstream tasks, the strongest and one of the metric-specific criteria.","marker":"Jones et al., 2024"},{"why":"Supplies TRIP, the physical-state-tracking task that is one of the two most consistently predicted criteria.","marker":"Storks et al., 2021"}],"fun_headline_variants":["Benchmarks predict LLM commonsense on just 2 of 8 tasks","Reworked commonsense benchmarks still miss most downstream","LLM commonsense benchmarks: narrow transfer, little fix","Commonsense benchmarks fail broad LLM skill prediction","Why benchmarks overstate LLM commonsense: 2 of 8"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The eight downstream tasks, several of which have very few items after manual filtering (CEI n=31, TimeDial n=60), are valid, representative operationalizations of the commonsense competence the benchmarks claim to measure; if those criteria are noisy or unrepresentative, the observed narrow transfer is an artifact of criterion quality rather than a property of the benchmarks.","fun_headline_variants_meta":{"raw":{"variants":["Benchmarks predict LLM commonsense on just 2 of 8 tasks","Reworked commonsense benchmarks still miss most downstream","LLM commonsense benchmarks: narrow transfer, little fix","Commonsense benchmarks fail broad LLM skill prediction","Why benchmarks overstate LLM commonsense: 2 of 8"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000143,"raw_usage":{"total_tokens":982,"prompt_tokens":691,"completion_tokens":291,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":435,"completion_tokens_details":{"reasoning_tokens":205}},"tokens_in":435,"tokens_out":291,"duration_ms":4025,"temperature":1.0,"reasoning_tokens":205,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T20:41:22.946397+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same benchmark-to-downstream prediction with downstream tasks at their original, unfiltered sizes (keeping all CEI and TimeDial items) and with at least 1,000 items per task; if CEI and Sarcasm correlations turn positive and leave-one-family-out R2 rises above zero, the narrow-transfer result is an artifact of manual filtering. Alternatively, if the HellaSwag–TRIP partial correlation survives FDR correction when family and size are matched more tightly, the claim of no incremental validity beyond general capability fails.","supporting_citations":[],"review_version":1}