Pith. sign in

REVIEW 5 major objections 6 minor 27 references

mSCoRe: a $M$ultilingual and Scalable Benchmark for $S$kill-based $Co$mmonsense $Re$asoning

T0 review · 5 major / 6 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read The paper introduces mSCoRe, a multilingual, skill-labelled commonsense benchmark that scales questions to higher complexity and reports that all eight evaluated LLMs—including reasoning-reinforced models—lose accuracy as complexity rises,

desk verdict A genuinely useful benchmark idea whose empirical claims currently hang on unverified LLM-generated data and a missing L4–L6 provenance trail; worth serious review if the authors supply validation and release. read the letter →

arxiv 2508.10137 v1 pith:YYSA6VKM submitted 2025-08-13 cs.CL cs.AI

classification cs.CLcs.AI
keywords commonsensereasoningmultilingualbenchmarksskillsLLMevaluationcomplexityscalingculturallargemodelschain-of-thought
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

mSCoRe is a multilingual benchmark that labels each step of a commonsense answer with one of ten reasoning skills and scales difficulty by adding context, plausible wrong options, and extra reasoning steps. The paper evaluates eight large language models—including reasoning-reinforced models built to think step by step—on 5,600 questions spanning general commonsense in five languages and cultural social commonsense, and finds a consistent accuracy drop as complexity rises: GPT-4o falls from 80.5% to 68.0% on English general questions between level 0 and level 6, while LLaMA-3.1-8B collapses toward chance. The paper's central interpretive claim is that current models reason rigidly: they overuse deductive logic, fail to diversify their skill mix on harder questions, and keep reasoning depth roughly constant. That, the paper argues, makes mSCoRe a reusable instrument for measuring which human reasoning skills models actually employ, not just whether they pick the right answer.

What carries the argument

The load-bearing machinery is the atomic reasoning step: an indivisible unit that predominantly uses one of ten skills and is recorded together with the answer options it eliminates. mSCoRe turns this concept into a generation and evaluation pipeline: every question receives a gold reasoning path of such steps, complexity scaling adds one plausible distractor and one new atomic step per level, and commonsense implicitation rewrites the context and question into a single question that hides the context, forcing the model to supply world knowledge on its own. The same step annotation is used to compare reference reasoning paths with model-generated reasoning paths, which is what allows the pap

What would settle it

Take a random sample of 200 general and 100 social mSCoRe items at levels 1–6, remove the generated reasoning paths, and have native-speaker annotators independently answer the questions and verify that the gold answers are unambiguous and that the added options are genuinely plausible. If more than a small fraction (say 5–10%) of items fail either check, the reported difficulty curve and skill analyses could be artifacts of generation errors rather than measurements of model reasoning.

Watch

Extended reading notes

Core claim

The central discovery is that a benchmark built from human-annotated seeds can make commonsense questions harder in a controlled way by requiring additional atomic reasoning steps—and that doing so exposes a specific weakness in current LLMs. Across the eight evaluated models, accuracy declines as complexity rises; the steepest drop occurs in the first scaling steps, and extending the English general and social subsets to level 6 continues the decline. Reasoning-process analysis shows why: the reference paths diversify into contextual, social, and ethical reasoning skills at higher levels, whereas models like o1 stay heavily deductive and produce roughly constant numbers of steps. The paper

Load-bearing premise

The load-bearing premise is that every LLM-generated context expansion, added distractor, and implicitation rewrite preserves the seed question's correct answer and intended atomic reasoning path; the paper reports no human verification on the 5,600 generated instances.

Editorial extensions

If this is right

  • If the benchmark's difficulty signal is genuine, accuracy at level 0 versus higher levels isolates the cost of added reasoning steps, giving future evaluations a built-in scaling axis that does not require new annotation.
  • The results imply that reasoning-reinforced training as currently practiced can hurt non-English and culturally specific commonsense, so improving those models may require training data and objectives beyond math and code.
  • Because model step counts stay nearly constant while the reference path grows, dynamic allocation of reasoning depth—more steps when the question demands them—should be a testable target for model improvement.
  • Prompting with the ten-skill taxonomy outperforms both plain chain-of-thought and coarser skill categories in the reported models, suggesting that the granularity of reasoning labels itself affects performance.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • An independent human audit of the generated level 1–6 instances would settle whether the declining scores reflect intended reasoning demands or LLM-generation artifacts; until such an audit is public, the skill-distribution conclusions are best treated as provisional.
  • The similar scores across English, German, French, Chinese, and Japanese suggest the taxonomy transfers across these medium-to-high-resource languages; a natural test is to apply the same pipeline to lower-resource languages and non-Western cultural settings, where implicitation may behave differently.
  • The constant step-count finding suggests a cheap intervention: instruct models to budget steps proportionally to the question's complexity level and measure whether accuracy recovers—if it does, step-length adaptation is a causal bottleneck rather than a correlate.
  • Because the generated distractors and expansions come from the same broad LLM family that is later evaluated, the benchmark's future-proofing claim depends on decoupling generation from the target model; evaluating models outside that family would test whether the difficulty curve survives the decoupling.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 6 minor

Summary. The paper proposes mSCoRe, a multilingual benchmark for skill-based commonsense reasoning, with two subsets: mSCoRe-G built from mCSQA (five languages) and mSCoRe-S built from CultureBank (TikTok/Reddit). A four-step LLM-based pipeline filters seed questions, generates a commonsense context and a skill-labeled reasoning process, scales complexity by expanding context, adding plausible distractors and reasoning steps, and then applies 'commonsense implicitation' to render the context implicit. The paper evaluates eight LLMs, reporting accuracy per complexity level, skill-usage distributions, comparisons of reasoning taxonomies, and reasoning-step efficiency. The central claim is that mSCoRe remains challenging for current models, with accuracy declining as complexity increases, and that the skill-based analysis reveals model limitations, especially for social/cultural commonsense.

Significance. If the generation pipeline preserves answer correctness and the skill labels are reliable, mSCoRe would be a useful instrument for fine-grained multilingual commonsense evaluation: it combines a ten-skill taxonomy, a complexity-scaling protocol, and a multilingual general/cultural split. The evaluation covers a broad set of models, including reasoning-reinforced and multilingual systems, and reports both accuracy and reasoning-process analyses. The main weaknesses are that the correctness invariant of the generated instances is not human-verified, the extended complexity levels (L4–L6) are not documented in the dataset statistics, skill labels are not validated, and no statistical uncertainty is reported. These issues are addressable, but they currently leave the benchmark's central claims partly unsupported.

major comments (5)
  1. [§3.2.1 (Steps 3–4), §3.2.2] The correctness invariant of the generation pipeline is unverified. Step 3 instructs the LLM to keep the correct answer semantically similar and to add a plausible distractor; Step 4 asks the model to confirm correctness. These are self-reports by the same generator. No human validation, spot-check, inter-annotator agreement, or error analysis is reported for the 5,600 instances. If any expansion changes the answer or leaks a clue, the L0-to-L6 accuracy drop in Tables 3–5 and the skill distributions in Fig. 6 partly measure generation artifacts. Please add human verification of a representative sample (at least for answer preservation and implicitation) and report agreement/error types.
  2. [§3.2.3 vs §5.1, Table 5] Dataset statistics only describe L0–L3: 200 examples per language and 800 per language for mSCoRe-G, 200 per source for mSCoRe-S, total 5,600. Table 5 and Fig. 6 report L4–L6, but the provenance of those levels is not described. It is unclear how many items per level, which languages, and whether the same pipeline was used. This is central to the scalability claim. Add statistics and generation details for extended levels.
  3. [§5.2, Table 1] Skill labels are not validated. The reference reasoning-process labels are produced by the pipeline's LLM, and model labels are output by the evaluated model. The conclusion that o1 'over-relies on deductive reasoning' and that reference distributions diversify at higher levels is valid only if labels are reliable. No inter-annotator agreement or human validation of the ten-skill taxonomy is reported. Provide label-reliability estimates (e.g., Cohen's kappa on a sample) and, ideally, an analysis of label ambiguity.
  4. [§4.2, Tables 3–5] No confidence intervals or significance tests are reported. Each cell has 200 examples; the standard error is about 3.5 percentage points. Many reported differences are within this range (e.g., GPT-4o English L2=72.5 vs L3=71.5; Table 5 shows non-monotonicity for Aya-32B English L4=58.0, L5=60.0, L6=54.5 and LLaMA-3.3-70B social L3=74.8, L4=75.5). The claim that 'every model accuracy continues to decline to L6' is not supported by Table 5. Report CIs and tests for the main comparisons, or soften the claims to observed trends.
  5. [Appendix A and data availability] The dataset and source code are only promised 'upon acceptance'; no link is provided. For a benchmark, the data artifact is the central contribution. Without access, the accuracy tables and skill-analysis claims cannot be independently checked or reused. Make the benchmark available (e.g., anonymous URL for review) and include dataset documentation, licenses, and the exact generation prompts for all languages.
minor comments (6)
  1. [Abstract] Typo: 'eights state-of-the-art LLMs' should be 'eight state-of-the-art LLMs'.
  2. [§3.2.2, Table 2] The text refers to 'Fig. 2' for the CultureBank example, but the example is Table 2.
  3. [Table 5] Row names are inconsistent: 'Deepseek-70B/8B' in Table 5 versus 'R1-70B/R1-8B' in Tables 3–4 and the text.
  4. [Fig. 6 and §5.2] The caption references panels 'A and C' and 'B and D', but the panels are not described in the text; please identify each panel clearly.
  5. [Appendix A] There is a dangling sentence fragment: 'Source code with specification of all dependencies, including external libraries:' followed by a sentence about release.
  6. [Abstract and §3.2.2] The abstract says 'multilingual general and cultural commonsense,' but mSCoRe-S appears to be built solely from CultureBank and no language selection or translation step is described. If CultureBank is English-only, please clarify that only mSCoRe-G is multilingual or revise the wording.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the benchmark's difficulty claims rest on external evaluations, not on fitted parameters or self-citations.

full rationale

mSCoRe is a benchmark-construction paper rather than a derivation from fitted quantities. The central claims—that scaled complexity lowers accuracy and that models show limited skill diversification—are empirical results obtained by running eight external models on the generated benchmark. No equation defines a predicted value in terms of a fitted parameter, and no self-citation is load-bearing. The only passage that could invite a circularity concern is in Section 4.2, where the paper notes, 'While this can be an artifact of the benchmark creation process where GPT-4o was used for data generation.' That is a transparency statement, not a derivation, and the paper immediately provides an independent counter-check: LLaMA-3.3-70B, which was not part of the generation pipeline, performs comparably to GPT-4o. The skill taxonomy is explicitly injected into both generation and evaluation prompts, so the skill-distribution analysis is conditional on that taxonomy; but this is a stated analytical framing rather than a hidden equivalence. The lack of human verification of generated variants and the underspecified provenance of L4-L6 data are validity/transparency risks, not circularity. Accordingly, no circular step is identified.

Assumptions & free parameters 5 free parameters · 6 assumptions · 2 invented entities

The central claim that mSCoRe is a challenging, scalable benchmark rests on design choices and assumptions about LLM-generated data quality. The most consequential are the unverified preservation of correctness in scaled questions and the unvalidated skill taxonomy, both ad hoc to this paper.

free parameters (5)
  • Number of seed examples per language/source = 200
    Hand-chosen, not justified by power analysis; determines the statistical resolution of all reported accuracies.
  • Number of complexity levels = 0 to 3, extended to 6 in analysis
    Hand-chosen; the scaling behavior is only measured on these levels and may saturate by design.
  • Ten-skill taxonomy composition = 10 skills across 3 categories
    Hand-selected from Wikipedia and Do et al.; not derived from data or validated by experts.
  • Selection thresholds for LLM-judge filtering = not specified
    The rubric has scores 1-5 or 1-3, but the paper never states the cutoff or sampling procedure used to get 200 examples per language.
  • Evaluation prompt requiring skill labels = full taxonomy prompt
    This prompt design affects both the reasoning process outputs and the skill distribution analyses; it is a modeling choice not varied across the main results.
assumptions (6)
  • domain assumption Seed datasets mCSQA and CultureBank contain correct answers and representative commonsense situations.
    Section 3.2 builds all derived questions on these seeds; if a seed label is wrong, all expanded variants inherit the error.
  • ad hoc to paper LLM-generated complexity scaling and implicitation preserve the correct answer and the original reasoning process.
    Steps 3-4 in Section 3.2.1 and Figures 12-13; no human annotation or agreement check is reported for the expanded questions.
  • ad hoc to paper The ten-skill taxonomy is a valid and exhaustive decomposition of commonsense reasoning.
    Section 3.1.1; assembled from Wikipedia and Do et al. with no inter-annotator agreement study or comparison to alternative taxonomies.
  • ad hoc to paper Atomic reasoning steps with a single skill label are a meaningful unit for model analysis.
    Section 3.1 and Figure 2; the definition is normative and there is no evidence that model reasoning decomposes this way.
  • domain assumption Flow Judge provides reliable quality scores for seed filtering.
    Appendix B.1; the judge's outputs are used to select seeds, but no human evaluation of the judge's decisions is reported.
  • standard math Accuracy on a 200-example cell is comparable across models without significance testing.
    Tables 3-5 report point accuracies; the paper interprets differences without confidence intervals, implicitly assuming the differences are not noise.
invented entities (2)
  • Atomic Reasoning Step
    purpose: Defines the unit of analysis for labeling model reasoning processes and for scaling complexity by adding steps.
    Introduced in Section 3.1 and Figure 2; no human annotation study establishes that model reasoning paths decompose into these units.
  • Ten-skill reasoning taxonomy
    purpose: Provides the skill labels used in generation prompts, evaluation prompts, and skill distribution analysis.
    Introduced in Section 3.1.1; the specific set of 10 skills and their definitions are not validated against human reasoning data or alternative taxonomies.

how reviews work

0 comments
Cite this review

Pith. "Pith review of mSCoRe: a $M$ultilingual and Scalable Benchmark for $S$kill-based $Co$mmonsense $Re$asoning." pith.science (2026). https://pith.science/paper/YYSA6VKM

@misc{pith2026250810137,
  author       = {Pith},
  title        = {Pith review of: mSCoRe: a $M$ultilingual and Scalable Benchmark for $S$kill-based $Co$mmonsense $Re$asoning},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/YYSA6VKM}},
  note         = {Machine review of arXiv:2508.10137}
}
read the original abstract

Recent advancements in reasoning-reinforced Large Language Models (LLMs) have shown remarkable capabilities in complex reasoning tasks. However, the mechanism underlying their utilization of different human reasoning skills remains poorly investigated, especially for multilingual commonsense reasoning that involves everyday knowledge across different languages and cultures. To address this gap, we propose a \textbf{M}ultilingual and Scalable Benchmark for \textbf{S}kill-based \textbf{Co}mmonsense \textbf{Re}asoning (\textbf{mSCoRe}). Our benchmark incorporates three key components that are designed to systematically evaluate LLM's reasoning capabilities, including: (1) a novel taxonomy of reasoning skills that enables fine-grained analysis of models' reasoning processes, (2) a robust data synthesis pipeline tailored specifically for commonsense reasoning evaluation, and (3) a complexity scaling framework allowing task difficulty to scale dynamically alongside future improvements in LLM abilities. Extensive experiments on eights state-of-the-art LLMs of varying sizes and training approaches demonstrate that \textbf{mSCoRe} remains significantly challenging for current models, particularly at higher complexity levels. Our results reveal the limitations of such reasoning-reinforced models when confronted with nuanced multilingual general and cultural commonsense. We further provide detailed analysis on the models' reasoning processes, suggesting future directions for improving multilingual commonsense reasoning capabilities.

Figures

Figures reproduced from arXiv: 2508.10137 by the authors.

Figure 1
Figure 1. Data Generation Process. The four-step data creation pipeline of mSCoRe. Each step builds upon the previous one to create progressively more challenging reasoning tasks while maintaining the underlying reasoning skills being evaluated. knowledge, causal relationships, and social interactions. Recent comprehensive benchmarks like MMLU (Hendrycks et al., 2021) and Big-Bench Hard (Suzgun et al., 2023) evaluate the gene… view at source ↗
Figure 2
Figure 2. Atomic Reasoning Step definition. Commonsense reasoning involves making inferences about unstated aspects of a scenario using implicit world knowledge – a capability ingrained in human behavior but still challenging for current LLMs. Unlike formal reasoning domains such as mathematics or logic, where rules are explicitly defined and conclusions follow deter￾minate paths, commonsense reasoning requires access to a va… view at source ↗
Figure 3
Figure 3. Four steps of data generation process. Commonsense-ness: Does answering the question rely solely on commonsense knowledge accessible to the general pop￾ulation, or does it require formal reasoning and specialized expertise beyond everyday understanding? Complexity: How difficult is the question to understand and answer? Does it require minimal reasoning or a complex, multi-step thought process to identify the correc… view at source ↗
Figures from the paper (13 more)
Figure 4
Figure 4. Figure 4: Three criteria of data filtering. Context Expansion: Add additional background or situa￾tional details to the Commonsense Context to increase depth and reasoning requirements to the question. Option Adjustment: Adjust the existing answer options to align with the new c…
Figure 5
Figure 5. Figure 5: Three sub-steps of Complexity Scaling. The overall data generation process is vi￾sualized in [PITH_FULL_IMAGE:figures/full_fig_p005_5.png]
Figure 6
Figure 6. Figure 6: The distribution of reasoning skills of reference reasoning process ( [PITH_FULL_IMAGE:figures/full_fig_p008_6.png]
Figure 7
Figure 7. Figure 7: Average number of reasoning steps in the reasoning [PITH_FULL_IMAGE:figures/full_fig_p009_7.png]
Figure 8
Figure 8. Figure 8: Reasoning skill details. 16 [PITH_FULL_IMAGE:figures/full_fig_p016_8.png]
Figure 9
Figure 9. Figure 9: Rubrics used for mCSQA Data Filtering Process 17 [PITH_FULL_IMAGE:figures/full_fig_p017_9.png]
Figure 10
Figure 10. Figure 10: Rubrics used for CultureBank Data Filtering Process 18 [PITH_FULL_IMAGE:figures/full_fig_p018_10.png]
Figure 11
Figure 11. Figure 11: Prompt for Structured Reasoning Generation step for mSCoRe-G (English). 19 [PITH_FULL_IMAGE:figures/full_fig_p019_11.png]
Figure 12
Figure 12. Figure 12: Prompt for Complexity Expansion step for mSCoRe-G (English). 20 [PITH_FULL_IMAGE:figures/full_fig_p020_12.png]
Figure 13
Figure 13. Figure 13: Prompt for Commonsense Implicitation step for mSCoRe-G (English). 21 [PITH_FULL_IMAGE:figures/full_fig_p021_13.png]
Figure 14
Figure 14. Figure 14: An example from mSCoRe-G for complexity level 0 to 3 (English). 24 [PITH_FULL_IMAGE:figures/full_fig_p024_14.png]
Figure 15
Figure 15. Figure 15: Structured Reasoning Generation Prompt for mSCoRe-S. 26 [PITH_FULL_IMAGE:figures/full_fig_p026_15.png]
Figure 16
Figure 16. Figure 16: An example from mSCoRe-S for complexity level 0 to 3 (English). 31 [PITH_FULL_IMAGE:figures/full_fig_p031_16.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

27 extracted references · 22 canonical work pages

  1. [1]

    Question Comprehension: Read the question carefully along with all the provided answer options

  2. [2]

    COMMONSENSE CONTEXT

    Adding The "COMMONSENSE CONTEXT": Expand on the original question by providing an additional "COMMONSENSE CONTEXT". Ensure that the added context is relevant and enriches the understanding of the question

  3. [3]

    REASONING PROCESS

    Describe your Step-by-Step "REASONING PROCESS" to arrive at the correct answer. Each "ATOMIC REASONING STEP" must following this sequence: 3.1. Choose a REASONING SKILL below to be used by the REASONING STEP: + inductive_reasoning: Drawing general conclusions from specific observations. + deductive_reasoning: Deriving specific conclusions from general pre...

  4. [4]

    commonsense_context

    Generate your output in the JSON format with the following structure: ```json { "commonsense_context": "context_text", "commonsense_question": "question_text", "options": { "A": "option_answer_text_A", ... }, "correct_answer": ["answer_option", "answer_text"], "reasoning_process": { "reasoning_step_1": { "reasoning_skill": "reasoning_skill_name", "reasoni...

  5. [5]

    REASONING PROCESS

    Reasoning Refinements: Refine the original "REASONING PROCESS" to fit the new context. The additional "ATOMIC REASONING STEP" must use one of the following "REASONING SKILLs": + inductive_reasoning: Drawing general conclusions from specific observations. + deductive_reasoning: Deriving specific conclusions from general premises. + abductive_reasoning: For...

  6. [6]

    commonsense_context

    Format the Output using JSON format with the following structure: ```json { "commonsense_context": "context_text", "commonsense_question": "question_text", "options": { "A": "option_answer_text_A", ... }, "correct_answer": ["answer_option", "answer_text"], "reasoning_process": { "reasoning_step_1": { "reasoning_skill": "reasoning_skill_name", "reasoning":...

  7. [10]

    Question Comprehension: Carefully read the given question and the context, and its answer options

  8. [11]

    COMMONSENSE CONTEXT

    Context Expansion: adding additional backgound or situaltional details to the "COMMONSENSE CONTEXT" to add depth and reasoning requirements to the question

Show all 27 references
  1. [12]

    EXPANDED COMMONSENSE CONTEXT

    Question Modificatioin: Utilize the "EXPANDED COMMONSENSE CONTEXT" to craft a more complex question while maintaining its core concept and commonsense

  2. [13]

    Option Adjustments: + Adjust the existing answer options to align with the new complex question + Ensure the correct answer option remains semantically similar to the original + Introduce an additional plausible but incorrect option to increase the complexity of the question +...

  3. [16]

    commonsense_context

    Analyze the provided "commonsense_context" to understand the underlying assumptions and implicit knowledge required for reasoning

  4. [17]

    commonsense_question

    Examine the "commonsense_question" and its associated "options" to identify key elements essential for answering the question

  5. [18]

    commonsense_question

    Rewrite the "commonsense_question" by combining the original context and question to create a more new "commonsense_question" with an "IMPLICITLY IMPLIED COMMONSENSE CONTEXT". Ensure that the new question remains clear and understandable

  6. [19]

    REASONING PROCESS

    Verify that the "REASONING PROCESS" remains unchanged in the transformed question, and confirm that the correct answer remains the same as in the original

  7. [20]

    Ensure that all answer options are reasonable, relevant, and maintain their original intent in the context of the rewritten question

  8. [21]

    reasoning

    Retain the structure and content of the "reasoning" section to reflect the logical steps supporting the correct answer. The "ATOMIC REASONING STEP" must use one of the following "REASONING SKILLs": + inductive_reasoning: Drawing general conclusions from specific observations. ...

  9. [22]

    Analyze the Provided Cultural Situation: Review the details of the cultural group, context, actor behaviors, and other descriptions to understand the key elements of the situation

  10. [23]

    COMMONSENSE CONTEXT

    Adding The "COMMONSENSE CONTEXT": Based on the context given in the input, A "COMMONSENSE CONTEXT" to the question refers to the background knowledge or additional details that are generally understood without requiring specialized knowledge, including factors such as time, pl...

  11. [24]

    Commonsense Question

    Create the "Commonsense Question": Combine the cultural context and the persona's inquiry to formulate a concise question. Ensure the question IMPLICITLY incorporates the original context without explicitly stating it. Create the correct answer option based on the "actor_behavior"

  12. [25]

    Two of which should be plausible options

    Provide Other Answer Options: Create 5 multiple-choice options (including the correct answer from the previous step). Two of which should be plausible options. The other two should be distractors that are relevant and reasonable but incorrect based on the cultural context

  13. [26]

    REASONING PROCESS

    Describe your Step-by-Step "REASONING PROCESS" to arrive at the correct answer. Each "ATOMIC REASONING STEP" must following this sequence: 5.1. Choose a "REASONING SKILL" below to be used by the "REASONING STEP": + inductive_reasoning: Drawing general conclusions from specific...

  14. [27]

    commonsense_context

    Generate your output in the JSON format with the following structure: ```json { "commonsense_context": "context_text", "commonsense_question": "question_text", "options": { "A": "option_answer_text_A", ... }, "correct_answer": ["answer_option", "answer_text"], "reasoning_proce...

  15. [102]

    Niklas Muennighoff, Zitong Yang, Weijia Shi, Xiang Lisa Li, Li Fei-Fei, Hannaneh Hajishirzi, Luke Zettlemoyer, Percy Liang, Emmanuel J

    URLhttps://aclanthology.org/2021.acl-long.102/. Niklas Muennighoff, Zitong Yang, Weijia Shi, Xiang Lisa Li, Li Fei-Fei, Hannaneh Hajishirzi, Luke Zettlemoyer, Percy Liang, Emmanuel J. Candès, and Tatsunori Hashimoto. s1: Simple test-time scaling.CoRR, abs/2501.19393, 2025. doi...

  16. [604]

    Mirac Suzgun, Nathan Scales, Nathanael Schärli, Sebastian Gehrmann, Yi Tay, Hyung Won Chung, Aakanksha Chowdhery, Quoc V

    URLhttps://aclanthology.org/2024.acl-long.604/. Mirac Suzgun, Nathan Scales, Nathanael Schärli, Sebastian Gehrmann, Yi Tay, Hyung Won Chung, Aakanksha Chowdhery, Quoc V . Le, Ed H. Chi, Denny Zhou, and Jason Wei. Challenging big-bench tasks and whether chain-of-thought can sol...

  17. [824]

    inductive_reasoning

    URLhttps://doi.org/10.18653/v1/2023.findings-acl.824. Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant. Commonsenseqa: A question answering challenge targeting commonsense knowledge. In Jill Burstein, Christy Doran, and Thamar Solorio (eds.),Proceedings of th...

  18. [2021]

    URLhttps://arxiv.org/abs/2110.14168. John Dang, Shivalika Singh, Daniel D’souza, Arash Ahmadian, Alejandro Salamanca, Madeline Smith, Aidan Peppin, Sungjin Hong, Manoj Govindassamy, Terrence Zhao, Sandra Kublik, Meor Amer, Viraat Aryabumi, Jon Ander Campos, Yi-Chern Tan, and e...

  19. [2024]

    URL https://doi.org/10.48550/arXiv.2408

    doi: 10.48550/ARXIV .2408.03314. URL https://doi.org/10.48550/arXiv.2408. 03314. Jiaxing Sun, Weiquan Huang, Jiang Wu, Chenya Gu, Wei Li, Songyang Zhang, Hang Yan, and Conghui He. Benchmarking Chinese commonsense reasoning of LLMs: From Chinese-specifics to reasoning-memorizat...

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.