Pith. sign in

REVIEW 1 cited by

FoundaBench: Evaluating Chinese Fundamental Knowledge Capabilities of Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.18359 v1 pith:DJQEGFIX submitted 2024-04-29 cs.CL cs.AI

classification cs.CLcs.AI
keywords foundabenchknowledgemodelschinesefundamentalllmscapabilitieslanguage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In the burgeoning field of large language models (LLMs), the assessment of fundamental knowledge remains a critical challenge, particularly for models tailored to Chinese language and culture. This paper introduces FoundaBench, a pioneering benchmark designed to rigorously evaluate the fundamental knowledge capabilities of Chinese LLMs. FoundaBench encompasses a diverse array of 3354 multiple-choice questions across common sense and K-12 educational subjects, meticulously curated to reflect the breadth and depth of everyday and academic knowledge. We present an extensive evaluation of 12 state-of-the-art LLMs using FoundaBench, employing both traditional assessment methods and our CircularEval protocol to mitigate potential biases in model responses. Our results highlight the superior performance of models pre-trained on Chinese corpora, and reveal a significant disparity between models' reasoning and memory recall capabilities. The insights gleaned from FoundaBench evaluations set a new standard for understanding the fundamental knowledge of LLMs, providing a robust framework for future advancements in the field.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MCEval: A Dynamic Framework for Fair Multilingual Cultural Evaluation of LLMs

    cs.CL 2025-07 conditional novelty 5.0 of 10

    A dynamic multilingual cultural evaluation framework shows that LLM cultural performance depends on both training data distribution and language-culture alignment, and that English-only evaluations hide severe cultura...

Pith tools