Pith. sign in

REVIEW 1 cited by

CJEval: A Benchmark for Assessing Large Language Models Using Chinese Junior High School Exam Data

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.16202 v2 pith:6UZAJS3I submitted 2024-09-24 cs.AI

classification cs.AI
keywords educationalbenchmarkcjevalllmsapplicationschineseeducationexam
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Online education platforms have significantly transformed the dissemination of educational resources by providing a dynamic and digital infrastructure. With the further enhancement of this transformation, the advent of Large Language Models (LLMs) has elevated the intelligence levels of these platforms. However, current academic benchmarks provide limited guidance for real-world industry scenarios. This limitation arises because educational applications require more than mere test question responses. To bridge this gap, we introduce CJEval, a benchmark based on Chinese Junior High School Exam Evaluations. CJEval consists of 26,136 samples across four application-level educational tasks covering ten subjects. These samples include not only questions and answers but also detailed annotations such as question types, difficulty levels, knowledge concepts, and answer explanations. By utilizing this benchmark, we assessed LLMs' potential applications and conducted a comprehensive analysis of their performance by fine-tuning on various educational tasks. Extensive experiments and discussions have highlighted the opportunities and challenges of applying LLMs in the field of education.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Pedagogy-R1: Pedagogically-Aligned Reasoning Model with Balanced Educational Benchmark

    cs.AI 2025-05 reject novelty 4.0 of 10

    Pedagogy-R1 distills pedagogical reasoning into small models from QwQ-32B and evaluates them with a new five-domain educational benchmark, but the gains are modest and the benchmark is only partially public.

Pith tools