Pith. sign in

REVIEW 1 cited by

Disce aut Deficere: Evaluating LLMs Proficiency on the INVALSI Italian Benchmark

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.17535 v1 pith:PZ36JD5L submitted 2024-06-25 cs.CL cs.AI

classification cs.CLcs.AI
keywords benchmarkllmsinvalsimodelsacrossautomatedcrucialcurrent
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Recent advancements in Large Language Models (LLMs) have significantly enhanced their ability to generate and manipulate human language, highlighting their potential across various applications. Evaluating LLMs in languages other than English is crucial for ensuring their linguistic versatility, cultural relevance, and applicability in diverse global contexts, thus broadening their usability and effectiveness. We tackle this challenge by introducing a structured benchmark using the INVALSI tests, a set of well-established assessments designed to measure educational competencies across Italy. Our study makes three primary contributions: Firstly, we adapt the INVALSI benchmark for automated LLM evaluation, which involves rigorous adaptation of the test format to suit automated processing while retaining the essence of the original tests. Secondly, we provide a detailed assessment of current LLMs, offering a crucial reference point for the academic community. Finally, we visually compare the performance of these models against human results. Additionally, researchers are invited to submit their models for ongoing evaluation, ensuring the benchmark remains a current and valuable resource.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. An Empirical Investigation of Gender Stereotype Representation in Large Language Models: The Italian Case

    cs.CL 2025-07 conditional novelty 4.0 of 10

    In Italian, both ChatGPT (gpt-4o-mini) and Gemini (gemini-1.5-flash) associate higher-status professional roles with male pronouns and subordinate roles with female pronouns, e.g., 97-100% of 'she' responses pointed t...

Pith tools