Pith. sign in

REVIEW 7 cited by

MEGA: Multilingual Evaluation of Generative AI

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2303.12528 v4 pith:54NMQOJW submitted 2023-03-22 cs.CL

classification cs.CL
keywords generativemodelsllmslanguagesperformancelanguagetasksacross
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Generative AI models have shown impressive performance on many Natural Language Processing tasks such as language understanding, reasoning, and language generation. An important question being asked by the AI community today is about the capabilities and limits of these models, and it is clear that evaluating generative AI is very challenging. Most studies on generative LLMs have been restricted to English and it is unclear how capable these models are at understanding and generating text in other languages. We present the first comprehensive benchmarking of generative LLMs - MEGA, which evaluates models on standard NLP benchmarks, covering 16 NLP datasets across 70 typologically diverse languages. We compare the performance of generative LLMs including Chat-GPT and GPT-4 to State of the Art (SOTA) non-autoregressive models on these tasks to determine how well generative models perform compared to the previous generation of LLMs. We present a thorough analysis of the performance of models across languages and tasks and discuss challenges in improving the performance of generative LLMs on low-resource languages. We create a framework for evaluating generative LLMs in the multilingual setting and provide directions for future progress in the field.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ChiKhaPo: A Large-Scale Multilingual Benchmark for Evaluating Lexical Comprehension and Generation in Large Language Models

    cs.CL 2025-10 conditional novelty 7.0 of 10

    ChiKhaPo is an 8-subtask benchmark that measures word-level comprehension and generation in 2,700+ languages and shows state-of-the-art models perform poorly on low-resource languages.

  2. Preperiodic points, finiteness, and structures of semigroups of algebraic morphisms

    math.NT 2025-08 unverdicted novelty 6.0 of 10

    The paper proves finiteness and structural results for preperiodic points of algebraic morphisms, including Burnside-type and Northcott-type theorems.

  3. Multilingual Question Answering in Low-Resource Settings: A Dzongkha-English Benchmark for Foundation Models

    cs.CL 2025-05 conditional novelty 6.0 of 10

    DZEN, a parallel Dzongkha-English benchmark of 5,161 school science exam questions, shows large LLM accuracy gaps in Dzongkha; adding English translations narrows the gap.

  4. How and Where to Translate? The Impact of Translation Strategies in Cross-lingual LLM Prompting

    cs.CL 2025-07 conditional novelty 5.0 of 10

    For multilingual RAG intent classification, the best translation strategy depends on the model and language; translating instructions into the user's language helps some models, while making the model answer in low-re...

  5. Beyond English: The Impact of Prompt Translation Strategies across Languages and Tasks in Multilingual LLMs

    cs.CL 2025-02 conditional novelty 5.0 of 10

    Selective pre-translation, translating only some prompt components into English, generally outperforms both full prompt translation and direct inference across tasks and languages, with the largest gains for low-resou...

  6. Leveraging Large Language Models for Bengali Math Word Problem Solving with Chain of Thought Reasoning

    cs.CL 2025-05 conditional novelty 4.0 of 10

    A new Bengali math word problem dataset translated from GSM8K is benchmarked with chain-of-thought prompting, yielding 88% accuracy with LLaMA-3.3 70B on a 1,000-sample test subset.

  7. Prompt, Translate, Fine-Tune, Re-Initialize, or Instruction-Tune? Adapting LLMs for In-Context Learning in Low-Resource Languages

    cs.CL 2025-06

Pith tools