Pith. sign in

Evaluating Large Language Models with fmeval

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

fmeval is an open source library to evaluate large language models (LLMs) in a range of tasks. It helps practitioners evaluate their model for task performance and along multiple responsible AI dimensions. This paper presents the library and exposes its underlying design principles: simplicity, coverage, extensibility and performance. We then present how these were implemented in the scientific and engineering choices taken when developing fmeval. A case study demonstrates a typical use case for the library: picking a suitable model for a question answering task. We close by discussing limitations and further work in the development of the library. fmeval can be found at https://github.com/aws/fmeval.

fields

cs.CL 1

years

2025 1

verdicts

CONDITIONAL 1

representative citing papers

The Science of Evaluating Foundation Models

cs.CL · 2025-02-12 · conditional · novelty 3.0

A survey-and-checklist proposal that organizes LLM evaluation into an ABCD framework (Algorithm, Big Data, Computation, Domain Expertise) for context-aware, documented assessment.

citing papers explorer

Showing 1 of 1 citing paper.

  • The Science of Evaluating Foundation Models cs.CL · 2025-02-12 · conditional · none · ref 77 · internal anchor

    A survey-and-checklist proposal that organizes LLM evaluation into an ABCD framework (Algorithm, Big Data, Computation, Domain Expertise) for context-aware, documented assessment.