Pith. sign in

REVIEW 2 cited by

LMentry: A Language Model Benchmark of Elementary Language Tasks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2211.02069 v2 pith:F7USRYE4 submitted 2022-11-03 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords languagelargelmentrymodelsbenchmarktaskscomplexhumans
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

As the performance of large language models rapidly improves, benchmarks are getting larger and more complex as well. We present LMentry, a benchmark that avoids this "arms race" by focusing on a compact set of tasks that are trivial to humans, e.g. writing a sentence containing a specific word, identifying which words in a list belong to a specific category, or choosing which of two words is longer. LMentry is specifically designed to provide quick and interpretable insights into the capabilities and robustness of large language models. Our experiments reveal a wide variety of failure cases that, while immediately obvious to humans, pose a considerable challenge for large language models, including OpenAI's latest 175B-parameter instruction-tuned model, TextDavinci002. LMentry complements contemporary evaluation approaches of large language models, providing a quick, automatic, and easy-to-run "unit test", without resorting to large benchmark suites of complex tasks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. TASE: Token Awareness and Structured Evaluation for Multilingual Language Models

    cs.CL 2025-08 unverdicted novelty 5.0 of 10

    TASE benchmark shows LLMs lag humans on token-level and structural language tasks across Chinese, English, and Korean despite strong high-level performance.

  2. Enhancing LLM Character-Level Manipulation via Divide and Conquer

    cs.CL 2025-02 conditional novelty 4.0 of 10

    ToCAD, a three-stage divide-and-conquer prompt, atomizes words into spaced letters, edits them, and reconstructs, sharply improving LLM exact-match accuracy on deletion, insertion, and substitution tasks.

Pith tools