Pith. sign in

REVIEW 4 cited by

ANLS* -- A Universal Document Processing Metric for Generative Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.03848 v10 pith:JY6RDLY5 submitted 2024-02-06 cs.CL cs.AI

classification cs.CLcs.AI
keywords anlsmetricmodelsgllmsdifferentevaluationgenerativetasks
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Traditionally, discriminative models have been the predominant choice for tasks like document classification and information extraction. These models make predictions that fall into a limited number of predefined classes, facilitating a binary true or false evaluation and enabling the direct calculation of metrics such as the F1 score. However, recent advancements in generative large language models (GLLMs) have prompted a shift in the field due to their enhanced zero-shot capabilities, which eliminate the need for a downstream dataset and computationally expensive fine-tuning. However, evaluating GLLMs presents a challenge as the binary true or false evaluation used for discriminative models is not applicable to the predictions made by GLLMs. This paper introduces a new metric for generative models called ANLS* for evaluating a wide variety of tasks, including information extraction and classification tasks. The ANLS* metric extends existing ANLS metrics as a drop-in-replacement and is still compatible with previously reported ANLS scores. An evaluation of 7 different datasets, and more than 20 different GLLMs together with 3 different prompting methods using the ANLS* metric is also provided, demonstrating the importance of the proposed metric. We also benchmark a novel approach to generate prompts for documents, called SFT, against other prompting techniques such as LATIN. In almost all cases, SFT outperforms other techniques and improves the state-of-the-art, sometimes by as much as $10$ percentage points. Sources are available at https://github.com/deepopinion/anls_star_metric

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. GDP.pdf: Benchmarking Grounded Multimodal Reasoning over Professional PDF Documents

    cs.CV 2026-07 conditional novelty 7.0 of 10

    On GDP.pdf, 100 expert-authored professional PDF tasks, seventeen frontier multimodal models pass at most 30.7% of items, and most failures come from missed footnotes, exclusions, tables, and spatial details.

  2. What Does Your Short-Answer VQA Score Actually Measure? Evaluator-Dependent Instability in Multimodal Short-Answer Benchmarks

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Official short-answer VQA scores undercount semantic success by several points on text-rich benchmarks because automatic evaluators reject acceptable surface-form variants, with sensitivity structured by answer-contract type.

  3. MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models

    cs.AI 2026-06 conditional novelty 5.0 of 10

    By sampling variance-inflated query vectors during prefilling, MM-ShiftKV selects prompt KV caches that better match decoding-time attention and outperforms prior prefill-only KV compression on multimodal benchmarks a...

  4. CAD2DMD-SET: Synthetic Generation Tool of Digital Measurement Device CAD Model Datasets for fine-tuning Large Vision-Language Models

    cs.CV 2025-08 conditional novelty 5.0 of 10

    A synthetic data pipeline for digital measurement devices plus a real-image benchmark improves LVLM reading performance from 32.92% to 96.04% ANLS for InternVL.

Pith tools