Pith. sign in

REVIEW 1 cited by

AgriBench: A Hierarchical Agriculture Benchmark for Multimodal Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.00465 v2 pith:NDIQBAJU submitted 2024-11-30 cs.CV cs.AI

classification cs.CVcs.AI
keywords agriculturelandcoverdatasetmm-llmsmultimodalagribenchbenchmark
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce AgriBench, the first agriculture benchmark designed to evaluate MultiModal Large Language Models (MM-LLMs) for agriculture applications. To further address the agriculture knowledge-based dataset limitation problem, we propose MM-LUCAS, a multimodal agriculture dataset, that includes 1,784 landscape images, segmentation masks, depth maps, and detailed annotations (geographical location, country, date, land cover and land use taxonomic details, quality scores, aesthetic scores, etc), based on the Land Use/Cover Area Frame Survey (LUCAS) dataset, which contains comparable statistics on land use and land cover for the European Union (EU) territory. This work presents a groundbreaking perspective in advancing agriculture MM-LLMs and is still in progress, offering valuable insights for future developments and innovations in specific expert knowledge-based MM-LLMs.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Towards Large Reasoning Models for Agriculture

    cs.LG 2025-05 conditional novelty 6.0 of 10

    A new 100-question agricultural reasoning benchmark and a 44.6K-question training dataset show current AI models score at most 36%, and fine-tuning small models on the dataset lifts them from 0-1% to 3-5%.

Pith tools