Pith. sign in

REVIEW 2 cited by

AgroGPT: Efficient Agricultural Vision-Language Model with Expert Tuning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.08405 v2 pith:5TMQRGLQ submitted 2024-10-10 cs.CV cs.AI

classification cs.CVcs.AI
keywords dataagrogptmodelsagricultureagriculturaldomainslargeavailable
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Significant progress has been made in advancing large multimodal conversational models (LMMs), capitalizing on vast repositories of image-text data available online. Despite this progress, these models often encounter substantial domain gaps, hindering their ability to engage in complex conversations across new domains. Recent efforts have aimed to mitigate this issue, albeit relying on domain-specific image-text data to curate instruction-tuning data. However, many domains, such as agriculture, lack such vision-language data. In this work, we propose an approach to construct instruction-tuning data that harnesses vision-only data for the agriculture domain. We utilize diverse agricultural datasets spanning multiple domains, curate class-specific information, and employ large language models (LLMs) to construct an expert-tuning set, resulting in a 70k expert-tuning dataset called AgroInstruct. Subsequently, we expert-tuned and created AgroGPT, an efficient LMM that can hold complex agriculture-related conversations and provide useful insights. We also develop AgroEvals for evaluation and compare {AgroGPT's} performance with large open and closed-source models. {AgroGPT} excels at identifying fine-grained agricultural concepts, can act as an agriculture expert, and provides helpful information for multimodal agriculture questions. The code, datasets, and models are available at https://github.com/awaisrauf/agroGPT.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. AgroBench: Vision-Language Model Benchmark in Agriculture

    cs.CV 2025-07 conditional novelty 6.0 of 10

    AgroBench is a new expert-annotated benchmark showing that current vision-language models, especially open-source ones, struggle with fine-grained agricultural identification such as weed species.

  2. Self-Consistency in Vision-Language Models for Precision Agriculture: Multi-Response Consensus for Crop Disease Management

    cs.CV 2025-07 conditional novelty 5.0 of 10

    A cosine-consistency voting scheme with a domain-adapted embedding raises VLM accuracy on maize disease diagnosis by 5.6 to 15.5 percentage points, but the gain is measured with the same LLM scorer used to create the ...

Pith tools