REVIEW 1 major objections 33 references
Multimodal LLMs generate up to 91 percent biologically inconsistent agricultural scenes from text prompts.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
Multimodal LLMs show 25-37% error in zero-shot agricultural image interpretation and up to 91% biologically inconsistent outputs in text-to-image generation tasks.
T0 review reviewed 2026-06-29 challenge →
load-bearing objection The 91% inconsistency claim for generated ag scenes rests on unvalidated subjective criteria with no reported sample sizes or rater agreement. the 1 major comments →
Hallucination Behavior in Multimodal LLMs Across Agricultural Image Interpretation and Generation Tasks
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
Core claim
The paper establishes that multimodal LLMs exhibit recurring hallucinations in agricultural image tasks, with text-to-image generation showing up to 91 percent biologically inconsistent outputs in advanced models under relaxed constraints, while image-to-text interpretation shows moderate accuracy that improves with prompting but retains false detections and missed infections.
What carries the argument
Domain-informed criteria that score outputs for biological inconsistency, contextual inaccuracy, and agronomic implausibility.
Load-bearing premise
The criteria for judging biological inconsistency, contextual inaccuracy, and agronomic implausibility are accurate and free of evaluator bias when applied to model outputs.
What would settle it
Independent agronomists re-scoring the same set of generated scenes with the paper's criteria and finding a substantially lower inconsistency rate.
If this is right
- Few-shot prompting reduces but does not eliminate hallucinations in image interpretation.
- Current models show fundamental weaknesses in producing biologically plausible agricultural scenes from text.
- Recurring error patterns appear across both interpretive and generative agricultural tasks.
- Reliability improvements are needed before LLM-based agricultural imaging platforms can be trusted.
Where Pith is reading between the lines
- Similar inconsistency rates may appear in non-agricultural image generation tasks that require domain knowledge.
- Prompt engineering alone may not close the gap without added domain-specific constraints or verification steps.
- Deployment in real farm decision systems would require human review loops to catch the observed error types.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper examines hallucination behaviors in multimodal LLMs for agricultural image interpretation (image-to-text) and generation (text-to-image) tasks. It reports modest zero-shot accuracy of 63-75% for models like Gemma, LLAVA, Qwen, and MiniCPM in interpreting crop stresses, improving to 86.8% with few-shot prompting, and up to 91% biologically inconsistent scenes generated by GPT-5 and Gemini 2.5 Flash under relaxed prompts, using domain-informed criteria for biological inconsistency, contextual inaccuracy, and agronomic implausibility.
Significance. If the reported hallucination rates and patterns could be reproduced with validated, reproducible evaluation protocols, the work would document concrete limitations of current multimodal LLMs in a high-stakes applied domain and could inform targeted mitigation strategies for agricultural imaging pipelines.
major comments (1)
- [Abstract] Abstract: The headline quantitative results (63–75% zero-shot accuracy, 86.8% few-shot accuracy, and up to 91% biologically inconsistent scenes) are stated without any accompanying sample sizes, image or prompt selection criteria, statistical tests, or inter-rater reliability statistics for the domain-informed criteria. Because the distinction between consistent and inconsistent agricultural scenes is not self-evident, these percentages cannot be assessed and constitute the central empirical claim.
Simulated Author's Rebuttal
We thank the referee for their detailed review and constructive feedback on the abstract. We agree that the headline results require additional context for proper assessment and will revise accordingly to strengthen the manuscript's clarity and reproducibility.
read point-by-point responses
-
Referee: [Abstract] Abstract: The headline quantitative results (63–75% zero-shot accuracy, 86.8% few-shot accuracy, and up to 91% biologically inconsistent scenes) are stated without any accompanying sample sizes, image or prompt selection criteria, statistical tests, or inter-rater reliability statistics for the domain-informed criteria. Because the distinction between consistent and inconsistent agricultural scenes is not self-evident, these percentages cannot be assessed and constitute the central empirical claim.
Authors: We agree that the abstract as currently written does not provide sufficient supporting details for the reported percentages. In the revised manuscript we will expand the abstract to include: (1) sample sizes for both the image-to-text (N images evaluated across models) and text-to-image (N prompts per model) experiments; (2) a concise description of image and prompt selection criteria; (3) mention of any statistical tests performed; and (4) inter-rater reliability metrics (e.g., Cohen’s kappa) for the domain-informed criteria of biological inconsistency, contextual inaccuracy, and agronomic implausibility. The full Methods section already defines these criteria with examples; we will ensure the abstract is self-contained while remaining within length limits. These changes directly address the concern that the central claims cannot be assessed without this information. revision: yes
Circularity Check
No significant circularity; purely observational empirical reporting
full rationale
The paper reports measured accuracies (63-75% zero-shot, up to 86.8% few-shot) and inconsistency rates (up to 91%) from direct evaluation of LLM outputs on agricultural images and prompts. No equations, parameter fitting, derivations, or self-citations appear in the provided text. The central claims rest on external data collection and application of domain-informed criteria rather than any reduction to the paper's own inputs by construction. This matches the default expectation of non-circularity for observational studies.
Axiom & Free-Parameter Ledger
axioms (1)
- domain assumption Domain-informed criteria for biological inconsistency and agronomic implausibility are reliable and unbiased when applied to LLM outputs.
Cite this review
Pith. "Pith review of Hallucination Behavior in Multimodal LLMs Across Agricultural Image Interpretation and Generation Tasks." pith.science (2026). https://pith.science/paper/OHDWPW2C
@misc{pith2026260527595,
author = {Pith},
title = {Pith review of: Hallucination Behavior in Multimodal LLMs Across Agricultural Image Interpretation and Generation Tasks},
year = {2026},
howpublished = {\url{https://pith.science/paper/OHDWPW2C}},
note = {Machine review of arXiv:2605.27595}
}
read the original abstract
Large Language Models (LLMs) are being rapidly adopted in agricultural imaging applications, ranging from crop interpretation to synthetic field image generation. However, these models frequently exhibit hallucinations outputs that appear confident yet deviate from biological or environmental reality potentially leading to misinformed agronomic insights. This study investigates such hallucinations in two complementary directions: image-to-text, where LLMs interpret crop or field imagery to describe conditions such as biotic and abiotic stresses, and text-to-image, where models generate synthetic agricultural scenes based on descriptive prompts. We examine errors involving biological inconsistency, contextual inaccuracy, and agronomic implausibility, evaluating the outputs under domain-informed criteria across multiple imaging modalities. Our analysis identifies recurring hallucination patterns within both interpretive and generative tasks. In image interpretation, LLMs (e.g., Gemma, LLAVA, Qwen, and MiniCPM) achieved modest zero-shot accuracy (63 to 75 percent), whereas few-shot prompting improved performance up to 86.8 percent, exhibiting false detections and missed infections, indicating residual hallucination effects. In text-to-image tasks, advanced models such as GPT-5 and Gemini 2.5 Flash generate up to 91 percent biologically inconsistent scenes under relaxed prompt constraints, revealing fundamental weaknesses in current LLMs. This systematic assessment of visual reasoning and generation offers critical insights toward enhancing the reliability and trustworthiness of LLM-based agricultural imaging platforms.
Figures
Reference graph
Works this paper leans on
-
[1]
Chang, X
Y. Chang, X. Wang, J. Wang, Y. Wu, L. Yang, K. Zhu, H. Chen, X. Yi, C. Wang, Y. Wang, et al., A survey on evaluation of large language models, ACM transactions on intelligent systems and technology 15 (3) (2024) 1–45
2024
-
[2]
Bhayana, Chatbots and large language models in radiology: a practical primer for clinical and research applications, Radiology 310 (1) (2024) e232756
R. Bhayana, Chatbots and large language models in radiology: a practical primer for clinical and research applications, Radiology 310 (1) (2024) e232756. 11
2024
-
[3]
Tzachor, M
A. Tzachor, M. Devare, C. Richards, P. Pypers, A. Ghosh, J. Koo, S. Johal, B. King, Large language models and agricultural extension services, Nature food 4 (11) (2023) 941–948
2023
-
[4]
L. Fang, W. Xiang, J. Jin, K. Liao, C. Liu, Y. Han, F. D. Salim, Y.-P. P. Chen, Agri-llm: Prompt-based large language model for emission data analytics in smart agriculture, IEEE Internet of Things Journal (2025)
2025
- [5]
-
[6]
B., et al., Tomato leaf disease detection dataset,https://www.kaggle.com/datasets/ kaustubhb999/tomatoleaf/data, accessed: 2025-10-23 (2025)
K. B., et al., Tomato leaf disease detection dataset,https://www.kaggle.com/datasets/ kaustubhb999/tomatoleaf/data, accessed: 2025-10-23 (2025)
2025
-
[7]
Y. Zhou, C. Cui, J. Yoon, L. Zhang, Z. Deng, C. Finn, M. Bansal, H. Yao, Analyzing and mitigating object hallucination in large vision-language models, in: NeurIPS 2023 Workshop on Instruction Tuning and Instruction Following
2023
-
[8]
Huang, W
L. Huang, W. Yu, W. Ma, W. Zhong, Z. Feng, H. Wang, Q. Chen, W. Peng, X. Feng, B. Qin, et al., A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions, ACM Transactions on Information Systems 43 (2) (2025) 1–55
2025
-
[9]
Z. Bai, P. Wang, T. Xiao, T. He, Z. Han, Z. Zhang, M. Z. Shou, Hallucination of multimodal large language models: A survey, arXiv preprint arXiv:2404.18930 (2024)
work page internal anchor Pith review Pith/arXiv arXiv 2024
-
[10]
H.-K. Ko, G. Park, H. Jeon, J. Jo, J. Kim, J. Seo, Large-scale text-to-image generation models for visual artists’ creative works, in: Proceedings of the 28th international conference on intelligent user interfaces, 2023, pp. 919–933
2023
-
[11]
Q. Yan, X. He, X. E. Wang, Med-hvl: Automatic medical domain hallucination evaluation for large vision-language models, in: AAAI 2024 Spring Symposium on Clinical Foundation Models, 2024
2024
- [12]
-
[13]
Jayakumar, A
S. Jayakumar, A. K. Abhangrao, R. A. Sarje, R. Gupta, S. Pathania, B. V. Sree, Critical analysis on effect of micronutrients on flowering plants: a review, Int J Plant Soil Sci 36 (2024) 776–782
2024
-
[14]
Taghvaeian, J
S. Taghvaeian, J. L. Ch´ avez, J. Altenhofen, T. Trout, K. DeJonge, et al., Remote sensing for evaluating crop water stress at field scale using infrared thermography: potential and limitations (2013)
2013
-
[15]
G. Team, R. Anil, S. Borgeaud, J.-B. Alayrac, J. Yu, R. Soricut, J. Schalkwyk, A. M. Dai, A. Hauth, K. Millican, et al., Gemini: a family of highly capable multimodal models, arXiv preprint arXiv:2312.11805 (2023)
work page internal anchor Pith review Pith/arXiv arXiv 2023
-
[16]
T. Wei, Z. Chen, X. Yu, Snap and diagnose: An advanced multimodal retrieval system for identifying plant diseases in the wild, in: Proceedings of the 6th ACM International Conference on Multimedia in Asia, 2024, pp. 1–3
2024
-
[17]
J. Liu, X. Wang, A multimodal framework for pepper diseases and pests detection, Scientific Reports 14 (1) (2024) 28973
2024
- [18]
-
[19]
Zheng, J
J. Zheng, J. Wang, Agrigpt: A strong plant disease detection model via visual-language model, in: International Conference on Intelligent Computing, Springer, 2025, pp. 239–250. 12
2025
-
[20]
Y. Wang, F. Wang, W. Chen, B. Lv, M. Liu, X. Kong, C. Zhao, Z. Pan, A large language model for multimodal identification of crop diseases and pests, Scientific Reports 15 (1) (2025) 21959
2025
-
[21]
J. Qing, X. Deng, Y. Lan, Z. Li, Gpt-aided diagnosis on agricultural image based on a new light yolopc, Computers and electronics in agriculture 213 (2023) 108168
2023
-
[22]
R. Yan, P. An, X. Meng, Y. Li, D. Li, F. Xu, D. Dang, A knowledge graph for crop diseases and pests in china, Scientific Data 12 (1) (2025) 222
2025
-
[23]
Huang, W
N. Huang, W. Dong, Y. Zhang, F. Tang, R. Li, C. Ma, X. Li, C. Xu, Creativesynth: Creative blending and synthesis of visual arts based on multimodal diffusion, CoRR (2024)
2024
-
[24]
L. C. Adams, F. Busch, D. Truhn, M. R. Makowski, H. J. Aerts, K. K. Bressem, What does dall-e 2 know about radiology?, Journal of medical Internet research 25 (2023) e43110
2023
-
[25]
Seneviratne, D
S. Seneviratne, D. Senanayake, S. Rasnayaka, R. Vidanaarachchi, J. Thompson, Dalle-urban: Capturing the urban design expertise of large text to image transformers, in: 2022 International Conference on Digital Image Computing: Techniques and Applications (DICTA), IEEE, 2022, pp. 1–9
2022
-
[26]
Sapkota, Z
R. Sapkota, Z. Meng, M. Karkee, Synthetic meets authentic: Leveraging llm generated datasets for yolo11 and yolov10-based apple detection through machine vision sensors, Smart Agricultural Technol- ogy 9 (2024) 100614
2024
-
[27]
Vayadande, S
K. Vayadande, S. Bhemde, V. Rajguru, P. Ugile, R. Lade, N. Raut, Ai-based image generator web application using openai’s dall-e system, in: 2023 international conference on recent advances in science and engineering technology (ICRASET), IEEE, 2023, pp. 1–5
2023
-
[28]
Y. Kim, H. Jeong, S. Chen, S. S. Li, M. Lu, K. Alhamoud, J. Mun, C. Grau, M. Jung, R. Gameiro, et al., Medical hallucinations in foundation models and their impact on healthcare, CoRR (2025)
2025
-
[29]
A. Pal, L. K. Umapathi, M. Sankarasubbu, Med-halt: Medical domain hallucination test for large language models, in: Proceedings of the 27th Conference on Computational Natural Language Learning (CoNLL), 2023, pp. 314–334
2023
-
[30]
W. Shi, X. Han, M. Lewis, Y. Tsvetkov, L. Zettlemoyer, W.-t. Yih, Trusting your evidence: Hallucinate less with context-aware decoding, in: Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 2: Short Papers), 2024, pp. 783–791
2024
-
[31]
Chuang, Y
Y.-S. Chuang, Y. Xie, H. Luo, Y. Kim, J. R. Glass, P. He, Dola: Decoding by contrasting layers improves factuality in large language models, in: The Twelfth International Conference on Learning Representations, 2023
2023
-
[32]
K. Li, O. Patel, F. Vi´ egas, H. Pfister, M. Wattenberg, Inference-time intervention: Eliciting truthful answers from a language model, Advances in Neural Information Processing Systems 36 (2023) 41451– 41530
2023
-
[33]
Ghose, A
P. Ghose, A. Bashir, Y. Wang, C. Bua, A. Zahid, Yolo-sam agriscan: A unified framework for ripe strawberry detection and segmentation with few-shot and zero-shot learning, Sensors 25 (24) (2025) 7678. 13
2025
This paper was first reviewed by grok-4.3 on June 29, 2026.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.