Pith. sign in

REVIEW 1 major objections 33 references

Multimodal LLMs generate up to 91 percent biologically inconsistent agricultural scenes from text prompts.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

Multimodal LLMs show 25-37% error in zero-shot agricultural image interpretation and up to 91% biologically inconsistent outputs in text-to-image generation tasks.

T0 review reviewed 2026-06-29 challenge →

load-bearing objection The 91% inconsistency claim for generated ag scenes rests on unvalidated subjective criteria with no reported sample sizes or rater agreement. the 1 major comments →

arxiv 2605.27595 v1 pith:OHDWPW2C submitted 2026-05-26 cs.CV cs.AI

Hallucination Behavior in Multimodal LLMs Across Agricultural Image Interpretation and Generation Tasks

classification cs.CV cs.AI
keywords hallucinationsmultimodal LLMsagricultural imagingimage interpretationtext-to-image generationbiological inconsistencycrop stress detection
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tests how large language models handle two agricultural imaging tasks: describing real crop images for stresses and generating new scenes from text descriptions. In interpretation tasks, models reach 63 to 75 percent zero-shot accuracy and improve to 86.8 percent with few-shot examples, yet still produce false detections and missed infections. In generation tasks, models such as GPT-5 and Gemini 2.5 Flash produce up to 91 percent scenes that violate biological or agronomic rules under relaxed prompts. The work uses domain-informed checks for inconsistency to map these error patterns across both directions.

Core claim

The paper establishes that multimodal LLMs exhibit recurring hallucinations in agricultural image tasks, with text-to-image generation showing up to 91 percent biologically inconsistent outputs in advanced models under relaxed constraints, while image-to-text interpretation shows moderate accuracy that improves with prompting but retains false detections and missed infections.

What carries the argument

Domain-informed criteria that score outputs for biological inconsistency, contextual inaccuracy, and agronomic implausibility.

Load-bearing premise

The criteria for judging biological inconsistency, contextual inaccuracy, and agronomic implausibility are accurate and free of evaluator bias when applied to model outputs.

What would settle it

Independent agronomists re-scoring the same set of generated scenes with the paper's criteria and finding a substantially lower inconsistency rate.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • Few-shot prompting reduces but does not eliminate hallucinations in image interpretation.
  • Current models show fundamental weaknesses in producing biologically plausible agricultural scenes from text.
  • Recurring error patterns appear across both interpretive and generative agricultural tasks.
  • Reliability improvements are needed before LLM-based agricultural imaging platforms can be trusted.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Similar inconsistency rates may appear in non-agricultural image generation tasks that require domain knowledge.
  • Prompt engineering alone may not close the gap without added domain-specific constraints or verification steps.
  • Deployment in real farm decision systems would require human review loops to catch the observed error types.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

1 major / 0 minor

Summary. The paper examines hallucination behaviors in multimodal LLMs for agricultural image interpretation (image-to-text) and generation (text-to-image) tasks. It reports modest zero-shot accuracy of 63-75% for models like Gemma, LLAVA, Qwen, and MiniCPM in interpreting crop stresses, improving to 86.8% with few-shot prompting, and up to 91% biologically inconsistent scenes generated by GPT-5 and Gemini 2.5 Flash under relaxed prompts, using domain-informed criteria for biological inconsistency, contextual inaccuracy, and agronomic implausibility.

Significance. If the reported hallucination rates and patterns could be reproduced with validated, reproducible evaluation protocols, the work would document concrete limitations of current multimodal LLMs in a high-stakes applied domain and could inform targeted mitigation strategies for agricultural imaging pipelines.

major comments (1)
  1. [Abstract] Abstract: The headline quantitative results (63–75% zero-shot accuracy, 86.8% few-shot accuracy, and up to 91% biologically inconsistent scenes) are stated without any accompanying sample sizes, image or prompt selection criteria, statistical tests, or inter-rater reliability statistics for the domain-informed criteria. Because the distinction between consistent and inconsistent agricultural scenes is not self-evident, these percentages cannot be assessed and constitute the central empirical claim.

Simulated Author's Rebuttal

1 responses · 0 unresolved

We thank the referee for their detailed review and constructive feedback on the abstract. We agree that the headline results require additional context for proper assessment and will revise accordingly to strengthen the manuscript's clarity and reproducibility.

read point-by-point responses
  1. Referee: [Abstract] Abstract: The headline quantitative results (63–75% zero-shot accuracy, 86.8% few-shot accuracy, and up to 91% biologically inconsistent scenes) are stated without any accompanying sample sizes, image or prompt selection criteria, statistical tests, or inter-rater reliability statistics for the domain-informed criteria. Because the distinction between consistent and inconsistent agricultural scenes is not self-evident, these percentages cannot be assessed and constitute the central empirical claim.

    Authors: We agree that the abstract as currently written does not provide sufficient supporting details for the reported percentages. In the revised manuscript we will expand the abstract to include: (1) sample sizes for both the image-to-text (N images evaluated across models) and text-to-image (N prompts per model) experiments; (2) a concise description of image and prompt selection criteria; (3) mention of any statistical tests performed; and (4) inter-rater reliability metrics (e.g., Cohen’s kappa) for the domain-informed criteria of biological inconsistency, contextual inaccuracy, and agronomic implausibility. The full Methods section already defines these criteria with examples; we will ensure the abstract is self-contained while remaining within length limits. These changes directly address the concern that the central claims cannot be assessed without this information. revision: yes

Circularity Check

0 steps flagged

No significant circularity; purely observational empirical reporting

full rationale

The paper reports measured accuracies (63-75% zero-shot, up to 86.8% few-shot) and inconsistency rates (up to 91%) from direct evaluation of LLM outputs on agricultural images and prompts. No equations, parameter fitting, derivations, or self-citations appear in the provided text. The central claims rest on external data collection and application of domain-informed criteria rather than any reduction to the paper's own inputs by construction. This matches the default expectation of non-circularity for observational studies.

Axiom & Free-Parameter Ledger

0 free parameters · 1 axioms · 0 invented entities

The central claim rests on the unverified validity of expert-defined hallucination criteria and the representativeness of the (unspecified) test images and prompts; no free parameters or invented entities are introduced.

axioms (1)
  • domain assumption Domain-informed criteria for biological inconsistency and agronomic implausibility are reliable and unbiased when applied to LLM outputs.
    Evaluation of both interpretation and generation tasks depends entirely on these criteria to classify outputs as hallucinations.

reviewed 2026-06-29 · how reviews work

0 comments
Cite this review

Pith. "Pith review of Hallucination Behavior in Multimodal LLMs Across Agricultural Image Interpretation and Generation Tasks." pith.science (2026). https://pith.science/paper/OHDWPW2C

@misc{pith2026260527595,
  author       = {Pith},
  title        = {Pith review of: Hallucination Behavior in Multimodal LLMs Across Agricultural Image Interpretation and Generation Tasks},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/OHDWPW2C}},
  note         = {Machine review of arXiv:2605.27595}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Large Language Models (LLMs) are being rapidly adopted in agricultural imaging applications, ranging from crop interpretation to synthetic field image generation. However, these models frequently exhibit hallucinations outputs that appear confident yet deviate from biological or environmental reality potentially leading to misinformed agronomic insights. This study investigates such hallucinations in two complementary directions: image-to-text, where LLMs interpret crop or field imagery to describe conditions such as biotic and abiotic stresses, and text-to-image, where models generate synthetic agricultural scenes based on descriptive prompts. We examine errors involving biological inconsistency, contextual inaccuracy, and agronomic implausibility, evaluating the outputs under domain-informed criteria across multiple imaging modalities. Our analysis identifies recurring hallucination patterns within both interpretive and generative tasks. In image interpretation, LLMs (e.g., Gemma, LLAVA, Qwen, and MiniCPM) achieved modest zero-shot accuracy (63 to 75 percent), whereas few-shot prompting improved performance up to 86.8 percent, exhibiting false detections and missed infections, indicating residual hallucination effects. In text-to-image tasks, advanced models such as GPT-5 and Gemini 2.5 Flash generate up to 91 percent biologically inconsistent scenes under relaxed prompt constraints, revealing fundamental weaknesses in current LLMs. This systematic assessment of visual reasoning and generation offers critical insights toward enhancing the reliability and trustworthiness of LLM-based agricultural imaging platforms.

Figures

Figures reproduced from arXiv: 2605.27595 by Al Bashir, Azlan Zahid, Partho Ghose, Prem Raj.

Figure 1
Figure 1. Figure 1: Representative samples of two tomato leaf diseases: Bacterial spot (left) and Septoria leaf spot [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: LLM’s generated images for the prompt: “Generate a healthy soybean canopy with extensive pest damage.” (a) GPT-5 hallucinates, producing a close-up image of leaves severely eaten by pests while still labeling the canopy as healthy. (b) Gemini 2.5 Flash generates a wide, visually coherent soybean field with green foliage but subtle pest damage, separating canopy vigor from visible stress more accurately. Th… view at source ↗
Figure 3
Figure 3. Figure 3: Model responses to the prompt: “a thermal image of the same tomato leaf, highlighting localized heat signatures from infection.” (a) Gemini Flash-2.5 (b) GPT-5. Gemini produced an artificial uniformity in surface temperature, and GPT produced sharp temperature gradients, exaggerating halo transitions around lesion margins. (a) Gemini 2.5 Flash (b) GPT-5 [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Model responses to the prompt: “generate a high-resolution RGB image of a ripe strawberry fruit infected with gray mold (Botrytis cinerea).” (a) Gemini Flash-2.5 (b) GPT-5. Gemini created an image containing unprompted soil and debris, and GPT lacked the natural irregularity of real Botrytis growth. 6 [PITH_FULL_IMAGE:figures/full_fig_p006_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Model responses to the prompt: “generate a multispectral image of a rice field affected by nitrogen deficiency.” (a) Gemini Flash-2.5 (b) GPT-5. Gemini contained unwanted environmental features such as irrigation channels and flooded areas, and GPT generated an artistic scene rather than an analytical appearance with an unprompted textual label. These unnecessary elements reflect how generative models ofte… view at source ↗
Figure 6
Figure 6. Figure 6: Model behavior under hallucination-inducing prompts: (a) Gemini 2.5 Flash hallucinates with [PITH_FULL_IMAGE:figures/full_fig_p008_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Confusion matrices for Q1(fruit presence acknowledgement) across (a) LLAVA, (b) Gemma, (c) Qwen, and (d) MiniCPM. The confusion matrices ( [PITH_FULL_IMAGE:figures/full_fig_p009_7.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

33 extracted references · 5 canonical work pages · 2 internal anchors

  1. [1]

    Chang, X

    Y. Chang, X. Wang, J. Wang, Y. Wu, L. Yang, K. Zhu, H. Chen, X. Yi, C. Wang, Y. Wang, et al., A survey on evaluation of large language models, ACM transactions on intelligent systems and technology 15 (3) (2024) 1–45

  2. [2]

    Bhayana, Chatbots and large language models in radiology: a practical primer for clinical and research applications, Radiology 310 (1) (2024) e232756

    R. Bhayana, Chatbots and large language models in radiology: a practical primer for clinical and research applications, Radiology 310 (1) (2024) e232756. 11

  3. [3]

    Tzachor, M

    A. Tzachor, M. Devare, C. Richards, P. Pypers, A. Ghosh, J. Koo, S. Johal, B. King, Large language models and agricultural extension services, Nature food 4 (11) (2023) 941–948

  4. [4]

    L. Fang, W. Xiang, J. Jin, K. Liao, C. Liu, Y. Han, F. D. Salim, Y.-P. P. Chen, Agri-llm: Prompt-based large language model for emission data analytics in smart agriculture, IEEE Internet of Things Journal (2025)

  5. [5]

    B. Yang, Y. Zhang, L. Feng, Y. Chen, J. Zhang, X. Xu, N. Aierken, Y. Li, Y. Chen, G. Yang, et al., Agrigpt: A large language model ecosystem for agriculture, arXiv preprint arXiv:2508.08632 (2025)

  6. [6]

    B., et al., Tomato leaf disease detection dataset,https://www.kaggle.com/datasets/ kaustubhb999/tomatoleaf/data, accessed: 2025-10-23 (2025)

    K. B., et al., Tomato leaf disease detection dataset,https://www.kaggle.com/datasets/ kaustubhb999/tomatoleaf/data, accessed: 2025-10-23 (2025)

  7. [7]

    Y. Zhou, C. Cui, J. Yoon, L. Zhang, Z. Deng, C. Finn, M. Bansal, H. Yao, Analyzing and mitigating object hallucination in large vision-language models, in: NeurIPS 2023 Workshop on Instruction Tuning and Instruction Following

  8. [8]

    Huang, W

    L. Huang, W. Yu, W. Ma, W. Zhong, Z. Feng, H. Wang, Q. Chen, W. Peng, X. Feng, B. Qin, et al., A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions, ACM Transactions on Information Systems 43 (2) (2025) 1–55

  9. [9]

    Z. Bai, P. Wang, T. Xiao, T. He, Z. Han, Z. Zhang, M. Z. Shou, Hallucination of multimodal large language models: A survey, arXiv preprint arXiv:2404.18930 (2024)

  10. [10]

    H.-K. Ko, G. Park, H. Jeon, J. Jo, J. Kim, J. Seo, Large-scale text-to-image generation models for visual artists’ creative works, in: Proceedings of the 28th international conference on intelligent user interfaces, 2023, pp. 919–933

  11. [11]

    Q. Yan, X. He, X. E. Wang, Med-hvl: Automatic medical domain hallucination evaluation for large vision-language models, in: AAAI 2024 Spring Symposium on Clinical Foundation Models, 2024

  12. [12]

    J.-Y. Wu, Z. Y. Poh, A. C. Patil, B. Park, G. Volpe, D. Urano, Analysis of plant nutrient deficien- cies using multi-spectral imaging and optimized segmentation model, arXiv preprint arXiv:2507.14013 (2025)

  13. [13]

    Jayakumar, A

    S. Jayakumar, A. K. Abhangrao, R. A. Sarje, R. Gupta, S. Pathania, B. V. Sree, Critical analysis on effect of micronutrients on flowering plants: a review, Int J Plant Soil Sci 36 (2024) 776–782

  14. [14]

    Taghvaeian, J

    S. Taghvaeian, J. L. Ch´ avez, J. Altenhofen, T. Trout, K. DeJonge, et al., Remote sensing for evaluating crop water stress at field scale using infrared thermography: potential and limitations (2013)

  15. [15]

    G. Team, R. Anil, S. Borgeaud, J.-B. Alayrac, J. Yu, R. Soricut, J. Schalkwyk, A. M. Dai, A. Hauth, K. Millican, et al., Gemini: a family of highly capable multimodal models, arXiv preprint arXiv:2312.11805 (2023)

  16. [16]

    T. Wei, Z. Chen, X. Yu, Snap and diagnose: An advanced multimodal retrieval system for identifying plant diseases in the wild, in: Proceedings of the 6th ACM International Conference on Multimedia in Asia, 2024, pp. 1–3

  17. [17]

    J. Liu, X. Wang, A multimodal framework for pepper diseases and pests detection, Scientific Reports 14 (1) (2024) 28973

  18. [18]

    K. I. Roumeliotis, R. Sapkota, M. Karkee, N. D. Tselikas, D. K. Nasiopoulos, Plant disease detec- tion through multimodal large language models and convolutional neural networks, arXiv preprint arXiv:2504.20419 (2025)

  19. [19]

    Zheng, J

    J. Zheng, J. Wang, Agrigpt: A strong plant disease detection model via visual-language model, in: International Conference on Intelligent Computing, Springer, 2025, pp. 239–250. 12

  20. [20]

    Y. Wang, F. Wang, W. Chen, B. Lv, M. Liu, X. Kong, C. Zhao, Z. Pan, A large language model for multimodal identification of crop diseases and pests, Scientific Reports 15 (1) (2025) 21959

  21. [21]

    J. Qing, X. Deng, Y. Lan, Z. Li, Gpt-aided diagnosis on agricultural image based on a new light yolopc, Computers and electronics in agriculture 213 (2023) 108168

  22. [22]

    R. Yan, P. An, X. Meng, Y. Li, D. Li, F. Xu, D. Dang, A knowledge graph for crop diseases and pests in china, Scientific Data 12 (1) (2025) 222

  23. [23]

    Huang, W

    N. Huang, W. Dong, Y. Zhang, F. Tang, R. Li, C. Ma, X. Li, C. Xu, Creativesynth: Creative blending and synthesis of visual arts based on multimodal diffusion, CoRR (2024)

  24. [24]

    L. C. Adams, F. Busch, D. Truhn, M. R. Makowski, H. J. Aerts, K. K. Bressem, What does dall-e 2 know about radiology?, Journal of medical Internet research 25 (2023) e43110

  25. [25]

    Seneviratne, D

    S. Seneviratne, D. Senanayake, S. Rasnayaka, R. Vidanaarachchi, J. Thompson, Dalle-urban: Capturing the urban design expertise of large text to image transformers, in: 2022 International Conference on Digital Image Computing: Techniques and Applications (DICTA), IEEE, 2022, pp. 1–9

  26. [26]

    Sapkota, Z

    R. Sapkota, Z. Meng, M. Karkee, Synthetic meets authentic: Leveraging llm generated datasets for yolo11 and yolov10-based apple detection through machine vision sensors, Smart Agricultural Technol- ogy 9 (2024) 100614

  27. [27]

    Vayadande, S

    K. Vayadande, S. Bhemde, V. Rajguru, P. Ugile, R. Lade, N. Raut, Ai-based image generator web application using openai’s dall-e system, in: 2023 international conference on recent advances in science and engineering technology (ICRASET), IEEE, 2023, pp. 1–5

  28. [28]

    Y. Kim, H. Jeong, S. Chen, S. S. Li, M. Lu, K. Alhamoud, J. Mun, C. Grau, M. Jung, R. Gameiro, et al., Medical hallucinations in foundation models and their impact on healthcare, CoRR (2025)

  29. [29]

    A. Pal, L. K. Umapathi, M. Sankarasubbu, Med-halt: Medical domain hallucination test for large language models, in: Proceedings of the 27th Conference on Computational Natural Language Learning (CoNLL), 2023, pp. 314–334

  30. [30]

    W. Shi, X. Han, M. Lewis, Y. Tsvetkov, L. Zettlemoyer, W.-t. Yih, Trusting your evidence: Hallucinate less with context-aware decoding, in: Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 2: Short Papers), 2024, pp. 783–791

  31. [31]

    Chuang, Y

    Y.-S. Chuang, Y. Xie, H. Luo, Y. Kim, J. R. Glass, P. He, Dola: Decoding by contrasting layers improves factuality in large language models, in: The Twelfth International Conference on Learning Representations, 2023

  32. [32]

    K. Li, O. Patel, F. Vi´ egas, H. Pfister, M. Wattenberg, Inference-time intervention: Eliciting truthful answers from a language model, Advances in Neural Information Processing Systems 36 (2023) 41451– 41530

  33. [33]

    Ghose, A

    P. Ghose, A. Bashir, Y. Wang, C. Bua, A. Zahid, Yolo-sam agriscan: A unified framework for ripe strawberry detection and segmentation with few-shot and zero-shot learning, Sensors 25 (24) (2025) 7678. 13

This paper was first reviewed by grok-4.3 on June 29, 2026.