{"id":"c5993173-0844-4ab3-913a-34e3dfc36f21","arxiv_id":"2507.20519","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"AgroBench is a new expert-annotated benchmark showing that current vision-language models, especially open-source ones, struggle with fine-grained agricultural identification such as weed species.","lead":"This paper introduces AgroBench, a benchmark of 4,342 expert-annotated questions across seven agricultural tasks for vision-language models. It finds open-source models perform near chance on weed identification, while GPT-4o achieves 73.5% overall.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Human validators score at chance on AgroBench's headline tasks (WID 20.0%, DID 25.0%), so label validity is unestablished; the claim that VLMs lack fine-grained agricultural knowledge needs an expert label-ability study to land.","rationale":"The strongest claim is a negative capability finding: open-source VLMs near chance on weed/disease identification implies they lack the fine-grained visual knowledge needed for practical agriculture. That inference is valid only if AgroBench is a sound measurement instrument — labels correct, images answerable, distractors plausible-but-distinct. The reader's weakest assumption identifies exactly this, and I agree with it; my independent reading of Section 3.2, Table 2, and Sections 4.1-4.2 makes it the single most load-bearing concern. The decisive evidence against the instrument is internal to the paper: humans with agricultural degrees scored 20.0% (WID) and 25.0% (DID), at the chance floor, and each question was answered by only two raters, so the validation cannot certify either label correctness or image answerability. The WID task specifically inherits labels from four external datasets (Section 3.2), so the 'expert-annotated' guarantee is weakest on the very task the abstract headlines. I also flag the anomalous random baselines (several below the 20% five-option floor, e.g., DMN 15.64%) as a calibration inconsistency that should be corrected regardless of outcome; if the true baseline is 20%, the qualitative gap between open- and closed-source models survives, so this is secondary. I considered but set aside the unfair model-versus-human comparison (different subsets and settings) and the text-only ablation (Section 4.3, Table 3), which shows models answer DMN/CMN/TM well above chance without images — that contradiction weakens the 'image reference necessary' claim but is peripheral to the identification headline. Credit is due for what is genuinely solid: a single expert agronomist curated most tasks, management QA pairs were written without LLM knowledge, the category distributions are reported in detail, and the evaluation is reproducible from released code if the WID crops are actually provided. Nothing here suggests bad faith; the issue is that one-person annotation plus an underpowered, at-chance human validation cannot carry the validity burden of a benchmark whose headline is a negative result. The proposed inter-annotator label-ability study settles the direction: expert-level agreement near 80%+ clears the benchmark and converts the conditional to acceptance; expert accuracy near chance means the near-random VLM scores are an artifact of the instrument, and the central claim is unsupported.","tokens_in":33423,"tokens_out":14946,"duration_ms":149453,"concrete_test":"Run a label-ability study: give three independent agronomists (weed-science and plant-pathology specializations, not authors) a stratified random sample of 100 WID and 100 DID items with the same images, bounding-box crops, and five-option format, blind to the published ground truth; also have the annotator-author re-answer the same 200 items. Report per-expert agreement with the published labels and Fleiss' kappa across experts. If independent experts match the published labels on ≥80% of items with κ ≥ 0.6, the labels are a valid yardstick and the verdict can move to ACCEPT. If expert accuracy is below 40% (barely above the 20% chance floor) or agreement with the published labels is below 65%, the near-random VLM scores are better explained by label ambiguity or unanswerable crops, and the headline claim about missing VLM knowledge is unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central finding — that most open-source VLMs perform close to random on fine-grained identification, especially weed identification — requires that AgroBench's ground-truth labels be correct and that its images contain enough evidence to answer. The paper's human validation does not establish this. Table 2 shows agriculture-degree participants scoring 20.0% on WID and 25.0% on DID, indistinguishable from the reported random baselines (17.9%, 21.8%). Since the protocol is five-option exact matching, the pure-chance floor is 20%; several reported baselines fall below it (DMN 15.64%, CMN 16.06%, WID 17.90%), so the reference numbers themselves are miscalibrated, and the 'close to random' comparison is uncalibrated. Two further text-grounded problems amplify the risk. First, the validators' expertise is inconsistently described: Section 3.2 says Ph.D./M.S. holders, Section 4.1 says '28 students, each holding at least a bachelor's degree'; each item was answered by only two raters, so the validation is underpowered and no per-item agreement is reported. Second, for WID — the headline task — Section 3.2 states images and labels are inherited from four external datasets, with the authors supplying code to crop and assign bounding boxes rather than evidence of per-image expert re-annotation. With a single annotator and no inter-annotator agreement, a systematic label error is not excluded. If cropped weed images are not species-discriminable, or distractors are near-synonyms, near-random VLM scores measure label ambiguity rather than missing VLM visual knowledge, and the claim that VLMs lack practical agricultural knowledge does not follow.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"AgroBench introduces a seven-task, five-option multiple-choice benchmark for evaluating vision-language models (VLMs) in agriculture, comprising 4,342 QA pairs over crop/disease/pest/weed identification and crop/disease management, machine usage, and traditional methods. The authors evaluate four closed-source and eight open-source VLMs in a zero-shot setting, report a human validation on a subset, and conclude that closed-source models outperform open-source ones and that fine-grained identification — especially weed identification (WID) — remains difficult, with open-source VLMs performing close to random. The dataset and evaluation code are released, and an error analysis attributes most failures to lack of knowledge and perceptual errors.","tokens_in":33709,"tokens_out":6160,"duration_ms":60763,"significance":"If the benchmark's labels and images are valid, AgroBench is a useful evaluation resource: it is broader in task coverage and category counts than prior synthetic agricultural VLM benchmarks, and the human-validation attempt and error taxonomy are valuable. The claim that current VLMs lack fine-grained agricultural knowledge, particularly for weed identification, is plausible and practically important. However, the validity of the ground-truth labels for the fine-grained identification tasks is not established, and the human validation itself is at odds with the central claim: agriculture-degree participants scored at or near chance on WID (20.0%) and DID (25.0%). The reported random baseline is also miscalibrated, with several task baselines below the 20% chance level. These issues mean the headline conclusion currently rests on an unverified yardstick. The paper's strengths — broad task design, expert-authored management questions, and model/error analysis — could survive a revision that establishes label quality and recalibrates the baselines.","major_comments":[{"comment":"The human validation does not establish that the fine-grained identification labels are usable. Table 2 reports human accuracy of 20.0% on WID and 25.0% on DID for participants with at least a bachelor's degree in agriculture, while the five-option chance level is 20%. With human experts at or near chance, the possibility that many WID/DID items are ambiguous or that the distractor sets are too confusable cannot be dismissed. The central claim in the abstract and §4.2 that VLMs perform close to random in weed identification and that this reflects a lack of fine-grained VLM knowledge therefore requires an expert label-ability study (for example, independent expert annotation with per-item agreement, or a larger human evaluation) to rule out the alternative that the benchmark is measuring label uncertainty rather than model competence. The current text does not report inter-annotator agreement, and each item was answered by only two raters.","section":"Table 2; §3.2, §4.1"},{"comment":"The Random Choice baseline is miscalibrated. For five-option questions the expected accuracy is exactly 20%, yet the reported Random Choice values range from 15.64% to 22.11%, including DMN 15.64%, WID 17.90%, CMN 16.06%, and TM 19.31%. These sub-20% values are used to support the 'close to random' conclusion for open-source VLMs on WID. A proper baseline should be 20% with sampling variability reported (e.g., binomial confidence intervals or a multi-seed Monte Carlo average). As reported, the comparison is not statistically grounded; for instance, EMU2Chat's 23.81% WID accuracy is numerically close to 20%, but the paper provides no error bars to establish what 'close' means relative to the 55.17% of Gemini 1.5-Pro.","section":"Table 2; §4.2"},{"comment":"The human validation protocol is described inconsistently, and the human–model comparison is not matched on the same items. Section 4.1 states that 28 students answered 20 questions each, yielding 280 questions with two responses per question (560 responses), while the Table 2 caption says validation was on 80 samples per task, which would be 560 samples for seven tasks. More importantly, human scores are computed on this small subset, whereas all model scores are computed on the full test sets (e.g., DID has 1,502 QA pairs); the §4.2 statement that 'closed-source VLMs achieve better results than humans' is therefore not a like-for-like comparison. The authors should either evaluate models on the identical human subset or collect human responses on a matched sample of the full set.","section":"§4.1, Table 2 caption, §4.2"},{"comment":"The dataset is described as covering '682 disease categories' in the abstract and Figure 1, but §3.1 states that DID provides 682 crop-disease combinations covering 370 disease categories. Appendix Figures 8–13 plot 'Plant-Disease Pair' counts, confirming that 682 is the number of crop-disease pairs, not distinct diseases. This overstates the category coverage and should be corrected throughout. In addition, §3.2 states that WID images and labels are inherited from existing datasets [11,23,24,42] with code that crops images and assigns bounding boxes; no expert re-annotation of weed species is described. Since WID is the headline task on which open VLMs are 'close to random', the provenance of its labels is directly relevant to the main conclusion and must be clarified, and ideally validated by expert re-annotation.","section":"Abstract; §3.1; §3.2; Figure 1; Figures 8–13"}],"minor_comments":[{"comment":"The caption says text-only performance is 'close to random', but the reported text-only overall accuracies are 29.71% (GPT-4o) and 37.49% (LLaVA-Next-8B) versus the random baseline of 19.03%; this phrasing is misleading and should be revised.","section":"Table 3 caption"},{"comment":"References [6] and [7] list the same paper ('Are we on the right way for evaluating large vision-language models?') and should be merged or one removed.","section":"References [6] and [7]"},{"comment":"There is a typo in the Figure 5 caption: 'T wo-shot' should be 'Two-shot'.","section":"Figure 5 caption"},{"comment":"The x-axis label 'T otal' contains an erroneous space and should be 'Total'.","section":"Figure 6"}],"recommendation":"major_revision","confidential_remarks":"The paper is a potentially useful benchmark contribution, but the headline conclusion about VLMs lacking fine-grained agricultural knowledge is only as strong as the label validity and baseline calibration. The internal inconsistency between the human validation (chance-level expert scores on WID/DID) and the random baseline (sub-20% values) suggests the evaluation protocol needs substantial rework before the central claim can be trusted. I would not reject, as these issues are addressable with additional experiments and analysis, but the revision needs to be thorough."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know three things about this paper. First, it is a real contribution: an expert-annotated, multi-task benchmark for agricultural VLMs, covering disease, pest, weed, crop management, machinery, and traditional methods, with far broader category coverage than the synthetic GPT-4-generated predecessors. Second, the main evaluation is straightforward and reproducible: 12 VLMs, zero-shot multiple-choice, exact-match scoring. Third, the paper's central claim—that open-source VLMs are near-random on weed identification—rests on label validity that the paper does not actually establish.\n\nWhat is new and good: the benchmark itself, especially the breadth and the expert annotation for the DID/DMN/CMN/MQA/TM tasks. The error analysis (lack-of-knowledge vs. perceptual errors) is a useful framework. The text-only ablation is honest and shows that several management tasks can be answered without images, which is itself a useful finding, though it undercuts the claim that image reference is necessary.\n\nThe soft spots are real but addressable. The human validation description is internally inconsistent: Section 3.2 says Ph.D./M.S. holders review questions, while Section 4.1 says 28 students with at least a bachelor's degree answered a subset, each item seen by only two raters. That is underpowered, and no per-item agreement is reported. The human scores on WID (20.0%) and DID (25.0%) are statistically indistinguishable from chance, which cuts two ways: either the images are genuinely that hard for non-specialists, or many questions are ambiguous. The paper needs an expert label-ability study or at least an inter-annotator agreement number to decide which. Also, the '682 disease categories' claim is actually 682 crop-disease combinations across 370 distinct diseases; that inflation should be fixed. The random baseline numbers (15.64, 16.06, etc.) are below the 20% chance floor for five-option exact match, so the 'close to random' comparison is miscalibrated. For WID specifically, images and labels are inherited from external datasets with only cropping code, not per-image expert re-annotation, which weakens the 'expert-annotated' selling point for the headline task.\n\nWho is this for? Anyone building or evaluating agricultural VLMs. It deserves a serious referee, but with major revisions: fix the validation protocol, report agreement, correct the category counts, recalibrate the random baselines, and be more careful about what 'expert-annotated' means for WID. My own verdict is conditional, not dismissive.","headline":"A genuinely useful expert-annotated agricultural VLM benchmark, but the human validation and label-validity evidence are weaker than the headline claims and need tightening before the 'VLMs lack agricultural knowledge' conclusion fully lands.","tokens_in":34280,"tokens_out":2063,"would_cite":true,"duration_ms":23435,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A new expert-annotated benchmark of 4,342 agricultural questions shows most open-source vision-language models score near chance on weed identification.","keywords":["AgroBench","vision-language models","agricultural computer vision","weed identification","crop disease identification","expert annotation","multiple-choice question answering","VLM evaluation"],"falsifier":"If an independent panel of agronomists—who have not seen the benchmark—took the weed and disease identification questions under the same five-option, forced-choice format and also scored near the 20 percent random baseline, that would indicate the benchmark is measuring label ambiguity or unanswerable questions rather than VLM capability. Since the paper's own human validation shows agricultural degree-holders scoring 20-25 percent on these tasks, this check is directly relevant; a larger panel reproducing near-chance human accuracy would force a reinterpretation of the claim that VLMs specifically lack fine-grained agricultural knowledge.","tokens_in":33225,"feed_emoji":"🌾","tokens_out":8356,"duration_ms":78992,"temperature":0.7,"pith_summary":"AgroBench is a new benchmark for evaluating vision-language models in agriculture, built from 4,342 multiple-choice questions across seven tasks: disease identification, pest identification, weed identification, crop management, disease management, machine usage, and traditional farming methods. Unlike earlier agricultural VLM datasets that relied on synthetic GPT-generated annotations, every question in AgroBench was created and reviewed by agronomists, and the image set covers 203 crop categories and 682 disease categories. The paper reports that closed-source models such as GPT-4o and Gemini outperform open-source models overall, but that all models lag on fine-grained identification, with most open-source models at or near random guessing on weed identification. Error analysis attributes the majority of failures (about 52 percent) to missing domain knowledge rather than perception or reasoning. If the benchmark results hold, they show that current VLMs are not yet reliable for practical agricultural identification and that AgroBench provides a yardstick for measuring future progress.","feed_headline":"Open-source vision models near random on weed ID","feed_subtitle":"An expert-annotated 4,342-question benchmark shows where agricultural AI models fall short.","key_machinery":"The central machinery is AgroBench itself, a benchmark organized as seven vision-language tasks, each using five-option multiple-choice questions with a single correct answer. Images were curated from licensed and public sources, with one author holding a Ph.D. in agriculture selecting clear images and another set of agricultural degree-holders reviewing the QA pairs. For the identification tasks, the benchmark deliberately includes misleading distractor options—diseases with similar symptoms, pests associated with the target crop, and weeds of similar appearance. AgroBench also computes task accuracy as an average across tasks rather than QA count, so that large categories do not dominate the overall score. This design is what lets the paper attribute failures to category-specific knowledge gaps.","core_discovery":"The central discovery of this paper is the construction and evaluation of AgroBench, an expert-annotated, multiple-choice benchmark of 4,342 QA pairs across seven agricultural tasks. On this benchmark, the authors establish that current vision-language models, and open-source models in particular, are far from reliable on fine-grained agricultural identification: weed identification proves the hardest task, with most open-source models performing near the random baseline, and disease identification also scores well below the management tasks. The paper's evaluation further shows that closed-source models outperform open-source models on average, and that the most common error type is lack of knowledge (51.92 percent), followed by perceptual error (32.69 percent). The authors argue that these results demonstrate the need for more fine-grained agricultural visual knowledge in VLMs and that AgroBench can serve as a comprehensive evaluation tool for future agricultural VLM development.","pith_inferences":["The near-chance human scores on weed and disease identification in the paper's own validation raise the possibility that many items are extremely difficult even for trained agronomists, so a directly usable AgriBench variant might need to distinguish 'model lacks knowledge' from 'question is borderline unanswerable' before being used to certify field-ready AI.","A natural next experiment is to measure how much accuracy improves when a model is given access to a weed or plant encyclopedia as retrieval context; since most errors are classified as lack of knowledge, retrieval-augmented answering is a testable extension of the paper's diagnosis.","If the benchmark were expanded with per-question expert confidence ratings or multiple independent labelers, it could be used to quantify label noise and produce a difficulty-calibrated subset for tracking VLM progress on truly diagnostic items."],"forward_implications":["Open-source vision-language models are not yet ready for weed identification in real farm settings; most perform near random guessing on the 108 weed species covered.","Improving fine-grained visual knowledge—especially for weeds, pests, and subtle disease symptoms—is a concrete next target for agricultural VLM development, since over half of errors are traced to missing knowledge rather than perception or reasoning.","Because closed-source models clearly outperform open-source ones on average, model scale and web-scale training appear to matter for agricultural expertise; progress in open-source agricultural AI can be measured against the closed-source scores in AgroBench.","The five-option multiple-choice format and per-task score averaging make AgroBench a stable, reproducible yardstick for comparing future models, including both identification and management-related tasks."],"supporting_citations":[{"why":"Supplies the weed images and bounding-box setup used in the weed identification task, the hardest task in the benchmark.","marker":"[42]"},{"why":"Synthetic GPT-generated agricultural benchmark whose annotation approach AgroBench positions itself against.","marker":"[20]"},{"why":"Synthetic agricultural VLM dataset that represents the prior state of the art in agricultural VLM evaluation.","marker":"[48]"},{"why":"The best-performing closed-source VLM in the paper's evaluation, setting the reference accuracy on all seven tasks.","marker":"[32]"},{"why":"An open-source VLM family whose near-random weed identification scores drive the paper's main negative result.","marker":"[18]"},{"why":"A general VLM benchmark whose protocol and error categories the paper adapts for its error analysis.","marker":"[55]"},{"why":"An existing plant disease dataset used to motivate the need for broader crop and disease coverage.","marker":"[29]"}],"fun_headline_variants":["Open-source AI near random at weed ID in ag benchmark","Expert-built AgroBench exposes open-source VLM gaps","Weed ID stumps most open-source vision models","AgroBench: open-source VLMs barely beat chance on weeds","New ag benchmark: open-source models fail fine-grained ID"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The correct answers supplied by the expert annotators are accurate and unambiguous, so that a low model score is evidence of a model limitation rather than a bad or debatable question.","fun_headline_variants_meta":{"raw":{"variants":["Open-source AI near random at weed ID in ag benchmark","Expert-built AgroBench exposes open-source VLM gaps","Weed ID stumps most open-source vision models","AgroBench: open-source VLMs barely beat chance on weeds","New ag benchmark: open-source models fail fine-grained ID"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00024,"raw_usage":{"total_tokens":1501,"prompt_tokens":912,"completion_tokens":589,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":528,"completion_tokens_details":{"reasoning_tokens":508}},"tokens_in":528,"tokens_out":589,"duration_ms":6563,"temperature":1.0,"reasoning_tokens":508,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T17:42:00.348513+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"If an independent panel of agronomists—who have not seen the benchmark—took the weed and disease identification questions under the same five-option, forced-choice format and also scored near the 20 percent random baseline, that would indicate the benchmark is measuring label ambiguity or unanswerable questions rather than VLM capability. Since the paper's own human validation shows agricultural degree-holders scoring 20-25 percent on these tasks, this check is directly relevant; a larger panel reproducing near-chance human accuracy would force a reinterpretation of the claim that VLMs specifically lack fine-grained agricultural knowledge.","supporting_citations":[{"cited_title":"The cropandweed dataset: A multi-modal learning approach for efficient crop and weed manipulation","cited_arxiv_id":null,"evidence_quote":"Supplies the weed images and bounding-box setup used in the weed identification task, the hardest task in the benchmark."},{"cited_title":"A multimodal bench- mark dataset and model for crop disease diagnosis","cited_arxiv_id":null,"evidence_quote":"Synthetic GPT-generated agricultural benchmark whose annotation approach AgroBench positions itself against."},{"cited_title":"Gpt-4o, 2024","cited_arxiv_id":null,"evidence_quote":"The best-performing closed-source VLM in the paper's evaluation, setting the reference accuracy on all seven tasks."},{"cited_title":"Visual instruction tuning","cited_arxiv_id":null,"evidence_quote":"An open-source VLM family whose near-random weed identification scores drive the paper's main negative result."},{"cited_title":"Mmmu: A massive multi-discipline multimodal understand- ing and reasoning benchmark for expert agi","cited_arxiv_id":null,"evidence_quote":"A general VLM benchmark whose protocol and error categories the paper adapts for its error analysis."},{"cited_title":"Mohanty, David P","cited_arxiv_id":null,"evidence_quote":"An existing plant disease dataset used to motivate the need for broader crop and disease coverage."}],"review_version":1}