{"id":"2deb1d5f-8b11-42d3-8286-2562d5d9f525","arxiv_id":"2506.03182","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A new dynamic benchmark using real field-collected images of rare insects shows frontier vision-language models are strong at coarse taxonomic classification but nearly fail at species-level identification and are inconsistent at abstaining on novel species.","lead":"This paper introduces TerraIncognita, a benchmark of around 200 insect species that combines expert-labeled images of known species with field photos of rare, potentially undescribed insects, and tests 12 leading vision-language models. The top models reach over 90% F1 at the insect Order level on known species but fall below 2% at the Species level, revealing a steep coarse-to-fine difficulty cliff.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Novelty claim for the 12 labeled 'Novel Genus' specimens may be unsupported, undermining the reported 55–88% discovery-accuracy spread.","rationale":"The reader's concern centers on training-data contamination: the 100 Novel specimens could be memorized by closed models. That is an unverifiable premise and is honestly acknowledged in §3.1; I agree it caps confidence but I do not think it is the most actionable issue. My check targets the discovery-accuracy metric itself. The abstract's second headline number — 'Discovery accuracy varies from 55% to 88%' — is presented as measuring abstention, but §4.1's definition is ambiguous for images with partial labels. The 12 Genus-labeled specimens are the only place this ambiguity bites, but 12/100 is enough to move a percentage point or two and, more importantly, there is no code or per-image scoring released to verify any of the table values. A benchmark whose central evidence is a derived score needs its scoring function to be unambiguous; currently the 'depending on label availability' clause admits two incompatible readings. This is a correctness risk internal to the paper, not a disagreement with outside consensus. The strong order-to-species F1 staircase on Known images (Table 4) is independently interesting and likely robust, so I would not reject; I would hold the paper to CONDITIONAL until the scoring and per-image records are public or the metric is redefined precisely.","tokens_in":15161,"tokens_out":1347,"duration_ms":11756,"concrete_test":"Release (or independently reconstruct from the public Hugging Face dataset) the per-image prediction records for the 12 Genus-labeled Novel images, and recompute Novel discovery accuracy under two explicit scorings: (a) abstain at Species only, and (b) abstain at a level matching the most specific available label (Genus for those 12, Species-or-higher for the rest). If the 55–88% spread shrinks or reorders under (a), the central discovery-accuracy claim depends on an unmotivated scoring choice.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The 'Novel' set is the crux of the discovery claim, but the paper's own Table 1 shows 12 of 100 novel specimens carry a Genus label. The metric definition in §4.1 says a correct Novel prediction is abstention at the Species level 'or higher levels, depending on label availability.' If the 12 Genus-labeled images are scored as correct only when the model abstains at Genus (i.e., emits 'Unknown' at Genus), then a model that happens to abstain strictly more often at finer levels gets inflated discovery accuracy, and a model that correctly guesses the expert-labeled genus is penalized. The reported novel discovery-accuracy values (55.27–87.76) are computed somewhere, but no per-image scoring rubric, no code, and no confidence intervals are provided, so the headline 'abstention behavior varies widely across models' is not reproducible. This is not about training-data contamination; it is about whether the number in the abstract measures what the paper says it measures.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces TerraIncognita, a benchmark for evaluating vision-language models on hierarchical taxonomic classification of insects, using two subsets: Known species (iNaturalist images with full labels) and Novel species (field-collected images of rare, potentially undescribed taxa with only partial labels). The authors evaluate 12 frontier multimodal models under a unified zero-shot prompting protocol, reporting order-level F1 above 90% for some models, species-level F1 below 2% for most, and a discovery-accuracy range of 55–88% on novel specimens. They also provide a qualitative expert review of model explanations, identifying failure modes such as morphological hallucination and taxonomic overreach, and commit to quarterly dataset updates.","tokens_in":15299,"tokens_out":4753,"duration_ms":44072,"significance":"If the reported numbers hold, TerraIncognita is a valuable addition to biodiversity AI evaluation: it uses real field-collected images of rare taxa rather than simulated novelty, involves expert entomological curation, evaluates a broad set of frontier models under a controlled prompt, and includes a structured qualitative analysis of model reasoning. The sharp drop from coarse- to fine-grained taxonomic accuracy is a clinically clear finding, and the public dataset plus longitudinal update plan are concrete strengths. However, the central quantitative claims currently rest on an underspecified metric and an unverifiable novelty assumption, so the benchmark's immediate usefulness depends on whether these issues are resolved.","major_comments":[{"comment":"The discovery-accuracy metric is not precisely defined. The text says that for Novel images, 'the correct behavior is abstention (i.e., predicting “Unknown”) at the Species or higher levels, depending on label availability.' Because Table 1 reports that the Novel set has 43 distinct genera and only 12 species-level labels, it is unclear whether a model must abstain at the finest available label or only at Species. Without a per-image scoring rule, the reported values (55.27–87.76) cannot be reproduced. Please provide the exact algorithm, the raw per-image predictions, and code that maps predictions to the reported numbers.","section":"§4.1, Table 3"},{"comment":"The claim that the Novel specimens are 'unseen by the model' is asserted rather than verified. Since the evaluated models are closed API systems, the paper cannot rule out that the 100 Novel images or near-duplicates appeared in training data. The Section 3.1 clarification appropriately softens the language to 'very rare, and for which almost surely few (or no) labeled images exist,' but the abstract and Section 5 still interpret the 55–88% numbers as open-world discovery accuracy. This is a load-bearing assumption for the paper's central message. I recommend reframing the results as measuring behavior on rare, expert-curated specimens, and adding a control experiment, such as testing on iNaturalist species that were collected after model training cutoffs, to bound the contamination risk.","section":"§3.1, §5"},{"comment":"The abstract's statement that species-level F1 drops 'below 2%' is contradicted by Table 4, where Gemini-2.5-Flash achieves 3.00% Species F1. The accurate claim is that all models are below 3%, and only one exceeds 2%. Please correct the abstract and the corresponding sentence in Section 5 so that the quantitative summary matches the reported data.","section":"Abstract, Table 4"},{"comment":"The evaluation reports point estimates without confidence intervals or significance tests. Given that the Novel subset has only 237 images and the Known subset 200, the spread among mid-range discovery-accuracy values (e.g., 75.11 vs. 75.53) is within binomial noise, and even the reported range of 55.27–87.76 would benefit from interval estimates. The paper states that 'All TerraIncognita data, results, and future updates are available,' but the current supplement does not include the raw per-image model outputs, the scoring script, or code to reproduce Tables 3 and 4. Please release these artifacts to make the benchmark reproducible.","section":"Tables 3 and 4"},{"comment":"There is an internal inconsistency about the number of Novel specimens with genus-level labels. Table 1 reports 43 distinct genera and 12 species-level labels in the Unknown category, while Section 7.3 states that 'we include five examples labeled up to Genus.' Since the discovery-accuracy rubric depends on available label depth, this discrepancy affects the denominator of the metric and must be reconciled.","section":"Table 1, §7.3"}],"minor_comments":[{"comment":"The column header 'F1.' is unexplained; please define how the F1 value combines Known and Novel discovery accuracy, or rename it to something unambiguous.","section":"Table 3"},{"comment":"The figure caption includes the stray text 'Loading [MathJax]/extensions/MathMenu.js', which appears to be a rendering artifact and should be removed.","section":"Figure 3"},{"comment":"The observation 'Gemini 2.5 models' is inaccurate because only one Gemini model (Gemini-2.5-Flash) is evaluated.","section":"§5, Comparative trends"},{"comment":"The prompt description in Figure 4 asks for explanations of 50–60 words per level, while Prompt 1 in Section 7.2 requests no explanations and Prompt 2 requests a comma-separated format. Please clarify which prompt was used for which quantitative result, since both are referenced as the evaluation framework.","section":"§4 and §7.2"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Read it. It's a genuine new evaluation resource: field-collected images of rare insects from Central/South America with partial expert labels, paired with iNaturalist knowns, and a quarterly refresh plan. The central quantitative trend—order-level F1 near 90–100% for several models falling to <2% at species level across all 12 models—looks robust and is probably the paper's most valuable result. The qualitative taxonomy of failure modes (morphological hallucination, speculative inference, taxonomic overreach) is also a useful framing.\n\nThe soft spots are real but mostly fixable. The discovery accuracy metric is under-specified. The definition \"abstention at the Species or higher levels, depending on label availability\" is ambiguous for the 12 novel specimens that carry genus labels. Without a per-image scoring rubric or released evaluation code, the reported 55–88% spread is not reproducible. The stress-test's specific worry—that a model correctly predicting an expert-labeled genus could be penalized for not abstaining—is one plausible reading; the opposite reading (where abstaining at genus is penalized) is also plausible. Either way, the authors need to specify exactly how each specimen was scored. The authors also provide no confidence intervals on a ~100-image benchmark, where differences between 55% and 88% are not all significant.\n\nThe unverifiable \"novel\" claim for closed API models is a structural limitation, but the authors openly acknowledge it in section 3.1, and the benchmark's value as a periodically refreshed testbed survives that caveat. The dataset is Lepidoptera-heavy, which limits the breadth of the conclusions.\n\nAll that said, the core finding—a sharp taxonomic difficulty gradient for frontier VLMs—doesn't depend on the contested metric. I'd send this to a serious referee. The authors have shipped a new dataset, an honest baseline, and a clear statement of limitations. The evaluation code and precise metric definitions are the missing pieces for the next version. I recommend engaging with it through full peer review.","headline":"Useful new benchmark with a robust coarse-to-fine performance drop, but the discovery-accuracy metric needs precise definition and released code before the headline spread is credible.","tokens_in":15882,"tokens_out":3653,"would_cite":true,"duration_ms":31888,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper introduces TerraIncognita, a benchmark that pairs images of well-known insect species with field-collected images of rare, possibly undescribed species, and asks twelve frontier vision-language models to classify each image down…","keywords":["insect biodiversity","benchmark dataset","vision-language models","open-set recognition","zero-shot classification","taxonomic hierarchy","out-of-distribution detection","species discovery"],"falsifier":"Conduct a reverse image search of the 237 novel images against public and commercial image corpora likely used in model training; if a substantial fraction return near-duplicate matches, the discovery accuracy results would measure memorization. Alternatively, run the identical prompt on an open-weight model whose training data is known to exclude these species and compare its discovery accuracy to the API models; if the open model performs similarly, the claim of open-world generalization is strengthened, but if it performs much worse, the API models' scores may reflect training contamination.","tokens_in":14973,"feed_emoji":"🦋","tokens_out":4664,"duration_ms":41845,"temperature":0.7,"pith_summary":"The paper introduces TerraIncognita, a benchmark that pairs images of well-known insect species with field-collected images of rare, possibly undescribed species, and asks twelve frontier vision-language models to classify each image down the taxonomic hierarchy (Order, Family, Genus, Species) while abstaining when uncertain. It finds a steep competence gradient: on known species, top models exceed 90 percent F1 at the Order level but fall below 2 percent at the Species level, and on novel specimens, the ability to correctly abstain (discovery accuracy) ranges from 55 to 88 percent across models. The benchmark is designed to test whether AI can move from classifying the known to flagging the unknown, something central to biodiversity conservation.","feed_headline":"Insect AI drops from 90% order accuracy to 2% species","feed_subtitle":"A new field-collected benchmark shows frontier models cannot yet support automated species discovery.","key_machinery":"The central object is the TerraIncognita dataset itself: 200 species (100 known, 100 novel) with roughly 450 high-resolution images, where known images come from iNaturalist research-grade entries and novel images come from light-trap field collection in Central and South America with partial expert labels. Evaluation uses a fixed zero-shot prompt asking for the four-level hierarchy with 'Unknown' for uncertain levels; metrics include discovery accuracy (correct abstention on novel and confident species-level on known), hierarchical F1, and expert-reviewed explanation categories such as morphological hallucination and taxonomic overreach.","core_discovery":"On its own terms, the paper establishes that current frontier vision-language models possess coarse taxonomic competence but not fine-grained species identification under zero-shot conditions. Across 100 known species and 100 novel, field-collected species, the same models that exceed 90 percent Order-level F1 score below 2 percent Species-level F1, and on novel specimens they either overcommit to unsupported genus or species labels or abstain, with no consistent strategy across model families. The paper interprets this as evidence that open-world species discovery is beyond current models, while coarse-level triage is feasible.","pith_inferences":["If the novel images are indeed absent from training corpora, the 55-88 percent discovery accuracy reflects genuine open-set detection; for closed API models that assumption remains unverifiable.","The contrast between near-perfect Order accuracy and near-zero Species accuracy on the same images suggests models rely on global shape or background cues rather than diagnostic anatomy, a hypothesis that could be tested with occlusion or cropping experiments.","The benchmark's design could be extended beyond insects to other hyperdiverse taxa, and its quarterly updates could double as a contamination monitor for future model releases.","A testable extension is to fine-tune an open model on iNaturalist species and re-evaluate on the same known set; if the Species-level gap does not close, the 2 percent result points to dataset bias rather than model limits."],"forward_implications":["Frontier models could serve as automatic pre-screeners that flag likely novel specimens for expert review, but cannot yet replace taxonomists at the species level.","The sharp order-to-species drop suggests model representations encode coarse morphological or contextual cues but not diagnostic fine-grained traits.","The wide variance in discovery accuracy (55-88 percent) shows abstention behavior is not reliable and needs targeted training or calibration.","Quarterly dataset refresh with new field specimens provides a way to longitudinally track whether future models improve or simply memorize familiar images.","The qualitative failure modes (hallucination, overreach) argue for explanation-aware evaluation in addition to label accuracy."],"supporting_citations":[{"why":"Supplies the research-grade iNaturalist images used as the Known species subset of the benchmark.","marker":"[12]"},{"why":"Identifies the Yanayacu Biological Station and field entomologist whose light-trap collection produced the Novel species images.","marker":"[16]"},{"why":"Defines the open-set recognition problem that the benchmark's discovery-accuracy metric is designed to evaluate.","marker":"[15]"},{"why":"Provides the prior open-set insect benchmark with simulated novelty that TerraIncognita contrasts against by using real field-collected specimens.","marker":"[29]"}],"fun_headline_variants":["AI identifies insect order 90% but species under 2%","Frontier AI flunks insect species ID: <2% F1","Species-level insect ID remains elusive for AI: <2%","90% order, <2% species: AI's steep drop on insects","Benchmark reveals AI's insect ID: great at order, awful at species"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the 100 field-collected Novel specimens are effectively unseen by the evaluated models; the paper acknowledges it cannot verify the training data of closed API models, so the discovery accuracy results could partly reflect familiarity instead of open-world generalization.","fun_headline_variants_meta":{"raw":{"variants":["AI identifies insect order 90% but species under 2%","Frontier AI flunks insect species ID: <2% F1","Species-level insect ID remains elusive for AI: <2%","90% order, <2% species: AI's steep drop on insects","Benchmark reveals AI's insect ID: great at order, awful at species"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000646,"raw_usage":{"total_tokens":2964,"prompt_tokens":937,"completion_tokens":2027,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":553,"completion_tokens_details":{"reasoning_tokens":1932}},"tokens_in":553,"tokens_out":2027,"duration_ms":13370,"temperature":1.0,"reasoning_tokens":1932,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T12:43:35.125416+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Conduct a reverse image search of the 237 novel images against public and commercial image corpora likely used in model training; if a substantial fraction return near-duplicate matches, the discovery accuracy results would measure memorization. Alternatively, run the identical prompt on an open-weight model whose training data is known to exclude these species and compare its discovery accuracy to the API models; if the open model performs similarly, the claim of open-world generalization is strengthened, but if it performs much worse, the API models' scores may reflect training contamination.","supporting_citations":[{"cited_title":"The inaturalist species classification and detection dataset","cited_arxiv_id":null,"evidence_quote":"Supplies the research-grade iNaturalist images used as the Known species subset of the benchmark."},{"cited_title":"Yanayacu biological station and center for creative studies.https: //www.facebook.com/Yanayacustation/, 2025","cited_arxiv_id":null,"evidence_quote":"Identifies the Yanayacu Biological Station and field entomologist whose light-trap collection produced the Novel species images."},{"cited_title":"Scheirer, Anderson de Rezende Rocha, Archana Sapkota, and Terrance E","cited_arxiv_id":null,"evidence_quote":"Defines the open-set recognition problem that the benchmark's discovery-accuracy metric is designed to evaluate."},{"cited_title":"Christian Schmidt, Aditya Jain, Yves Basset, Sara Beery, Maxim Larrivée, and David Rolnick","cited_arxiv_id":null,"evidence_quote":"Provides the prior open-set insect benchmark with simulated novelty that TerraIncognita contrasts against by using real field-collected specimens."}],"review_version":1}