Pith. sign in

REVIEW 4 major objections 6 minor 19 references

Beyond Visual Understanding: Introducing PARROT-360V for Vision Language Model Benchmarking

T0 review · 4 major / 6 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read A new visual-puzzle benchmark reports that top vision-language models score only 28% to 56%.

desk verdict A genuinely new puzzle-based VLM benchmark with a public dataset, but the headline claims about reasoning failure are undercut by a misdefined hallucination metric and missing baselines. read the letter →

arxiv 2411.15201 v1 pith:SEZ3OBPR submitted 2024-11-20 cs.CV cs.AI

classification cs.CVcs.AI
keywords PARROT-360Vvision-languagemodelsvisualreasoningbenchmarkmulti-stephallucinationJumblepuzzleschain-of-thoughtevaluationdatacontamination
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

PARROT-360V is a benchmark of 2,487 Jumble-style visual puzzles that asks a vision-language model to unscramble four words, read circled letters from the image, interpret a cartoon clue, and assemble a final bonus answer. The paper reports that GPT-4o, Claude-3.5-Sonnet, and Gemini-1.5-Pro score roughly 56%, 50%, and 28% on this test, far below their averages of 72% to 80% on popular benchmarks such as MMMU, ChartQA, AI2D, and MathVista. The authors' intended point is that current benchmarks overstate real-world visual reasoning because their questions can often be answered with text knowledge or single-step image matching. If the benchmark holds, it identifies a concrete weakness in current VLMs: multi-step visual integration, fine-grained letter perception, and resistance to hallucination.

What carries the argument

The central object is PARROT-360V itself: a dataset of 2,487 Jumble puzzles scraped from the internet, with ground-truth labels drawn from the solved puzzles. Each entry contains a question screenshot, four scrambled words, circled-letter positions, a visual cartoon clue, and the final bonus answer; models are prompted to work through the puzzle step by step using a chain-of-thought plan. The scoring system gives 10 points for each unscrambled word, 10 points for extracting the circled characters, and 20 points for the final bonus answer, for a maximum of 70 per puzzle; wrong or missing answers subtract 5 points, negative totals are clipped to zero, and the normalized score is reported. Hallucination rate is defined as one minus that score, so the benchmark's headline numbers bundle perception errors, reasoning errors, and invented characters into a single measure.

What would settle it

Build a fresh set of same-format Jumble puzzles that cannot be in the models' training data and compare scores with the published PARROT-360V set; if the fresh set closes the performance gap, the reported 28% to 56% scores are explained by memorization rather than by a genuine visual-reasoning limit.

Watch

Extended reading notes

Core claim

On the paper's own terms, the discovery is a performance differential. State-of-the-art models that average 0.72 to 0.80 on MMMU, MathVista, AI2D, and ChartQA score 0.28 to 0.57 on PARROT-360V (GPT-4o at 0.57, Claude-3.5-Sonnet at 0.50, Gemini-1.5-Pro at 0.28; the abstract states the range as 28% to 56%). The benchmark's design attributes this drop to tasks that cannot be answered by recall: each puzzle requires sequential word unscrambling, extraction of circled letters from the image, interpretation of a cartoon clue, and synthesis of a final bonus answer. The paper reports hallucination rates of 43%, 50%, and 72%, defined as one minus the benchmark score, and interprets the gap as evidence that conventional leaderboards inflate what VLMs can actually do in complex visual tasks.

Load-bearing premise

The load-bearing premise is that the Jumble puzzles scraped from a public source were not already memorized by the evaluated models, so their low scores measure perception and reasoning rather than failure to recall a seen puzzle.

Editorial extensions

If this is right

  • A high score on MMMU, ChartQA, AI2D, or MathVista should no longer be read as evidence that a model can handle multi-step visual reasoning; the same model can fall by more than half on PARROT-360V.
  • The benchmark's per-component scoring lets developers see exactly where a model fails, such as circled-letter recognition versus final bonus synthesis, making it a diagnostic tool rather than just a leaderboard.
  • The reported hallucination rates, however defined, point to a practical failure mode: these models introduce letters and answers not supported by the input image when forced to do fine-grained visual work.
  • If current state-of-the-art models sit at 28% to 56%, then the field's progress claims on complex visual reasoning should be re-examined, and evaluation suites should include stepwise visual puzzles of this kind.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A natural decomposition experiment the paper does not run would feed the model the four unscrambled words and the circled letters as text, leaving only the cartoon clue to interpret; if scores stay low, the bottleneck is visual-clue integration, while a large jump would implicate letter-level perception.
  • Because all three evaluated models are proprietary, PARROT-360V is hard to audit externally; running the same benchmark on open-weight models would make contamination checks and score verification possible for independent researchers.
  • The daily publication cycle of Jumble puzzles means the benchmark could be re-generated continuously; a rolling version would keep the test unseen and automatically retire the contamination assumption the fixed 2,487-puzzle set depends on.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper introduces PARROT-360V, a benchmark of 2487 Jumble-style visual puzzles collected from public internet sources, and evaluates three state-of-the-art VLMs (GPT-4o, Claude-3.5-Sonnet, Gemini-1.5-Pro) on them. The authors report large performance drops relative to popular benchmarks, with scores of 57%, 50%, and 28% respectively, and interpret the drop as evidence that current VLMs lack complex multi-step visual reasoning ability. The paper also defines a hallucination rate as 1 minus the benchmark score and reports hallucination rates of 43%, 50%, and 72% for the three models.

Significance. The benchmark resource is potentially valuable: it addresses a real gap by requiring step-by-step visual reasoning, the dataset is publicly released, and the task is genuinely multi-modal rather than a multiple-choice proxy. If the performance numbers were properly validated, the conclusion that standard benchmarks overestimate VLM reasoning would be an important contribution to the field. However, the current manuscript does not yet secure the validity of its own metrics: the hallucination rate is not measured as defined, the scoring protocol is brittle and unreported in its components, and there are no human or chance baselines. The significance is therefore conditional on substantial additional validation.

major comments (4)
  1. [Section 4.2, Eq. (4)] The hallucination rate does not measure hallucination as defined in the text. The paper defines hallucination as the frequency with which models introduce information not present in the input, but Eq. (4) sets HallucinationRate = 1 − PARROT360VScore, which counts every incorrect or missing answer component as a hallucination. Consequently, the claims in Section 5.2 that GPT-4o hallucinated 43%, Claude-3.5-Sonnet 50%, and Gemini-1.5-Pro 72% of the time are not supported by the reported computation. To draw conclusions about hallucination, the authors need a per-response annotation of whether the model actually generated content absent from the input, rather than a direct transform of the aggregate score.
  2. [Section 4.2, Eq. (3), and Section 5] The scoring protocol is too brittle to support the reported absolute scores without additional diagnostic information. Each incorrect or missing component receives −5, and the total is clipped at zero; because the bonus clue and puzzle answer depend on the four unscrambled words, a single failure can make the remaining components impossible and drive the whole puzzle to zero. The paper reports no per-component breakdown, no parsing protocol for free-form model outputs, and no inter-annotator agreement on judging correctness. Section 8 itself concedes that task complexity may obscure whether failures come from reasoning difficulty or task intricacy. Without component-level results and a human ceiling or chance-level baseline, the 28–57% scores could reflect exact-match brittleness, parsing artifacts, or output-format mismatch rather than the intended reasoning deficit.
  3. [Section 5 and Table 2] The cross-benchmark performance gap is not statistically grounded. Table 2 lists MMMU, MathVista, AI2D, and ChartQA scores without confidence intervals, number of runs, or evaluation-protocol details, and PARROT-360V scores in Figure 2 are shown without error bars. The comparison is also confounded by format: the reference benchmarks are largely multiple-choice or short-answer tasks, while PARROT-360V requires multi-step free-form answers with a rigid scoring rule. The claim that the performance drop is 'significant' therefore needs either matched statistical testing on the same models or a human baseline on the same PARROT-360V protocol.
  4. [Sections 3 and 4] The data-contamination safeguard is not sufficient as stated. Section 3 reports that puzzles were scraped from public internet sources, and Section 4 claims the setup is 'entirely novel' because the prompting is novel. Novel prompting does not make the puzzle images unseen: the same Jumble puzzles could easily appear in web-scale training corpora. The paper provides no contamination check, no temporal split, and no evidence that the evaluated models were not exposed to these exact images. If a model had encountered a puzzle, high scores could reflect memorization, whereas if the model had not, low scores could reflect unfamiliarity with the puzzle format; either way the central performance-gap interpretation requires an explicit contamination analysis.
minor comments (6)
  1. [Section 2.1] There are several grammatical issues, for example 'Rather benchmarking for VLMs should evaluate perception' and 'not adequately capturing the abilities of the model'; these should be rewritten.
  2. [Section 4] The spelling of the benchmark name is inconsistent: 'PARROT360V' appears in Section 4.1 while 'PARROT-360V' is used elsewhere.
  3. [Section 5.1] The phrase 'with-in' should be 'within', and 'required higher-order detail to reasoning' is ungrammatical; the intended meaning should be stated clearly.
  4. [Section 4.2] The sentence 'And Hallucination rate is the error rate, i.e. the proportion of incorrect predictions given by an VLM' should be rewritten; 'an VLM' should be 'a VLM', and the sentence should not begin with 'And'.
  5. [References] The paper cites Wei et al. 2022 and Wei et al. 2023 as though they were two separate papers, but the reference list entries are the same arXiv paper; this should be corrected.
  6. [Abstract] The phrase 'scored between 28 to 56 percentage' should be 'scored between 28% and 56%' or 'between 28 and 56 percentage points'.

Circularity Check

1 steps flagged · score 2.0 of 10

Hallucination-rate findings are definitional restatements of the benchmark score, while the central benchmark comparison is independent.

  1. self definitional [Section 4.2, Eq. (4)]
    "Hallucination within PARROT-360V benchmarking is relates to the frequency with which models introduced information not present in the input. And Hallucination rate is the error rate, i.e. the proportion of incorrect predictions given by an VLM: HallucinationRate = 1−(P ARROT360Vscore) (4)"

    Section 4 defines hallucination as the model introducing information not present in the input, but Eq. (4) defines HallucinationRate as 1 minus the aggregate PARROT-360V score. That score is computed from exact-match components with -5 penalties for wrong or missing answers (Eqs. 1-3), so it already counts every error type. The reported hallucination rates of 43%, 50%, and 72% are therefore exactly the complements of the reported scores of 57%, 50%, and 28% by construction. The claim that models hallucinate at these rates reduces to the scoring formula rather than to any independent measurement of invented content.

full rationale

The benchmark's central performance-gap claim (28-56% on PARROT-360V versus higher scores on MMMU, ChartQA, AI2D) is an external evaluation: the puzzles are scraped, the scores are computed from exact-match components, and no fitted parameter is relabeled as a prediction. The only definitional collapse is the hallucination metric: Eq. (4) sets hallucination rate equal to 1 minus the benchmark score, so the hallucination findings in Section 5.2 are tautological restatements of the score rather than independent measurements. This does not affect the raw benchmark scores, but it does mean the hallucination claims cannot support conclusions about models inventing unseen characters. No load-bearing self-citation or uniqueness argument appears, so the overall circularity is minor and localized.

Assumptions & free parameters 4 free parameters · 3 assumptions · 1 invented entities

The benchmark introduces no theoretical entities. The main hand-chosen numbers are the scoring weights and penalty, which directly shape the reported scores. The axioms are the standard assumptions of a benchmark paper: reliable ground truth, comparable external baselines, and an objective matching rule.

free parameters (4)
  • Score weight for each unscrambled word = 10
    Hand-chosen in Section 4.1; no justification for relative importance of word unscrambling versus other components.
  • Score weight for synthesizing answer characters = 10
    Hand-chosen in Section 4.1; arbitrary weighting of the letter extraction step.
  • Score weight for the final puzzle solution = 20
    Hand-chosen in Section 4.1; gives the final bonus answer double weight without empirical justification.
  • Penalty for incorrect or missing answer = -5
    Hand-chosen in Section 4.2; the magnitude of the penalty affects the total score and therefore the reported performance percentages.
assumptions (3)
  • domain assumption The ground-truth annotations extracted from the solved puzzles are correct and unambiguous.
    Section 3 states that ground truth labels are obtained from the solved puzzle, but there is no discussion of annotation errors or ambiguous answers.
  • domain assumption The comparison scores for MMMU, MathVista, AI2D, and ChartQA cited in Table 2 are accurate and directly comparable to the PARROT-360V scores.
    Table 2 lists performance values without citing sources for each number or describing the evaluation settings used to obtain them.
  • domain assumption Exact case-insensitive matching is a sufficient measure of correctness for the model outputs.
    Section 4.2 defines correctness as exact matching after case-insensitive comparison, assuming no valid alternative phrasings exist.
invented entities (1)
  • PARROT-360V benchmark dataset independent evidence
    purpose: A new collection of 2,487 visual word puzzles used to evaluate VLMs on multi-step visual reasoning.
    The dataset is publicly released on Hugging Face, providing an external handle for verification and replication.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Beyond Visual Understanding: Introducing PARROT-360V for Vision Language Model Benchmarking." pith.science (2026). https://pith.science/paper/SEZ3OBPR

@misc{pith2026241115201,
  author       = {Pith},
  title        = {Pith review of: Beyond Visual Understanding: Introducing PARROT-360V for Vision Language Model Benchmarking},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/SEZ3OBPR}},
  note         = {Machine review of arXiv:2411.15201}
}
read the original abstract

Current benchmarks for evaluating Vision Language Models (VLMs) often fall short in thoroughly assessing model abilities to understand and process complex visual and textual content. They typically focus on simple tasks that do not require deep reasoning or the integration of multiple data modalities to solve an original problem. To address this gap, we introduce the PARROT-360V Benchmark, a novel and comprehensive benchmark featuring 2487 challenging visual puzzles designed to test VLMs on complex visual reasoning tasks. We evaluated leading models: GPT-4o, Claude-3.5-Sonnet, and Gemini-1.5-Pro, using PARROT-360V to assess their capabilities in combining visual clues with language skills to solve tasks in a manner akin to human problem-solving. Our findings reveal a notable performance gap: state-of-the-art models scored between 28 to 56 percentage on our benchmark, significantly lower than their performance on popular benchmarks. This underscores the limitations of current VLMs in handling complex, multi-step reasoning tasks and highlights the need for more robust evaluation frameworks to advance the field.

Figures

Figures reproduced from arXiv: 2411.15201 by the authors.

Figure 1
Figure 1. Sample from the PARROT-360V Dataset. et al., 2023). 3 PARROT-360V Dataset The PARROT-360V Benchmark dataset was care￾fully curated by scraping Jumble puzzles from the internet to challenge VLMs in solving complex jumbles (Redblock.ai, 2024). Each scraped puz￾zle combines various elements representing an in￾stance of gameplay, as shown in [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Performance of State-of-The-Art VLMs on PARROT-360V the image often serves merely as a backdrop to a question that could just as easily be presented as pure text, reducing the need for genuine visual understanding. PARROT-360V, by contrast, involves complex tasks like word unscrambling, bonus clue extrac￾tion, and interpreting visual elements, all requiring deep integration of visual and textual information. GPT-4o’… view at source ↗
Figure 3
Figure 3. Snapshot of the Puzzle [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: Snapshot of the Solved puzzle [PITH_FULL_IMAGE:figures/full_fig_p008_4.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

19 extracted references · 3 canonical work pages

  1. [1]

    online" 'onlinestring :=

    ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRING...

  2. [2]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...

  3. [3]

    Sotiris Anagnostidis and Jannis Bulian. 2024. https://arxiv.org/abs/2408.11865 How susceptible are llms to influence in prompts? Preprint, arXiv:2408.11865

  4. [4]

    Haonan Duan, Adam Dziedzic, Nicolas Papernot, and Franziska Boenisch. 2023. https://proceedings.neurips.cc/paper_files/paper/2023/file/f26119b4ffe38c24d97e4c49d334b99e-Paper-Conference.pdf Flocks of stochastic parrots: Differentially private prompt learning for large language models . In Advances in Neural Information Processing Systems, volume 36, pages ...

  5. [5]

    Tyna Eloundou, Sam Manning, Pamela Mishkin, and Daniel Rock. 2024. https://doi.org/10.1126/science.adj0998 Gpts are gpts: Labor market impact potential of llms . Science, 384(6702):1306--1308

  6. [6]

    Aniruddha Kembhavi, Mike Salvato, Eric Kolve, Minjoon Seo, Hannaneh Hajishirzi, and Ali Farhadi. 2016. https://arxiv.org/abs/1603.07396 A diagram is worth a dozen images . Preprint, arXiv:1603.07396

  7. [7]

    Ahmed Masry, Do Xuan Long, Jia Qing Tan, Shafiq Joty, and Enamul Hoque. 2022. https://arxiv.org/abs/2203.10244 Chartqa: A benchmark for question answering about charts with visual and logical reasoning . Preprint, arXiv:2203.10244

  8. [8]

    Yoav Mintz and Ronit Brodie. 2019. https://doi.org/10.1080/13645706.2019.1575882 Introduction to artificial intelligence in medicine . Minimally Invasive Therapy & Allied Technologies, 28(2):73--81. PMID: 30810430

Show all 19 references
  1. [9]

    Redblock.ai. 2024. Parrot-360v benchmark. Available at: https://huggingface.co/datasets/RedBlock/parrot360v

  2. [10]

    Vinay Samuel, Yue Zhou, and Henry Peng Zou. 2024. https://arxiv.org/abs/2409.09927 Towards data contamination detection for modern large language models: Limitations, inconsistencies, and oracle challenges . Preprint, arXiv:2409.09927

  3. [11]

    Sivan Schwartz, Avi Yaeli, and Segev Shlomov. 2023. https://arxiv.org/abs/2308.05391 Enhancing trust in llm-based ai automation agents: New considerations and future challenges . Preprint, arXiv:2308.05391

  4. [12]

    Yahan Tu, Rui Hu, and Jitao Sang. 2024. https://arxiv.org/abs/2409.09318 Ode: Open-set evaluation of hallucinations in multimodal large language models . Preprint, arXiv:2409.09318

  5. [13]

    Yiqi Wang, Wentao Chen, Xiaotian Han, Xudong Lin, Haiteng Zhao, Yongfei Liu, Bohan Zhai, Jianbo Yuan, Quanzeng You, and Hongxia Yang. 2024. https://arxiv.org/abs/2401.06805 Exploring the reasoning abilities of multimodal large language models (mllms): A comprehensive survey on...

  6. [14]

    Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le, and Denny Zhou. 2023. https://arxiv.org/abs/2201.11903 Chain-of-thought prompting elicits reasoning in large language models . Preprint, arXiv:2201.11903

  7. [15]

    Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022. Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems, 35:24824--24837

  8. [16]

    Sadler, Dinesh Manocha, and Amrit Singh Bedi

    Xiyang Wu, Souradip Chakraborty, Ruiqi Xian, Jing Liang, Tianrui Guan, Fuxiao Liu, Brian M. Sadler, Dinesh Manocha, and Amrit Singh Bedi. 2024. https://arxiv.org/abs/2402.10340 Highlighting the safety concerns of deploying llms/vlms in robotics . Preprint, arXiv:2402.10340

  9. [17]

    Enneng Yang, Li Shen, Guibing Guo, Xingwei Wang, Xiaochun Cao, Jie Zhang, and Dacheng Tao. 2024 a . https://arxiv.org/abs/2408.07666 Model merging in llms, mllms, and beyond: Methods, theories, applications and opportunities . Preprint, arXiv:2408.07666

  10. [18]

    Qian Yang, Weixiang Yan, and Aishwarya Agrawal. 2024 b . https://arxiv.org/abs/2407.07840 Decompose and compare consistency: Measuring vlms' answer reliability via task-decomposition consistency comparison . Preprint, arXiv:2407.07840

  11. [19]

    Xiang Yue, Yuansheng Ni, Kai Zhang, Tianyu Zheng, Ruoqi Liu, Ge Zhang, Samuel Stevens, Dongfu Jiang, Weiming Ren, Yuxuan Sun, Cong Wei, Botao Yu, Ruibin Yuan, Renliang Sun, Ming Yin, Boyuan Zheng, Zhenzhu Yang, Yibo Liu, Wenhao Huang, Huan Sun, Yu Su, and Wenhu Chen. 2024. htt...

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.