{"id":"ec786ac9-354d-4b2f-8593-ec299c38e784","arxiv_id":"2501.13964","paper_version":3,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Vision-language models can spot obvious virtual objects in AR photos but frequently miss seamlessly integrated ones, with performance dropping sharply as scene complexity increases.","lead":"The paper tests whether commercial vision-language models (GPT, Gemini, Claude) can look at augmented reality photos and correctly spot the virtual objects. It introduces a new 318-image AR dataset, finds the models often succeed on obvious virtual content but fail on realistically blended objects, and compares them with five human viewers.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Complexity labels in §4.2 are defined by how easy virtual content is to distinguish, so the main 'decline with complexity' result partly restates the rubric rather than testing VLMs independently.","rationale":"The reader flagged single-annotator complexity labels as the weakest assumption. My stress-test sharpens this: the problem is not only annotator subjectivity but construct validity. The easy/medium/hard rubric in §4.2 is defined in terms of how distinguishable virtual content is, so a monotone decline in VLM detection accuracy across these levels is expected by construction and does not independently establish that VLMs specifically struggle with seamless content. The user study in §5.3 does show humans also decline, which lends some convergent validity, but because the same labels are used, this does not break the circularity. The internal inconsistency between the Introduction's TPR numbers and the Results section (e.g., GPT's TPRP under prompt G is reported as 71.4%→11.4% in §5.1 but the Introduction claims 75.8%→22.8% for perception) adds a reporting-integrity concern. These issues do not necessarily falsify the central claim; the data may still support the applied conclusion, but the paper should either (a) provide objective, property-based complexity labels, or (b) demonstrate inter-annotator agreement and show that the decline persists when images are stratified by individual objective characteristics such as shadow correctness or physical law adherence. Given that the reader's verdict is already CONDITIONAL, my concern reinforces that condition rather than moving the verdict to reject or accept. I therefore recommend UNCHANGED, with the caveat that the conditional acceptance should explicitly require the objective-label robustness analysis described in concrete_test.","tokens_in":9967,"tokens_out":5453,"duration_ms":57423,"concrete_test":"Have three independent annotators re-label all 298 AR images using the §4.2 rubric and compute Fleiss' kappa. Then re-label the same images using a pre-registered objective rubric scored from Table 1 attributes (e.g., shadow direction/absence, object-ground intersection, render-quality proxy, size/placement plausibility) without any reference to 'obviousness' or 'seamlessness'. Recompute TPRP and TPRD for each model under prompt T stratified by the objective labels. If the monotone decline across easy/medium/hard survives with kappa ≥ 0.6 and using objective labels, the central claim is robust; if kappa is low or the decline flattens/reverses under objective labeling, the reported trend is an artifact of the subjective complexity construct.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 4.2 defines Easy as 'obvious virtual content ... easily distinguishable from the real world' and Hard as 'high-quality virtual content seamlessly integrated ... making them more challenging to distinguish as virtual.' The central empirical result, that VLM perception/description TPR drops as complexity rises (Figures 3-4), is therefore partly circular: the independent variable is defined as the very difficulty the dependent variable measures. A single annotator assigned these labels, with no inter-rater reliability reported, so the threshold between levels is unvalidated. Even if the labels are consistent, the claim that VLMs 'struggle with seamlessly integrated content' restates the labeling criterion; any detector, including the human participants in §5.3, would be expected to show some decline under this rubric. The applied conclusion—that VLMs cannot be trusted to catch well-integrated virtual content—requires that the Hard subset be representative of seamless AR integration by objective criteria (e.g., correct shadows, physical plausibility) independent of human detectability. The paper does not provide such objectivity, and the Introduction's TPR values (75.8%→22.8%, 97.8%→34.2%) contradict the Results section (e.g., GPT 71.4%→11.4% under prompt G), further undermining confidence in the exact magnitudes.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces DiverseAR, a 318-image dataset of AR and non-AR scenes collected from commercial apps, research prototypes, and custom-built applications. It evaluates three commercial VLMs (GPT-4o, Claude-3.5-Sonnet, Gemini-1.5-Pro) on two tasks—perceiving whether an image contains AR content and describing the virtual elements—under a general captioning prompt and a task-aware prompt. The authors define three AR scene complexity levels, categorize model responses into five qualitative groups, and report perception/description true positive rates across complexity levels, plus a five-participant user study. The central claim is that VLMs are quite capable on obvious AR content (perception TPR up to 93%, description TPR around 71–74%) but their performance degrades sharply on seamlessly integrated virtual content, implying that current VLMs can screen for obvious AR failures but cannot be trusted as automated evaluators of high-quality AR scenes.","tokens_in":10174,"tokens_out":7419,"duration_ms":77612,"significance":"If the headline findings hold, this paper provides a useful first benchmark for VLM-based AR content evaluation. The dataset is publicly released, spans diverse devices and rendering conditions, and the five-category error taxonomy gives actionable insight into where VLMs confuse real and virtual objects. The paper also makes a concrete applied recommendation—VLMs are suitable for catching obvious virtual-object failures but not well-integrated content—that could inform AR quality-assurance practice. However, the main decline-with-complexity result rests on a single annotator's detectability-based complexity labels, and the reported numerical claims lack uncertainty quantification. These issues are load-bearing for the paper's central message, so the benchmark is not yet fully reliable as a reference result.","major_comments":[{"comment":"The Easy/Medium/Hard complexity definitions in Section 4.2 are stated directly in terms of how distinguishable the virtual content is: Easy is 'obvious virtual content ... easily distinguishable from the real world,' and Hard is 'high-quality virtual content seamlessly integrated ... making them more challenging to distinguish as virtual.' The central result that VLM perception/description TPR declines across these levels therefore partly restates the labeling criterion rather than independently measuring VLM performance against an objective notion of AR scene complexity. To support the applied conclusion that VLMs 'cannot be trusted to catch well-integrated content,' the Hard subset needs to be characterized by objective rendering properties (e.g., shadow direction, lighting consistency, physical plausibility) collected from multiple annotators with reported inter-rater reliability, and the analysis should be re-run with such labels or with human detectability explicitly controlled.","section":"§4.2, Figs. 3–4"},{"comment":"The paper reports model rankings and per-complexity trends without confidence intervals or significance tests. For example, Section 5.1 states that Gemini achieves 93.3% TPR_P under prompt T and that GPT's TPR_D is 73.8% versus Gemini's 71.8%, and Figure 4 shows GPT 'outperforming' the other models on medium and hard scenes. With only 79 hard-level AR images, a 2% TPR_D difference is well within binomial sampling error. The same issue affects the user-study comparison, where five participants are used to claim large human advantages on hard scenes. Provide bootstrap confidence intervals, McNemar tests for paired model comparisons, or per-item error tables before asserting that one model outperforms another or that performance significantly declines with complexity.","section":"§5.1, §5.3"},{"comment":"The reported magnitudes are internally inconsistent. The Introduction states that as complexity increases, perception TPR drops from 75.8% to 22.8% and description TPR from 97.8% to 34.2%, while Section 5.1 reports GPT's prompt-G perception decline as 71.4% to 11.4% and gives GPT's overall prompt-G description TPR as 38.9%. The abstract says description TPR 'up to 71%,' but Section 5.1 reports 73.8% for GPT under prompt T. More importantly, under the Section 4.4 definitions TPR_D = N1/(N1+N2+N3) is always less than or equal to TPR_P = (N1+N2)/(N1+N2+N3), so a description easy-level value of 97.8% alongside a perception easy-level value of 75.8% is arithmetically impossible on the same subset. Clarify whether the Introduction reports a different prompt/model aggregation, and align all reported numbers.","section":"Introduction (contributions bullet 3) vs §5.1 and §4.4"},{"comment":"The human comparison is not a clean head-to-head evaluation. Participants were explicitly informed that they would see 'a mix of AR and non-AR images,' which gives them knowledge of the base rate and primes them to search for AR content; the paper even attributes human false positives to this priming. The VLMs, by contrast, are not given this meta-information. This design inflates human TPRs and makes the reported human advantages (e.g., 27.1% higher TPR_D on hard-level images) difficult to interpret. Report the exact participant instructions, or match the information conditions, or present the user study as a primed-detection condition rather than a general human baseline.","section":"§5.3"}],"minor_comments":[{"comment":"The sentence 'Could Vision-Language Models (VLMs) offer a solution for the automated evaluation of AR-generated scenes?' appears twice in the abstract; remove the duplicate.","section":"Abstract"},{"comment":"Section 5.1 refers to a Description True Negative Rate (TNR_D) and reports it as 100%, but Section 4.4 only defines TNR_P. Add the corresponding definition or state explicitly that TNR_D is the non-AR analogue of TNR_P.","section":"§4.4, §5.1"},{"comment":"The phrase 'unrealistic attributes like informal size or placement' appears to contain a typo; 'informal size' should likely be 'unusual size' or 'abnormal size.'","section":"§4.2"},{"comment":"Figures 3 and 4 do not show numerical values or error bars, so the per-complexity TPRs cannot be checked from the figures alone. Add a table with exact per-level values and, ideally, confidence intervals.","section":"Figs. 3–4"},{"comment":"The paper says the task-aware prompt without explicitly referencing AR was chosen 'for consistency and simplicity,' but it does not state whether any reported result uses the explicit-AR prompt (prompt T is described as 'task-aware' in Figures 3–4). Clarify which prompt variant generated the numbers in Section 5.","section":"§4.1"},{"comment":"The attribution of specific VLM behaviors to the transformer self-attention mechanism (citing a general vision-transformer survey) is speculative and unsupported by the experiments; it would be safer to describe these as observed failure patterns rather than architectural explanations.","section":"§5.2"}],"recommendation":"major_revision","confidential_remarks":"This is a useful empirical benchmark paper, and the public dataset release is a positive feature. The main risk is that the headline complexity-decline result is partly built into the detectability-based complexity labels, and the statistical evidence for model rankings is thin. Both are fixable with objective or multi-annotator labels, confidence intervals, and aligned numbers across the abstract, introduction, and results. I do not see grounds for rejection, but the revisions described in the major comments are necessary before the paper can serve as a reliable reference for VLM-based AR evaluation."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know: (1) This is a genuinely useful first-step benchmark — a new 318-image AR dataset with released code, and a clean complexity-stratified evaluation of three commercial VLMs. (2) The central 'decline with complexity' result is partly baked into the labeling rubric, and the paper's own intro numbers don't match the results section. Still worth engaging, with revisions.\n\nWhat's new: DiverseAR, first dataset specifically for VLM evaluation of AR images, with images from multiple sources including their own AR apps. They test GPT, Gemini, Claude with general vs task-aware prompts, report perception/description TPR, and run a small human comparison. The observation that task-aware prompts help a lot (e.g., Gemini perception from 61.4% to 93.3%) is useful and practical. The dataset is released on GitHub, so others can build on it.\n\nWhere it's soft: The complexity levels (easy/medium/hard) are defined by how easily virtual content can be distinguished from real — 'Hard: seamlessly integrated... making them more challenging to distinguish.' That means the finding that VLMs do worse on hard images partly restates the definition of hard. Any detector, human or model, would be expected to drop on images defined by their indistinguishability. The labels come from a single annotator with no inter-rater check, so we don't know how robust the levels are. A more defensible approach would be to label hard by objective properties (correct shadows, physical plausibility) independent of human detectability. Also, the introduction cites TPR drops 'from 75.8% to 22.8% and from 97.8% to 34.2%,' but the results section gives different numbers (e.g., GPT under general prompt: 71.4% to 11.4%). That inconsistency needs fixing. The user study has only 5 participants; it's a pilot, not a strong comparison. And the 'first to apply VLMs to AR evaluation' claim is undercut by their own ViDDAR reference [23], which is exactly a VLM-based AR content detector. That should be reworded.\n\nWho it's for: Researchers working on automated AR/XR quality assessment or VLM evaluation. The benchmark is a reasonable starting point, and the released data makes it citable. The paper deserves a serious referee, but it needs revision before acceptance — chiefly, clarifying the label circularity, fixing the numeric inconsistency, and softening the novelty claim.\n\nRecommendation: Send to peer review, not desk reject. The dataset and prompt effects are real contributions; just demand the authors address the circularity and the intro/results mismatch.","headline":"A useful first-step AR benchmark with a real circularity flaw in the complexity labels and a numeric mismatch; worth reviewing with revisions.","tokens_in":10733,"tokens_out":2481,"would_cite":true,"duration_ms":22575,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Vision-language models can see obvious augmented-reality objects but routinely miss seamlessly integrated virtual content, a new evaluation shows.","keywords":["Augmented reality","Vision-language models","Scene understanding","Content evaluation","AR dataset","Virtual content perception","Scene complexity","Multimodal evaluation"],"falsifier":"Have multiple independent annotators re-label the 298 AR images without seeing the VLM results; if the easy–medium–hard ordering does not reproduce, or if a larger sample of hard images yields perception and description true-positive rates comparable to easy images, the central complexity-decline claim would collapse. Alternatively, demonstrating that a VLM fine-tuned on AR images maintains description true-positive rate above 90% on seamlessly integrated content would directly contradict the claim that current VLMs cannot judge polished AR.","tokens_in":9741,"feed_emoji":"👁️","tokens_out":9303,"duration_ms":85298,"temperature":0.7,"pith_summary":"This paper asks whether commercial vision-language models (VLMs) can automatically evaluate augmented-reality (AR) scenes — that is, whether they can see virtual content placed in real photos and say what it is. To answer it, the authors built DiverseAR, a dataset of 318 AR and non-AR images spanning three levels of scene complexity, and tested GPT, Gemini, and Claude with two kinds of prompts. They report that the models identify obvious virtual objects well, reaching a perception true-positive rate as high as 93% and a description true-positive rate up to 71%, but that performance collapses as virtual content becomes seamlessly integrated with realistic shadows and lighting, with reported ranges falling to 22.8% for perception and 34.2% for description on the hardest scenes. The paper concludes that today's VLMs could screen for glaring AR artifacts but are not yet dependable automated judges of polished AR experiences.","feed_headline":"VLMs ace simple AR scenes, fail on seamlessly blended ones","feed_subtitle":"New DiverseAR dataset shows GPT, Gemini, and Claude catch obvious virtual objects but miss realistic ones.","key_machinery":"The argument is carried by two newly designed instruments. The first is DiverseAR, a dataset of 318 images collected from commercial AR apps, research prototypes, and public sources, with each AR image labeled as easy, medium, or hard according to the obviousness of its virtual content. The second is a five-category response taxonomy—accurate AR recognition, partial recognition, missed recognition, false detection, and accurate non-AR recognition—that converts each VLM answer into a countable class. Two prompts are used, one general ('what is happening?') and one task-aware (asking whether virtual content is superimposed), and performance is scored with true-positive rates for perception (is it AR?) and for description (is the virtual content correctly identified and named?). This apparatus produces the paper's main empirical curve: both metrics decline monotonically as complexity rises.","core_discovery":"The central claim, stated on the paper's own terms, is that state-of-the-art commercial VLMs are generally capable of perceiving and describing AR scenes, yet their competence is sharply conditional on how well the virtual content is integrated. On DiverseAR, the perception true-positive rate reaches 93% and the description true-positive rate reaches 71% when the best-performing model and a task-aware prompt are used, but both metrics fall steeply as scenes move from 'easy' (transparent or low-fidelity overlays) to 'hard' (high-quality virtual objects with proper shadows, plausible size, and physical consistency). The paper characterizes the failure modes: with seamlessly blended content, models miss the AR elements or misclassify real objects as virtual when the scene already contains other AR objects; with obvious content, they succeed readily. A user study shows humans also decline with complexity but outperform the best VLM on the hardest scenes, especially in describing what is virtual. The paper's takeaway is that VLMs have potential as AR quality evaluators but currently fall short of what reliable automated high-quality evaluation would require.","pith_inferences":["Because 'hard' is defined as seamless integration, part of the observed decline is built into the label; a sharper test would ask whether VLMs can identify the specific rendering cue (shadow direction, lighting mismatch, depth inconsistency) that reveals a virtual object, separating perceptual ability from label design.","The finding that hands or QR codes in the scene improve recognition suggests these models rely on contextual priors more than pixel-level rendering analysis; this predicts that VLM-based AR detectors could be improved substantially by giving the model interactive cues or a two-stage 'what looks off?' reasoning prompt, rather than only by adding more training images.","The dataset's current size (318 images, 79 of them 'hard') makes the hard-scene numbers sensitive to a few dozen examples; a scaled-up independently labeled collection would likely shift the precise percentages even if the qualitative decline persists."],"forward_implications":["A task-aware prompt that explicitly asks about virtual content raises perception true-positive rate by 30–47 percentage points across models, so prompt design is a major lever in any VLM-based AR screening tool.","Because all three models maintain a 100% true-negative rate, non-AR images are never flagged as AR; the practical bottleneck is missed detection, not false alarms.","Human and VLM performance decline in the same direction with complexity, and the images that trip up people largely overlap with those that trip up VLMs, suggesting shared cues drive both successes and failures.","For hard, seamlessly integrated AR scenes, human viewers beat the best VLM by about 8 percentage points in perception and 27 in description, so an AR quality-checking pipeline that relies on VLMs alone would currently underperform a human reviewer on precisely the content that needs the most checking."],"supporting_citations":[{"why":"Supplies a benchmark of synthetic, atypical images that motivates testing VLMs on AR content.","marker":"[2]"},{"why":"Cited for VLMs' weak depth perception, which the paper uses to explain why seamlessly integrated virtual objects evade detection.","marker":"[3]"},{"why":"Describes the self-attention mechanism the paper credits for VLMs' ability to capture spatial relationships on easier AR scenes.","marker":"[8]"},{"why":"Documents VLM hallucination, the failure mode the paper accounts for in its partial-recognition category.","marker":"[20]"},{"why":"Shows VLMs struggle with reasoning beyond visual common sense, used to explain missed physical-law violations.","marker":"[25]"},{"why":"Establishes detection of task-detrimental AR content as a research problem this work extends.","marker":"[23]"}],"fun_headline_variants":["VLMs ace obvious AR, but realistic blends trip them up","Seamless AR fools VLMs: they miss what humans see","VLMs excel at obvious AR, but realistic shadows fool them","When AR is photorealistic, VLMs fail to see virtual","AR evaluation: VLMs shine on simple, stumble on realistic"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The complexity labels (easy, medium, hard) were assigned by a single annotator and are treated as ground truth, so the finding that VLM performance declines with complexity depends entirely on those labels being reproducible.","fun_headline_variants_meta":{"raw":{"variants":["VLMs ace obvious AR, but realistic blends trip them up","Seamless AR fools VLMs: they miss what humans see","VLMs excel at obvious AR, but realistic shadows fool them","When AR is photorealistic, VLMs fail to see virtual","AR evaluation: VLMs shine on simple, stumble on realistic"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000573,"raw_usage":{"total_tokens":2740,"prompt_tokens":1014,"completion_tokens":1726,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":630,"completion_tokens_details":{"reasoning_tokens":1638}},"tokens_in":630,"tokens_out":1726,"duration_ms":11878,"temperature":1.0,"reasoning_tokens":1638,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T17:04:45.956362+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have multiple independent annotators re-label the 298 AR images without seeing the VLM results; if the easy–medium–hard ordering does not reproduce, or if a larger sample of hard images yields perception and description true-positive rates comparable to easy images, the central complexity-decline claim would collapse. Alternatively, demonstrating that a VLM fine-tuned on AR images maintains description true-positive rate above 90% on seamlessly integrated content would directly contradict the claim that current VLMs cannot judge polished AR.","supporting_citations":[{"cited_title":"Bitton-Guetta, Y","cited_arxiv_id":null,"evidence_quote":"Supplies a benchmark of synthetic, atypical images that motivates testing VLMs on AR content."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Describes the self-attention mechanism the paper credits for VLMs' ability to capture spatial relationships on easier AR scenes."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Documents VLM hallucination, the failure mode the paper accounts for in its partial-recognition category."},{"cited_title":"ROME: Evaluating Pre-trained Vision-Language Models on Reasoning beyond Visual Common Sense","cited_arxiv_id":"2310.19301","evidence_quote":"Shows VLMs struggle with reasoning beyond visual common sense, used to explain missed physical-law violations."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes detection of task-detrimental AR content as a research problem this work extends."}],"review_version":1}