{"id":"50886d5c-123b-4fba-a10e-4297ee1f43de","arxiv_id":"2608.10195","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A four-task behavioral battery scores 45 vision models against published human perception data and shows that task accuracy and human-like grouping dissociate across model families.","lead":"The authors tested 45 vision models on four visual grouping tasks to see if they organize images the way humans do, scoring them against human data from published psychology studies. They found that a model's accuracy often says little about whether it sees patterns the way people do, and some top-performing closed models are less human-like than their accuracy suggests.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Series-counting effect replication is unscorable: the 6–12 band is not a per-trial human curve, so 'a(k) unconstrained by the band' is not a tested claim.","rationale":"After reading the full manuscript and appendix, I agree with the reader's weakest assumption as the load-bearing issue. The silhouette-based dissociation (Fig. 3, Table A3) is the strongest part of the paper: it uses per-trial human data, a convention-invariant ceiling (Appendix D.2), and bootstrap intervals. The series-counting effect-replication claim is the soft spot because the human anchor is not per-trial. The paper explicitly says no per-trial human data exists for this task (Sec. 3.1) and yet asserts a cross-family failure of effect replication (Sec. 4). Appendix D.1 even restricts the task to 'capacity-band membership,' which is not the same as a fitted a(k) test. Because the 6–12 band is qualitative and comes from different stimulus types, no criterion for 'bends where human capacity does' is defined; the claim is unfalsifiable as stated. This does not require doubting the authors' integrity; it requires either collecting a human a(k) curve on the exact stimuli or softening the claim. Since the central dissociation does not depend on series counting, the verdict remains CONDITIONAL rather than REJECT.","tokens_in":17149,"tokens_out":6166,"duration_ms":63488,"concrete_test":"Run a small human validation on the exact 40 series-counting stimuli (e.g., 20 naive observers, same rendering and prompt), producing per-count human accuracy a_h(k). Then re-run the Behavioral Effect Replication analysis: define a pre-registered criterion for 'constrained by the human profile' (e.g., model a(k) within the bootstrap CI of a_h(k) at each k, or a slope/interaction test). If the empirical human curve is flat across K=2–12, the 6–12 band is not a valid anchor and the series-counting claim should be dropped. If the human curve bends near 6–12, test whether any model's a(k) lies within the human envelope; if none do, the claim survives with a real reference.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's conclusion that 'Behavioral Effect Replication fails in every family' (Sec. 4) rests on the weakest human anchor in the battery. Table 1 and Sec. 3.1 state that color-series counting is scored against a 'derived 6–12 capacity band [12, 35, 29]' explicitly 'because no per-trial human data exists.' Sec. 3.3 defines Behavioral Effect Replication as a(k), accuracy versus rendered series count, and says replication means 'the profile bends where human capacity does.' But the 6–12 band is a qualitative range of set-size limits from other studies, not a per-count human accuracy curve for these exact 40 stimuli. The paper never specifies a test statistic, a null, or a criterion by which a model's a(k) is 'constrained by' the band. Appendix D.1 itself limits series counting to 'capacity-band membership' and gives a CI half-width of 0.155 for n=40, yet Sec. 4 uses the task to assert a strong cross-family failure. This is load-bearing because the claimed failure of effect replication is part of the evidence that human-likeness is task-specific; without a defined human profile, 'unconstrained' is unfalsifiable. The remaining dissociation results (silhouette accuracy vs. kappa) are much better supported by per-trial human data, so the central dissociation survives; the series-counting replication claim should be removed or downgraded until a human a(k) reference exists.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces a behavioral battery that scores vision models against published human psychophysics on four Gestalt tasks: silhouette recognition (closure), mark-color odd-one-out, color-series counting, and object odd-one-out (similarity). It evaluates 15 encoders and 30 foundation models, computing Behavioral Agreement (B), trial-level Behavioral Error Consistency (kappa), and Behavioral Effect Replication (a(k)). The central finding is that silhouette accuracy and human error consistency dissociate: many models exceed human accuracy while failing to share human error patterns, and only a few closed frontier models approach or exceed the human-human consistency ceiling. The authors conclude that human-likeness is task-specific, not a global trait, and propose the battery as a reusable, user-study-free yardstick for auditing models in visualization pipelines. The manuscript includes reproducibility details, per-trial records, bootstrap CIs, and a validation of the kappa ceiling, making the core dissociation claim transparent and testable.","tokens_in":17458,"tokens_out":6556,"duration_ms":63126,"significance":"If the core dissociation result holds, this is a valuable contribution to model evaluation in visualization and beyond. The paper introduces a general methodology for reusing published perception data as evaluation targets, which is cheap and principled. The silhouette dissociation is supported by per-trial human data, a chance-corrected kappa with a validated ceiling, and bootstrap confidence intervals. The manuscript also ships reproducible stimulus generation, a schema-checked record format, and per-trial records for every model, which is exemplary practice. The main weakness is the series-counting Behavioral Effect Replication claim, which rests on a qualitative human capacity band rather than a per-trial human curve; this specific claim is unfalsifiable as stated. The remaining evidence for task-specific human-likeness is credible, and the methodological contribution is significant.","major_comments":[{"comment":"The claim in Sec. 4 that \"Behavioral Effect Replication fails in every family\" is not supported by the evidence presented. The human anchor for color-series counting is a \"derived 6–12 capacity band\" (Table 1, Sec. 3.1), which is a qualitative range from other studies, not a per-count human accuracy curve. Sec. 3.3 states that replication means \"the profile bends where human capacity does,\" but the paper never defines a test statistic, a null distribution, or a criterion for deciding that a model's a(k) is \"unconstrained by the band.\" Moreover, Appendix D.1 explicitly limits series counting to \"capacity-band membership\" and reports a bootstrap CI half-width of 0.155 for n=40, which is inconsistent with the strong, cross-family failure asserted in Sec. 4. Without a defined human a(k) profile, the claim is unfalsifiable. This is load-bearing because it is used as evidence for the task-specificity of human-likeness, so it should be removed or downgraded to a descriptive observation until a per-trial human reference or a clearly specified criterion is available.","section":"Sec. 4; Table 1; Appendix D.1"}],"minor_comments":[{"comment":"The phrase \"benchmark accuracy\" is ambiguous. The only accuracy measure used in the accuracy-vs-human-likeness dissociation is within-battery silhouette class accuracy, not a general-purpose benchmark such as ImageNet or chart-QA accuracy. Please replace \"benchmark accuracy\" with \"silhouette accuracy\" or \"within-battery accuracy\" to avoid overstatement.","section":"Abstract; Sec. 4"},{"comment":"The contrastive vision-language family is denoted with a bullet \"•\" in the abstract and text, but with a filled circle \"●\" in the Fig. 1 legend. Please use one symbol consistently across the paper.","section":"Fig. 1 legend; Sec. 1; Abstract"},{"comment":"The metric a(k) is described verbally but never defined with an equation, and the paper does not show a plot or table of accuracy versus series count k. Adding an explicit definition (even a simple one) and a supplementary figure would allow readers to assess the claimed threefold spread across encoder objectives.","section":"Sec. 3.3"}],"recommendation":"major_revision","confidential_remarks":"The central dissociation result on silhouettes is well supported and the reproducibility practices are exemplary. The series-counting effect-replication claim is a fixable overreach: the authors should remove it or clearly limit it to a capacity-band membership observation. With that change, the paper would likely be acceptable for publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things you should know. The central dissociation is real: on the Geirhos silhouette set, accuracy against class labels and error-consistency against human observers rank models almost independently (rho=0.38, 37% discordant pairs), and 34 of 45 models beat mean human accuracy while only 4 reach human–human consistency. That is a clean, bootstrapped result, and it undercuts any simple 'higher accuracy = more human-like' story. The second thing is that one headline claim is not actually supported. The series-counting 'effect replication' is scored against a qualitative 6–12 capacity band, not a per-trial human curve, and the paper never defines what 'constrained by the band' means. The appendix itself says series counting should be read only as capacity-band membership, yet Section 4 uses it to assert a cross-family failure. That claim needs to be removed or demoted until a real human a(k) profile exists.\n\nWhat is genuinely new: error-consistency scoring from Geirhos, applied to chart-based Gestalt stimuli (Demiralp kernels, THINGS triplets, silhouettes) across 45 models including closed foundation models. That combination is not in the literature, and the methodological care is real. Floors use empirical position frequencies rather than naive 1/3; the kappa derivation as rescaled covariance is clean; the ceiling validation (leave-one-out observer 0.487 vs pairwise 0.489) is exactly the kind of check that should be standard. No fitting to targets, no circularity: probes are fit on disjoint splits and human references are fixed.\n\nSoft spots beyond the series-counting overclaim: the abstract says 'benchmark accuracy' when only silhouette class accuracy is measured; the wording should change. And there is no code or data release. The paper claims end-to-end reproducibility, but without artifacts that is just a promise. For a paper whose pitch is a reusable yardstick, releasing the battery is not optional in the long run. Minor: the open-VLM tier is small, but the closed tier partially compensates.\n\nThis paper deserves a serious referee. The flaws are surgical, not existential. Send it to review, ask for the series-counting claim to be dropped or reframed, the abstract softened, and the artifacts released. I would bring it to our reading group and cite it once the series-counting fix is in.","headline":"A genuinely reusable Gestalt battery with a real silhouette dissociation, held back by an overclaimed series-counting result and no released artifacts.","tokens_in":17955,"tokens_out":3083,"would_cite":true,"duration_ms":30997,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Accuracy does not buy human-like grouping in 45 vision models","keywords":["Graphical perception","Gestalt grouping","vision science","psychophysics","behavioral evaluation","error consistency","vision models","human–model alignment"],"falsifier":"Run a fresh psychophysics experiment on the identical color-series counting stimuli, collecting per-trial human responses; if human accuracy by series count does not show a capacity-limited decline or plateau across the 6–12 range, then the claimed failure of Behavioral Effect Replication in every model family is not measurable against the human reference the paper assumes.","tokens_in":16970,"feed_emoji":"👁️","tokens_out":5550,"duration_ms":50979,"temperature":0.7,"pith_summary":"This paper argues that a vision model's benchmark accuracy says little about whether it organizes visual input the way humans do, and that agreement with published human perception data is a measurable, reusable yardstick for that difference. Across 45 models on four Gestalt grouping tasks, accuracy and human-likeness dissociate: 34 of 45 models recognize silhouettes more accurately than human observers, yet only four match human–human error consistency, and the two measures rank models discordantly in 37% of pairs (ρ=0.38). The authors build a behavioral battery from existing perception studies—silhouette recognition, mark-color odd-one-out, color-series counting, and object odd-one-out—so no new user study is needed. The upshot for visualization research is concrete: models entering chart-reading pipelines cannot be assumed to see marks and objects the way their human audience does, and which model looks human-like depends on which grouping task is measured.","feed_headline":"Accuracy doesn't buy human-like grouping in 45 vision models","feed_subtitle":"Most models beat human accuracy yet fail to share human errors; only a few frontier models do both.","key_machinery":"The carrying object is the behavioral battery itself: four forced-choice tasks reusing stimuli and human data from published perception studies, with two natural-image control tasks (silhouette closure, object odd-one-out) and two chart-track tasks (mark-color odd-one-out, color-series counting). Three metrics do the scoring: Behavioral Agreement B (how often the model's response equals the human reference), trial-level Behavioral Error Consistency κ (a chance-corrected covariance between model correctness and per-item human accuracy, scored against the human–human ceiling), and Behavioral Effect Replication a(k) (whether accuracy as a function of series count bends where human ensemble capacity does). The battery makes encoders and foundation models directly comparable by reducing every output to one scored response, and it anchors each score to source-study floors and ceilings rather than to thresholds chosen by the authors.","core_discovery":"The central claim is that human-likeness in perceptual organization is a distinct, task-specific property from task performance, and that it can be measured behaviorally by scoring models against published human responses. On their battery, human agreement B and error consistency κ against a human–human ceiling reveal dissociations invisible to accuracy: the best color-similarity encoder is mid-pack on semantic odd-one-out, the best semantic encoder is unremarkable on color, and open-weight VLMs express little of the grouping their encoder backbones carry. At the frontier, however, a few closed foundation models pair above-human accuracy with error consistency at or beyond the human–human ceiling, showing that human-like grouping is achievable, but it is earned by training recipe rather than by scale or accuracy.","pith_inferences":["The paper leaves open whether the language-alignment stage actively destroys perceptual organization; a testable follow-up would probe intermediate encoder layers of the open VLMs to see where grouping information is lost.","If the dissociation generalizes to complete charts with axes and legends, visualization guidelines validated on human viewers may need model-specific versions, not just a single human-likeness score.","The rank discordance between accuracy and κ suggests human-aligned grouping could be optimized as its own training target, with the prediction that such models would score higher on this battery without necessarily changing benchmark accuracy.","The color-series capacity band is the weakest anchor; a fresh per-trial human study on the exact stimuli would tighten or refute the Behavioral Effect Replication result."],"forward_implications":["Benchmark accuracy alone cannot certify a model as a human proxy for design evaluation; a task-indexed behavioral audit is needed.","A model that groups chart marks human-likely on one task can diverge sharply on another, so human-perception guidelines do not automatically transfer to machine readers.","Open-weight VLMs can lose the perceptual organization of their encoder towers, so the instruction-tuning stage itself may be reshaping grouping behavior.","The best closed foundation models achieve human-level error consistency on silhouettes, showing the dissociation is not an inherent limit of current architectures.","Published psychophysics data can be reused as a no-new-user-study yardstick, and the same construction extends to other Gestalt laws and chart-context stimuli."],"supporting_citations":[{"why":"Supplies the silhouette stimuli and per-trial responses of ten observers that define the human reference for closure.","marker":"[8]"},{"why":"Defines trial-level error consistency κ and the chance-correction used throughout the battery.","marker":"[6]"},{"why":"Supplies the perceptual-kernel data for mark-color odd-one-out, including the human odd-one-out target.","marker":"[5]"},{"why":"Provides the human similarity-judgment data and test–retest ceiling for object odd-one-out.","marker":"[14]"},{"why":"Provides the natural-image set and behavioral ceiling used for the object odd-one-out control task.","marker":"[13]"},{"why":"Supplies the chance-corrected agreement coefficient underlying the κ estimator.","marker":"[4]"},{"why":"Fixes the 6–12 categorical-color capacity band used as the human anchor for series counting.","marker":"[35]"},{"why":"One of the sources for the 6–12 categorical-color capacity band referenced by the series-counting task.","marker":"[12]"}],"fun_headline_variants":["Human-like vision is not about accuracy, study of 45 models shows","Most vision models beat humans but don't see like them","Accuracy alone misses how vision models organize scenes","Only a few frontier models see the way humans do","Vision models' grouping defies their benchmark scores"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the published human data remain valid for the exact stimuli shown to the models, most fragilely the 6–12 color-series capacity band, which rests on qualitative findings with no per-trial human responses; if that band is not the right signature of human ensemble segmentation, the Behavioral Effect Replication result loses its anchor.","fun_headline_variants_meta":{"raw":{"variants":["Human-like vision is not about accuracy, study of 45 models shows","Most vision models beat humans but don't see like them","Accuracy alone misses how vision models organize scenes","Only a few frontier models see the way humans do","Vision models' grouping defies their benchmark scores"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000119,"raw_usage":{"total_tokens":1048,"prompt_tokens":872,"completion_tokens":176,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":488,"completion_tokens_details":{"reasoning_tokens":98}},"tokens_in":488,"tokens_out":176,"duration_ms":2638,"temperature":1.0,"reasoning_tokens":98,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T04:10:40.861348+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a fresh psychophysics experiment on the identical color-series counting stimuli, collecting per-trial human responses; if human accuracy by series count does not show a capacity-limited decline or plateau across the 6–12 range, then the claimed failure of Behavioral Effect Replication in every model family is not measurable against the human reference the paper assumes.","supporting_citations":[{"cited_title":"Geirhos, P","cited_arxiv_id":null,"evidence_quote":"Supplies the silhouette stimuli and per-trial responses of ten observers that define the human reference for closure."},{"cited_title":"Geirhos, K","cited_arxiv_id":null,"evidence_quote":"Defines trial-level error consistency κ and the chance-correction used throughout the battery."},{"cited_title":"Demiralp, M","cited_arxiv_id":null,"evidence_quote":"Supplies the perceptual-kernel data for mark-color odd-one-out, including the human odd-one-out target."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the human similarity-judgment data and test–retest ceiling for object odd-one-out."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the natural-image set and behavioral ceiling used for the object odd-one-out control task."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the chance-corrected agreement coefficient underlying the κ estimator."},{"cited_title":"Ware.Information Visualization: Perception for Design","cited_arxiv_id":null,"evidence_quote":"Fixes the 6–12 categorical-color capacity band used as the human anchor for series counting."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"One of the sources for the 6–12 categorical-color capacity band referenced by the series-counting task."}],"review_version":1}