{"id":"7bf7615d-933c-4d6f-96d6-e816042a6482","arxiv_id":"2607.18237","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A text-prompted perceptual metric (TPIPS) trained on a new human-judgment dataset matches human aspect-conditioned similarity choices better than existing VLMs and prior metrics.","lead":"This paper introduces TPIPS, a text-prompted image similarity metric trained on a new dataset of about one million human judgments over image triplets, each annotated with free-form visual aspects. It shows that conditioning on an aspect (like lighting or pose) makes similarity judgments less ambiguous and that the metric beats existing VLMs and perceptual metrics at matching human choices.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Aspect space is VLM-proposed in both training and OOD evaluation, so the 'many senses' and OOD-generalization claims are not yet tested on human-authored aspects; a human-aspect eval is needed.","rationale":"The reader's weakest assumption and my read align. The paper is honest and contributes a valuable benchmark, but the central 'many senses' claim rests on an aspect space that the authors do not fully control. The dataset generation (Section 3.1/A.1) asks an LLM to propose variation axes and GPT-5.2 to refine them; annotators then judge only the proposed aspects and can mark 'can't tell.' Human pruning removes proposed aspects that are not visually evident, but it cannot recover aspects that were never proposed. The OOD 2AFC set (Section A.2) uses GPT-5.2-proposed aspects for editing and researcher-chosen lists for the other three task families; in no case are the aspects generated by naive human annotators. Thus the benchmark measures performance on a VLM-filtered aspect distribution, and the reported percentage-point gaps to human consensus are conditional on that distribution. Section 6 explicitly concedes the limitation. Table 5's QARE-Bench result provides partial reassurance, but QARE's aspects are fixed, common, and fully verbalizable, so it does not settle coverage of aspects a VLM cannot name. A dedicated human-authored aspect evaluation would settle the point. This is a genuine limitation of the scope of the claims, not a reason to reject the paper: the narrower claims about a benchmark for VLM-proposable aspects and a metric that outperforms baselines on that benchmark are well supported. I therefore keep the reader's conditional verdict unchanged.","tokens_in":28550,"tokens_out":14010,"duration_ms":118760,"concrete_test":"Collect a new evaluation set from the same 2AFC image pool (or a fresh sample of the four task families) in which, for each triplet, the aspect list is authored by independent human annotators who freely list the visual attributes that differ between the reference and the two candidates, without being shown any LLM/VLM proposal. Apply the same 5-vote can't-tell/tie protocol and recompute agreement for TPIPS, Qwen3-VL-Embedding-8B, and the GPT-5.4 triplet baseline. If TPIPS retains its top ranking with a gap to human consensus comparable to the current 2AFC result, the proposal-vocabulary objection is resolved; if its advantage shrinks or disappears on aspects that did not appear in the VLM proposal lists, the 'many senses' claim is bounded by the proposer's vocabulary.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that TPIPS captures 'many senses' of similarity and 'generalizes reliably beyond the training distribution' requires the metric to be evaluated on a representative distribution of text-specified aspects. Instead, the aspect space in both the training odd-one-out set and the out-of-distribution 2AFC set is produced by LLM/VLM proposals (Qwen3-4B for prompt variation, GPT-5.2 for refinement and for per-sample editing aspects) or by fixed researcher-selected lists (compositing, NVS, 3D), then pruned by a human 'can't tell' vote. The can't-tell filter can delete proposed aspects that are not visibly varying, but it cannot add an aspect that the proposing model failed to verbalize. Consequently, TPIPS is trained and evaluated only on the intersection of the proposer's visual vocabulary and human detectability. The OOD 2AFC set shifts the image distribution (real editing/compositing/NVS/3D outputs) but does not shift the aspect-distribution mechanism, so the reported gaps to human consensus (2.8 and 10.0 percentage points) do not test whether TPIPS can handle a human-authored aspect outside the proposing VLM's vocabulary. Section 6 explicitly concedes: 'Our aspect proposals are VLM-generated, and we can systematically miss aspects that VLMs cannot capture.' This does not invalidate the narrower benchmark claims, but it is load-bearing for the advertised 'many senses' and 'reliable generalization' claims.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces TPIPS, a text-prompted image perceptual similarity metric, together with a large-scale human-annotation dataset. The dataset consists of about 25K image triplets generated by text-to-image models, with over 1M human odd-one-out judgments collected across 257K triplet-aspect conditions. Aspect conditions are proposed by LLM/VLM pipelines (Qwen3 for prompt variation, GPT-5.2 for refinement) and then pruned by a human 'can't tell' filter. The authors benchmark a broad set of existing metrics, embedding models, and VLMs, showing a gap to human consensus. They then fine-tune Qwen3-VL-Embedding-8B with late-, mid-, and early-fusion architectures, achieving the best model-human agreement on both the in-distribution odd-one-out test and a separately collected out-of-distribution 2AFC set built from image editing, compositing, novel view synthesis, and single-image 3D outputs. Finally, they demonstrate aspect-conditioned retrieval and compositional retrieval as applications.","tokens_in":28962,"tokens_out":5357,"duration_ms":51057,"significance":"The dataset and benchmark are a substantial contribution: the odd-one-out annotation protocol with multiple free-form aspects per triplet goes beyond existing scalar perceptual similarity datasets, and the paper's evaluation is careful in several respects — separate test set, standard errors, human-consensus ceiling, multiple baselines, ablations over fusion architectures and data scale, and an out-of-distribution 2AFC set spanning real vision algorithms. The authors also commit to releasing code, data, and trained models, which materially strengthens the work. If the main claims hold, TPIPS would be a useful tool for aspect-conditioned image retrieval and for fine-grained evaluation of generative models. The architecture analysis, especially the construction of symmetry and identity properties in the early-fusion model, is thoughtful and clearly described. However, as discussed below, the breadth of the 'many senses' and 'reliable generalization' claims is currently limited by the way aspect conditions are generated and evaluated.","major_comments":[{"comment":"The claim that TPIPS captures 'many senses' of similarity and 'generalizes reliably beyond the training distribution' is not fully supported for the aspect space. In the odd-one-out dataset, aspect proposals come from LLM/VLM pipelines (Qwen3-4B for prompt variation, GPT-5.2 for refinement) and are only pruned by a human 'can't tell' filter; this filter can delete proposed aspects that are not visibly varying, but it cannot add aspects that the proposing VLM failed to verbalize. In the OOD 2AFC set, the editing subset uses GPT-5.2-proposed aspects and the other subsets use fixed researcher-selected lists; none of the aspects are elicited from naive human annotators in an open-ended manner. Thus the OOD evaluation shifts the image distribution but not the aspect-generation mechanism. The paper's own limitation statement in Section 6 admits this: 'our aspect proposals are VLM-generated, an","section":"Section 3.1, Table 2, Section 6"},{"comment":"The 2AFC evaluation removes aspects whose post-filter human votes are exactly tied, described as 'not useful for evaluation.' This selection removes ambiguous comparisons from the test set, which raises both the human-consensus ceiling and model agreement. The reported 10.0 percentage-point gap to human consensus on the OOD set is therefore computed on a subset of aspects that are resolvable by majority vote. The practice is defensible, but because it affects the headline generalization number, the paper should either report the tied-aspect rate and its effect on agreement, or evaluate on the full aspect set with an appropriate chance-level handling of ties. Without this, the reader cannot tell how much of the OOD performance depends on excluding the most difficult cases.","section":"Section 3.2 / Appendix A.2"},{"comment":"The in-distribution odd-one-out test shares the synthetic text-to-image generation pipeline with the training set. This is acknowledged, and the OOD 2AFC set is a reasonable mitigation. However, the two evaluation sets differ in protocol (odd-one-out vs. 2AFC), image source, and aspect source simultaneously, so it is difficult to attribute the OOD improvement to a single factor. I do not see this as a fatal flaw, but the Discussion should explicitly state that generalization is demonstrated on OOD image distributions plus a mix of VLM-proposed and researcher-fixed aspects, not on OOD aspect distributions. A cleaner decomposition — e.g., a 2AFC set on synthetic images with the same protocol, or an odd-one-out set on algorithm outputs — would strengthen the causal claim that the model learned a transferable conditional similarity function.","section":"Section 5, Figure 5 / Table 3"}],"minor_comments":[{"comment":"Typos: 'concensus' in the abstract and 'syntheic' in Section 1 should be corrected.","section":"Abstract / Section 1"},{"comment":"The notation for the mid-fusion similarity in Eq. (3) uses w_l(c) as text-conditioned channel weights, but the text says 'averaged across image tokens, and summed over layers.' The equation's placement of the 1/T average inside the sum is clear, but a parenthetical defining T and d in the main text would help readers who skip Appendix C.","section":"Section 4.2 / Figure 3"},{"comment":"The table is dense and the 'Ours' row appears under multiple fusion families with different numbers. It would be helpful to bold the single best TPIPS variant per column and to clarify in the caption which 'Ours' rows are late, early, and mid fusion, since the main text says the late-fusion model is the default but Figure 5 highlights the early-fusion result.","section":"Table 3"},{"comment":"The identity property f_theta(x,x|c)=1 is argued by symmetry of the attention mask and tied registers. The reasoning is sound, but it relies on the two image segments being token-for-token identical after preprocessing; this should be stated explicitly in the identity paragraph, since resolution/padding differences between the two copies would break the argument.","section":"Appendix C.2"},{"comment":"For the image-editing subset, the aspect list is 'predicted by GPT-5.2 per sample and human-pruned.' It would be useful to report how many aspects survive the human-pruning step and how often the pruning changes the majority vote, to quantify the human contribution to the final aspect labels.","section":"Section 3.2 / Table 2"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a strong dataset-and-benchmark paper with a well-executed evaluation and a useful model. My main concern is scope: the advertised 'many senses' and 'reliable generalization' claims outrun the current aspect-generation protocol, which is VLM-dependent in both training and OOD evaluation. This is fixable either by adding a human-authored aspect evaluation or by conservative reframing. I would not reject the paper, but I would require this issue to be addressed before publication. The other major comment on tie removal in the 2AFC set is also a quantitative-reporting issue that should be resolved with additional analysis."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"What you should know first: this is a serious, carefully executed paper. It contributes a new large-scale dataset of about one million human triplet similarity judgments conditioned on free-form visual aspects, and a fine-tuned VLM metric (TPIPS) that narrows the gap to human consensus from 9.1% to 2.8% on the odd-one-out task and beats all baselines on an out-of-distribution 2AFC set. The evaluation is unusually thorough: separate test split, standard errors, human-consensus ceiling, many baselines, ablations, and honest reporting of the remaining gaps.\n\nThe genuinely new thing is the data: human perceptual judgments collected per aspect, not mined captions or VLM-generated labels. That is a real contribution, and the paper is careful to position it against GeneCIS, FocalLens, and Omni-Attribute. The architecture study—late, mid, and early fusion—is solid engineering, and the early-fusion construction with symmetry and identity built in by attention masking is a nice touch. They ship code, data, and models, which counts for a lot.\n\nThe main soft spot is exactly the one flagged in your stress test. The aspect space in both the training triplets and the OOD 2AFC set is proposed by LLMs/VLMs and then human-pruned. The \"can't tell\" filter can remove proposals that aren't visibly varying, but it cannot add an aspect the proposer failed to verbalize. So the \"many senses\" claim is bounded by the proposing model's vocabulary. The OOD set shifts the image distribution but not the aspect-proposal mechanism, so it does not test whether TPIPS handles human-authored aspects outside that vocabulary. The paper admits this in Section 6. I see this as a real limitation but not a fatal one; the right response is a follow-up evaluation with human-authored or otherwise broader aspects.\n\nTwo smaller things: the in-distribution test shares the synthetic T2I pipeline with training, which makes the 2AFC result the more meaningful one. And the 2AFC set drops tied aspects post hoc, which is worth noting but is not a serious flaw.\n\nNo circularity problem. The metric is trained on human annotations and evaluated on held-out annotations; nothing reduces to a fitted constant. The hyperparameters were selected on validation.\n\nBottom line: the central claims are supported on the paper's own benchmarks, and the limitations are stated rather than hidden. This paper deserves a serious referee and, if the aspect-coverage issue is addressed in revision or discussed properly, will be a useful contribution to the perceptual-metrics community.\n\nRecommendation: send to peer review.","headline":"A careful, well-executed benchmark-and-metric paper; the core claims hold, but the 'many senses' claim is bounded by VLM-proposed aspects, which the authors themselves concede.","tokens_in":29422,"tokens_out":1459,"would_cite":true,"duration_ms":14645,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues that human visual similarity is not one thing: whether two images look alike depends on which aspect — color, shape, lighting, background — is being compared, and that a perceptual metric should let a user name that aspect","keywords":["visual similarity","perceptual metric","text-prompted conditioning","vision-language models","odd-one-out triplet dataset","image retrieval","generative model evaluation","human-aligned embeddings"],"falsifier":"Build a fresh odd-one-out set in which the aspects are written by human annotators (no VLM in the loop), including attributes that are hard to verbalize or culturally specific, render triplets that vary only along those aspects, and measure whether TPIPS recovers human choices above chance. If human consensus is high but TPIPS's agreement is near chance — or if TPIPS ranks pairs nearly identically under two aspects that humans treat as opposite (e.g., 'shape' vs. 'silhouette') — the claim that the text prompt genuinely steers the comparison fails.","tokens_in":28497,"feed_emoji":"🖼️","tokens_out":15688,"duration_ms":123386,"temperature":0.7,"pith_summary":"The paper argues that human visual similarity is not one thing: whether two images look alike depends on which aspect — color, shape, lighting, background — is being compared, and that a perceptual metric should let a user name that aspect in free-form text. To make this concrete, it collects a large dataset of human odd-one-out judgments over image triplets, each annotated under multiple aspects, totaling over a million votes. It fine-tunes a vision-language model on these judgments to produce TPIPS, a text-prompted image similarity metric. On the odd-one-out test set, TPIPS narrows the gap to human consensus from 9.1% to 2.8%, and on an out-of-distribution set built from four real vision algorithms it beats all baselines in its comparison groups. If the claim holds, image similarity becomes a steerable tool for retrieval, compositional search, and generative-model evaluation, each conditioned on the specific sense of similarity that matters.","feed_headline":"2.8% gap to human agreement left by text-prompted image metric","feed_subtitle":"A vision-language model, fine-tuned on a million human votes, matches humans across aspects and transfers to new tasks.","key_machinery":"The load-bearing object is the text-conditioned pairwise similarity function f(x1, x2 | c), trained end-to-end from odd-one-out judgments: within a triplet, the excluded pair must score highest. The training signal is a softmax over the three pairwise scores matched to human vote distributions with cross-entropy. The paper compares three fusion architectures — late (independent embeddings compared by cosine), mid (per-layer activation distances with text-conditioned channel weights), and early (both images fed into one vision-language model with two tied register tokens and a symmetric attention mask, which guarantees symmetry f(x1,x2|c)=f(x2,x1|c) and identity f(x,x|c)=1 by construction). T","core_discovery":"The central claim is that visual similarity is not a single scalar but a family of distances, one per aspect (color, shape, lighting, background, pose, etc.), and that a metric can be made to pick the right family member from a free-form text prompt. To support this, the paper builds a dataset of 24,342 synthetic image triplets with 257,391 triplet-aspect conditions and 1,044,495 human votes, each vote identifying the odd-one-out under a named aspect. Fine-tuning a vision-language embedding model on this data with a softmax choice model yields TPIPS, which on the in-distribution odd-one-out test reaches 64.1–64.7% rater agreement against a 67.5% human consensus — narrowing the gap from 9.1 t","pith_inferences":["The paper's data-scaling curve (performance plateaus at 60% of training data) suggests the bottleneck is aspect diversity, not triplet count; a testable extension is to invest in harder or more diverse aspect sets rather than more repetitions.","Because aspect proposals come from the VLM, the dataset cannot include similarities that lack a verbal label; a human-written aspect corpus would directly test how much of human similarity space TPIPS actually covers.","Compositional retrieval by simple score addition is a linear approximation of logical conjunction; negative constraints or weighted aspect clauses are a natural, unexplored extension.","The identity property f(x,x|c)=1 for all c means the model cannot express that an image is more self-similar under some aspects than others; applications measuring aspect salience would need a different normalization."],"forward_implications":["A single metric can serve multiple senses of similarity: the same query image returns different nearest neighbors when prompted with 'object color' versus 'background' versus 'camera distance'.","Compositional search works by adding aspect-conditioned scores: combining a subject-matter query with a brushwork query retrieves paintings that satisfy both, and swapping one component swaps the result.","Generative models can be audited along chosen visual axes rather than by one overall number, so a model may be flagged better in lighting but worse in texture.","Training on synthetic triplets transfers to outputs of real algorithms — image editing, compositing, novel-view synthesis, and single-image 3D — so the metric is usable beyond its training distribution.","The 'overall' condition also improves generic perceptual benchmarks (BAPPS, NIGHTS), indicating the collected data captures shared perceptual structure, not just aspect-specific quirks."],"fun_headline_variants":["Text-prompted metric matches human perception across visual aspects","One metric, many senses of similarity: text-prompted perceptual distance","TPIPS: Ask for 'shape' or 'color', get a human-aligned similarity","Visual similarity isn't one number - this metric lets you pick the aspect","New metric: prompt a visual aspect, get a human-like similarity"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that every similarity sense that matters to humans can be named as a short English noun phrase by the proposing VLM; the paper concedes (Section 3.1, Section 6) that aspects a VLM cannot verbalize are systematically missing from the data, and TPIPS inherits that blind spot.","fun_headline_variants_meta":{"raw":{"variants":["Text-prompted metric matches human perception across visual aspects","One metric, many senses of similarity: text-prompted perceptual distance","TPIPS: Ask for 'shape' or 'color', get a human-aligned similarity","Visual similarity isn't one number - this metric lets you pick the aspect","New metric: prompt a visual aspect, get a human-like similarity"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001112,"raw_usage":{"total_tokens":4476,"prompt_tokens":756,"completion_tokens":3720,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":500,"completion_tokens_details":{"reasoning_tokens":3634}},"tokens_in":500,"tokens_out":3720,"duration_ms":23623,"temperature":1.0,"reasoning_tokens":3634,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T15:35:46.055142+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Build a fresh odd-one-out set in which the aspects are written by human annotators (no VLM in the loop), including attributes that are hard to verbalize or culturally specific, render triplets that vary only along those aspects, and measure whether TPIPS recovers human choices above chance. If human consensus is high but TPIPS's agreement is near chance — or if TPIPS ranks pairs nearly identically under two aspects that humans treat as opposite (e.g., 'shape' vs. 'silhouette') — the claim that the text prompt genuinely steers the comparison fails.","supporting_citations":[],"review_version":1}