{"id":"5d7fdb9c-9a36-4ae5-93d9-66f8d17aa607","arxiv_id":"2508.17417","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":5,"one_line_summary":"Richer text prompts from LLM synonyms and cleaner image regions from activation maps improve zero-shot vision-language classification.","lead":"This paper describes a way to improve zero-shot image recognition by making text prompts richer and image crops cleaner before matching them. A smart generalist might care because it is a low-cost, training-free recipe that could make vision-language models more reliable without labeled examples.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"TGSSG's persistent-homology selection has no proven link to semantic completeness; if it prunes a rare class sense, the 'comprehensive' text-prompt set collapses and zero-shot accuracy can drop.","rationale":"The reader's weakest assumption is exactly the load-bearing point: TGSSG's LLM-generated synonymous set must be semantically complete and not omit or distort class senses. In the abstract, the only support for that completeness is the combination of semantic ambiguity entropy and persistent homology, but neither is shown to track human semantic categories. This is a concrete, testable threat to the central claim, not merely a disagreement with consensus. Because the actual full text is unavailable (the supplied text is corrupted and even contains a different arXiv identifier), I cannot check whether the authors already validate this with experiments. That does not make the paper fraudulent; it makes the evidence base unverifiable. I therefore keep the reader's UNVERDICTED verdict. The proposed test would settle the concern: if sense recall is high and restoring pruned senses does not improve accuracy, the concern is resolved. If the test fails, the method's key contribution is undermined.","tokens_in":3594,"tokens_out":6153,"duration_ms":70485,"concrete_test":"Select 50 ImageNet classes with two or more visually distinct senses (e.g., bat, crane, bank, seal). Run TGSSG with the paper's default LLM and persistence parameters. Have three annotators label each generated synonym as matching sense A, sense B, or neither. Measure sense-level recall: the fraction of known senses that appear in the final selected prompt set. If any known sense is missing for any class, the 'comprehensive' claim is falsified. Then restore the missing sense and rerun the reported matching/accuracy evaluation; if accuracy improves on those classes, the pruning step is actively harmful.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim depends on TGSSG producing semantically complete textual prompts: for every visual sense of a class that the VLM may encounter, at least one matching text prompt must survive selection. The proposed pipeline filters LLM-generated synonyms using semantic ambiguity entropy and persistent homology. Persistent homology is a geometric/topological criterion, not a semantic one: it keeps clusters that persist across a scale range in embedding space. There is no reason a rare but discriminative sense (e.g., 'crane' as a machine vs. a bird, 'seal' as an animal vs. an emblem) will form a persistent cluster at the chosen threshold. A small sense may be pruned as noise, or two senses may merge if their embeddings are close, so the resulting set is not comprehensive. If that happens for any nontrivial fraction of classes, the set-to-set alignment can match images to incomplete or misleading text, and the claimed zero-shot improvement can degrade. The supplied full text is corrupted and carries a header for arXiv:2508.17418v2, so no experiments, ablations, or baselines are available to demonstrate that the persistence cutoff preserves semantic coverage.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes Constrained Prompt Enhancement (CPE), a method intended to improve zero-shot generalization of vision-language models. It consists of two components: TGSSG, which uses LLM-generated synonymous text prompts per category filtered by semantic ambiguity entropy and persistent homology, and CADRS, which selects discriminative image regions using activation maps. A set-to-set matching strategy through test-time adaptation and optimal transport is introduced. The abstract claims improved visual-textual alignment and zero-shot generalization. However, the supplied full text is severely corrupted and includes a header for an unrelated arXiv paper; no experimental results, benchmark tables, ablations, error bars, or implementation details are present. The central empirical claim is therefore not verifiable from the manuscript as provided.","tokens_in":3930,"tokens_out":4905,"duration_ms":56312,"significance":"If substantiated, the method addresses real limitations in VLM prompting: hand-crafted prompts can be semantically incomplete, and random crops introduce visual noise. The use of persistent homology for semantic text-prompt filtering is a novel angle, and the combination of LLM-generated synonyms with activation-map region selection is a reasonable pipeline. However, the manuscript supplies no evidence: there are no benchmarks, baselines, ablations, or error bars, and no code or data release is indicated. The significance of the contribution is entirely hypothetical at this stage. The paper's potential is visible only in the abstract, not in any demonstrable result.","major_comments":[{"comment":"The central claim — that CPE 'improves zero-shot generalization of VLMs' — is unsupported. The supplied manuscript contains no experimental results: no benchmark tables, no comparisons to existing prompt-based methods, no ablations, no error bars, and no hyperparameter settings. Since this is an empirical methods paper, the absence of validation is a load-bearing failure. Without experiments, the claimed improvement cannot be assessed.","section":"Overall/Abstract"},{"comment":"The body of the manuscript is corrupted (mojibake) and carries the header 'arXiv:2508.17418v2 [physics.chem-ph] 8 Jan 2026', which is inconsistent with the claimed paper identity. This prevents verification of any derivations, algorithm descriptions, figures, or references. It also raises a submission-integrity concern: if this file is the manuscript under review, the submission is not in a reviewable state.","section":"Full Text"},{"comment":"The abstract asserts that TGSSG constructs 'comprehensive textual prompts' based on semantic ambiguity entropy and persistent homology. No theoretical or empirical support links persistent-homology persistence scale to semantic completeness. A rare but discriminative sense of a polysemous class (e.g., 'crane' as a machine vs. a bird) may not form a persistent cluster at the chosen threshold and could be pruned as noise. An ablation comparing PH-based selection with simpler filtering (e.g., clustering, entropy-only, or random sampling) is necessary to justify this design choice and its effect on zero-shot accuracy.","section":"TGSSG"},{"comment":"Key implementation details are missing. The activation-map region selection threshold, the number of LLM-generated synonyms per category, the entropy and persistence thresholds, and the test-time adaptation and optimal-transport hyperparameters are not specified anywhere in the supplied text. Even if experiments existed, the method would not be reproducible without these details.","section":"CADRS / set-to-set matching"}],"minor_comments":[{"comment":"The phrase 'and so improve zero-shot generalization' is grammatically awkward; consider 'thereby improving zero-shot generalization of VLMs.'","section":"Abstract"},{"comment":"The mathematical notation and equations are largely unreadable due to encoding corruption. The authors should ensure that the PDF is rendered with proper character encoding before resubmission.","section":"Full Text"},{"comment":"The manuscript lacks a limitations section and a broader-impact statement, which are expected in a complete submission. More importantly, no statement of data/availability or reproducibility plan is provided.","section":"Full Text"}],"recommendation":"reject","confidential_remarks":"I want to flag for the editor that the full text supplied to me is not legible and appears to contain a header from a different paper (arXiv:2508.17418v2, physics.chem-ph). This may be a technical error in the review pipeline, but if it reflects the actual submission, the paper is not in a reviewable state. I recommend rejection because the central empirical claim has no visible support in the manuscript as provided. If the correct full text exists, a revised submission containing complete experimental validation and implementation details would be worth reconsidering."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nYou should know two things: the abstract describes a coherent, training-free method that combines LLM-generated synonym prompts, persistent-homology-based selection, activation-map region cropping, and set-to-set matching (OT and TTA); and the full text we were sent is garbled and carries a header for a different arXiv paper, so every empirical claim in the abstract is currently unverified.\n\nWhat's actually new: the combination itself. LLM prompt generation, region cropping, and OT matching are all known, but using persistent homology to prune synonym sets for semantic coverage is not something I've seen in the related work cited in the abstract. That's a real novelty, and the motivation—incomplete text and noisy visual prompts—is standard enough that the approach is easy to explain. If the gains are real, this could be a useful plug-in for zero-shot VLM evaluation.\n\nSoft spots, in proportion. First, we have no experiments. The corrupted PDF means no benchmark tables, ablations, error bars, or implementation details. That's a load-bearing gap: the central claim 'so improve[s] zero-shot generalization' has no visible support in the material we received. Second, the stress-test worry about TGSSG is legitimate. Persistent homology is a topological criterion; it keeps clusters that persist in embedding space, but persistence is not coextensive with semantic importance. A rare but discriminative sense (crane-as-machine vs. crane-as-bird) could be pruned as noise, making the 'comprehensive' prompt set incomplete. The abstract doesn't address that, and we can't check whether the body does. That's a genuine risk, not a manufactured one. Third, there are several thresholds—entropy cutoff, persistence threshold, region selection—and we can't tell whether these were tuned on the target benchmarks; a careful version should report sensitivity. Minor by comparison: TTA technically uses unlabeled test data, so 'zero-shot' is transductive; that's standard in this literature, but the authors should say so.\n\nWho's this for: anyone working on prompt engineering or test-time adaptation for CLIP-like models. The idea is interesting enough that I'd want to see the full paper. What I'd recommend to the editor: ask the authors to resubmit a clean PDF, then send it to review. The current artifact can't be evaluated, but the abstract warrants referee time.","headline":"A plausible and genuinely novel training-free prompt/region enhancement recipe for zero-shot VLMs, but the supplied PDF is corrupted, so the central accuracy claims can't be checked; the persistent-homology selection is the piece to probe.","tokens_in":4322,"tokens_out":3414,"would_cite":false,"duration_ms":38331,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Constrained prompt enhancement improves zero-shot vision-language generalization by replacing incomplete text prompts and noisy image crops with semantically selected prompt sets.","keywords":["zero-shot generalization","vision-language models","prompt ensembling","persistent homology","test-time adaptation","optimal transport","discriminative region selection","semantic alignment"],"falsifier":"Run a benchmark with class names that have multiple distinct senses (for instance, 'seal', 'crane', 'bank') and compare CPE against a variant where a human manually adds the missing sense; if accuracy improves, the TGSSG set was incomplete. Also compare activation-map region selection against random crops on a dataset with small, off-center objects; if accuracy does not drop when regions are chosen randomly, the noise-filtering claim is not load-bearing.","tokens_in":3562,"feed_emoji":"🎯","tokens_out":3216,"duration_ms":38088,"temperature":0.7,"pith_summary":"This paper argues that zero-shot vision-language models fail mainly because their text prompts are incomplete and their visual prompts are noisy, and that both defects can be fixed from the semantic side without retraining. It proposes constrained prompt enhancement: generate a broad set of synonymous descriptions per class with a large language model, then prune them by semantic ambiguity and persistent homology to keep a comprehensive but compact text set; separately, select discriminative image regions using activation maps to replace random crops. The text set and region set are then matched as sets, either by test-time adaptation or optimal transport. If the method works as claimed, zero-shot classification accuracy on standard benchmarks improves beyond existing prompt-ensembling and region-alignment approaches.","feed_headline":"Two prompt sets, pruned and matched, lift zero-shot accuracy","feed_subtitle":"LLM-written synonyms and activation-map regions are aligned as sets, cutting noise without retraining.","key_machinery":"The two named mechanisms are TGSSG (Topology-Guided Synonymous Semantic Generation) and CADRS (Category-Agnostic Discriminative Region Selection). TGSSG generates synonymous descriptions per class, scores them by semantic ambiguity entropy, and applies persistent homology—a topological way of tracking when clusters of meanings appear and disappear—to select a comprehensive yet compact text set. CADRS derives activation maps from a frozen vision encoder and selects the discriminative regions, producing compact visual prompts. These sets are aligned with set-to-set matching, either by test-time adaptation or by optimal transport.","core_discovery":"The central claim is that visual-textual alignment in vision-language models improves when both sides are treated as sets selected by semantic constraints. TGSSG builds a synonymous semantic set for each class with a large language model, then uses semantic ambiguity entropy and persistent homology to choose the most informative descriptions, producing comprehensive textual prompts. CADRS uses activation maps from a pre-trained vision encoder to pick compact discriminative regions, filtering out the noise that random cropping introduces. The final set-to-set matching, via test-time adaptation or optimal transport, aligns these prompt sets and improves zero-shot generalization.","pith_inferences":["A natural extension is to apply the text-side pipeline alone to any vision-language model and the region-selection side alone to any cropping-based method; the paper's reported gains may decompose into two independent contributions.","Persistent homology may be a sufficient but not necessary selection tool; a cheaper clustering-based pruning might reproduce the same text sets, which would suggest the topological step is a means rather than the essential mechanism.","The approach could transfer to image-text retrieval or open-vocabulary detection, where aligning multiple text descriptions with multiple image regions is also the core operation.","A stress test that would expose the method's limits is a benchmark where class names have several distinct senses—for example 'seal' or 'crane'—because the assumption that the LLM-generated synonym set covers all relevant meanings is exactly what TGSSG relies on."],"forward_implications":["Zero-shot classification accuracy on standard benchmarks should rise without any model fine-tuning, because both text and image side noise are reduced before matching.","The method should beat simple prompt ensembling because it targets missing class senses rather than just adding more descriptions.","Activation-map region selection should make visual prompts more robust than random cropping, particularly on datasets where object location and size vary.","Set-to-set matching with optimal transport provides a principled way to align text and image sets even when the two sets have different sizes.","The same two-stage design could be applied to other vision-language tasks that rely on matching a query to multiple candidate representations."],"supporting_citations":[],"fun_headline_variants":["Pruned prompt sets align text and vision for zero-shot gains","Semantic sets prune text and visuals to boost zero-shot VLMs","Set-to-set matching of curated prompts improves zero-shot","LLM and activation maps clean VLMs' zero-shot prompts"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The method assumes the LLM-generated set of synonyms for each class covers every important meaning of the class name, and that the topological pruning never drops a meaning that matters for the downstream task.","fun_headline_variants_meta":{"raw":{"variants":["Pruned prompt sets align text and vision for zero-shot gains","Semantic sets prune text and visuals to boost zero-shot VLMs","Set-to-set matching of curated prompts improves zero-shot","LLM and activation maps clean VLMs' zero-shot prompts"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000858,"raw_usage":{"total_tokens":3575,"prompt_tokens":770,"completion_tokens":2805,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":514,"completion_tokens_details":{"reasoning_tokens":2735}},"tokens_in":514,"tokens_out":2805,"duration_ms":22707,"temperature":1.0,"reasoning_tokens":2735,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T16:51:58.303090+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a benchmark with class names that have multiple distinct senses (for instance, 'seal', 'crane', 'bank') and compare CPE against a variant where a human manually adds the missing sense; if accuracy improves, the TGSSG set was incomplete. Also compare activation-map region selection against random crops on a dataset with small, off-center objects; if accuracy does not drop when regions are chosen randomly, the noise-filtering claim is not load-bearing.","supporting_citations":[],"review_version":1}