{"id":"ff6341f4-1fd0-4c62-a7fd-5b11e241c226","arxiv_id":"2506.19644","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A user-controlled loop of LLM-suggested attributes, CLIP-based verification histograms, and probabilistic prompt sampling enables non-experts to steer the diversity of AI-generated image sets toward their own goals.","lead":"Varif.ai is an interactive tool that lets users control how varied AI-generated images are by setting the proportions of attributes like color, style, or ethnicity. In a 20-person study, it produced more diverse image sets than plain prompt writing and matched user-specified diversity targets better than automatic prompt-diversification baselines.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The RQ1 diversity gain may be confounded by the augmentation step in Sec. 7.1.1, which appears to generate 50 images from the participant's final attribute distribution for Varif.ai but from a single prompt for Prompt-only.","rationale":"The reader's weakest assumption concerned CLIP's ability to capture user-relevant diversity dimensions. That concern is real and is partially addressed by the paper's sensitivity analysis (Sec. 7.3), which reports that actual-label alignment is robust to CLIP accuracy in a small auxiliary study. In contrast, the augmentation protocol in Sec. 7.1.1 is completely unexamined and directly affects the primary quantitative result: the span difference between Varif.ai and Prompt-only. If the augmentation generates Varif.ai samples from a user-defined distribution and Prompt-only samples from a single prompt, the comparison conflates the system's interactive process with the structure of the final specification. The paper does not provide enough detail to rule this out. Because this is an addressable methodological ambiguity rather than a demonstrated fatal flaw, the conditional verdict remains appropriate, but the manuscript should clarify the augmentation step and report diversity on the actual user-generated images.","tokens_in":25681,"tokens_out":10056,"duration_ms":109121,"concrete_test":"Recompute the RQ1 span statistic using only the images each participant actually generated in the summative study (no augmentation), or using a fixed-size sample (e.g., the first 10 images of the final iteration) with the same seed model. If the Varif.ai-vs-Prompt-only difference in span is no longer significant at p<.01, the headline diversity claim is an artifact of the augmentation protocol. Also request the authors to specify the augmentation prompts for both conditions.","verdict_should_be":"UNCHANGED","load_bearing_attack":"In Sec. 7.1.1, the authors state: 'We augment each participant’s image collection to 50 total images with the original attribute specification.' For Varif.ai, this specification is a probability distribution over attribute labels (the user's final histograms); for Prompt-only, it is a text prompt. Sampling 50 images from a multi-label distribution yields high CLIP span by construction, whereas sampling 50 images from a single prompt yields near-duplicates. If the augmentation used these asymmetric specifications, the reported span difference (Varif.ai M=0.65 vs Prompt-only M=0.45, p<.0001) does not isolate the verify-and-vary interaction; it may only demonstrate that sampling from a broad distribution produces broader embeddings than sampling from a point. The paper does not disclose how Prompt-only sets were augmented or whether diversity was also computed on the actual user-generated images before augmentation. This is load-bearing because the central claim is that the user-in-the-loop process, rather than prompt augmentation per se, produces the gains. A symmetric evaluation on the actual generated sets (or fixed-size samples from the final iteration) is needed.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces Varif.ai, an interactive system for user-driven diversity in text-to-image generation, implementing a generate-verify-vary loop: users specify attributes and label distributions via histograms, the system verifies coverage using CLIP-based classification, and varies generation by probabilistically sampling labels appended to the prompt. The authors report an elicitation study (8 participants) to identify diversity needs, a formative study (8 participants) to evaluate usability, and a controlled summative study (20 participants) comparing Varif.ai to a prompt-only baseline and to automatic diversification baselines (Promptist, GPT-4o). The central quantitative claim is that Varif.ai yields significantly higher image diversity (span M=0.65 vs. 0.45, p<.0001) than prompt-only prompting, and better alignment to user-specified target distributions.","tokens_in":25864,"tokens_out":3353,"duration_ms":34839,"significance":"If the results hold, the paper makes a useful contribution to human-AI interaction for generative image tools, providing a concrete interface for user-controlled diversity, with open-source code and a model-agnostic architecture. The study includes both qualitative and quantitative evaluations and attempts to cover multiple diversity specification degrees. The main claims are plausible and the paper is generally well-written. However, the credibility of the headline quantitative comparison hinges on the evaluation methodology, which contains a potentially serious confound in the augmentation step and a possible circularity in the diversity alignment metric.","major_comments":[{"comment":"The augmentation step is asymmetric between conditions and may confound RQ1a. The paper states: 'We augment each participant’s image collection to 50 total images with the original attribute specification.' For Varif.ai, the 'attribute specification' is the user-defined probability distribution over attribute labels (e.g., 20% red, 40% blue), so sampling 50 images from this multi-label distribution will by construction produce a high CLIP span. For Prompt-only, the specification is a single text prompt, and sampling 50 images from a single prompt yields near-duplicate outputs. The paper does not disclose how Prompt-only sets were augmented or whether the diversity metric was also computed on the actual user-generated images before augmentation. Because the reported span difference (M=0.65 vs. 0.45, p<.0001) is computed on augmented sets, it may reflect the difference between sampling from a broad distribution versus a point, rather than the verify-and-vary interaction that is the paper's central claim. The authors should recompute the comparison on the actual generated sets per participant, or use a symmetric augmentation procedure (e.g., sample the same number of images from the final iteration for both conditions).","section":"Sec. 7.2.2, Diversity Alignment"},{"comment":"The diversity alignment result (attribute-label-specific M=0.79 vs. attribute-specific M=0.74, p<.001) appears to be computed between the user-specified target distribution and the distribution that Varif.ai's own CLIP classifier reports after regeneration. If this is the case, the metric partly measures whether the system's internal measurement is consistent with the target, not whether the generated images actually contain the specified attribute proportions. This is a circularity concern because the verification step uses the same CLIP model that defines the measured distribution. The paper should clarify which labels were used for this analysis; if CLIP labels were used, the authors should provide a version of Fig. 9b computed on manually annotated labels for the participants' final image sets, or at least report the CLIP accuracy for the attributes involved. The sensitivity analysis in Sec. 7.3 is reassuring for a uniform-distribution case, but it does not directly address the alignment values reported for the actual study data.","section":"Sec. 7.1, Computed Image Diversity and DV definitions"},{"comment":"The primary dependent variable (span) is computed from CLIP embeddings, but the paper provides no validation that CLIP span correlates with human-perceived diversity. Given that the paper's own sensitivity analysis (Sec. 7.3) documents substantial CLIP misclassification rates and that the formative study reports user-visible inconsistencies between CLIP histograms and perceived attributes, the central RQ1a claim would be substantially strengthened by a human-rated diversity check—for example, having independent judges or the participants themselves rate the diversity of a subset of the final image sets. Without such validation, the reported span advantage may overstate the benefit to users.","section":"Sec. 7.1"}],"minor_comments":[{"comment":"The paper uses 'elicitation study' and 'formative study' inconsistently; Section 3 is called 'Elicitation User Study' while Section 6 is 'Formative User Study', but the abstract refers to a 'pilot validation'. Please align the terminology.","section":"Abstract and Sec. 3/6"},{"comment":"There is a typo: 'wamted' should be 'wanted' in the sentence about E6 wanting frogs in different environments.","section":"Sec. 3.2.3"},{"comment":"The phrase '3× 2condition setup' lacks spaces and a multiplication symbol; please format as '3 × 2 condition setup'.","section":"Sec. 7.1.1"},{"comment":"In Fig. 8 and the accompanying text, the p-value threshold for 'very significant' (p<.01) and 'extremely significant' (p<.0001) should be defined in the figure caption or text to avoid ambiguity.","section":"Sec. 7.2.1"},{"comment":"The comparison with automatic baselines uses only a single run per baseline per scenario, and the manual annotation is performed by the first author. The lack of error bars for the baseline measurements should be at least mentioned as a limitation in the text.","section":"Sec. 7.2.3"}],"recommendation":"major_revision","confidential_remarks":"The paper is a solid HCI contribution with a well-motivated system and a thorough qualitative evaluation. However, the central quantitative comparison in RQ1 is potentially confounded by the asymmetric augmentation described in Sec. 7.1.1, and the diversity alignment metric in Sec. 7.2.2 may be partly circular. These are fixable with reanalysis or additional reporting, but they are load-bearing for the main claims, so major revision is appropriate. The authors should also consider whether the CLIP span metric requires human-validation to support the paper's conclusions about user-perceived diversity."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The interesting thing here is not any single component—LLM label suggestion, CLIP histogram verification, probabilistic prompt sampling all exist—it's the packaging into one user loop that lets non-experts specify, check, and steer diversity. The formative work is also a plus: the elicitation findings about attributes, labels, and proportions feel real and clearly shape the design.\n\nThe paper does several things well. It cites the prior diversification work it builds on (Promptist, ITI-Gen, PromptCharm), the system is model-agnostic in principle, and the authors include a sensitivity analysis that directly addresses their own CLIP-accuracy weakness. That analysis shows actual-label diversity alignment is not significantly affected when CLIP accuracy drops, which is a meaningful honesty point and partially answers the circularity worry for the generation loop itself.\n\nNow the soft spots, in proportion. The stress-test note has real traction. Section 7.1.1 says they augment each participant's image collection to 50 images \"with the original attribute specification.\" For Varif.ai, that specification is a probability distribution over attribute labels; for Prompt-only, it is a text prompt. Sampling 50 images from a multi-label distribution will inflate CLIP span almost by construction, while sampling 50 from a single prompt will produce near-duplicates. The paper does not disclose exactly how Prompt-only sets were augmented, and it does not report diversity on the actual user-generated images before augmentation. That makes the headline RQ1 result (M=0.65 vs 0.45, p<.0001) ambiguous: it may show that broad prompt distributions yield broad embeddings, not that the verify-and-vary interaction is what helps. This is load-bearing and needs to be fixed with a symmetric evaluation on the actual sets or fixed-size samples from the final iteration.\n\nThe RQ2 alignment result has a milder version of the same issue because it partly uses the system's own CLIP classifier, but the manual annotation in Figure 12 and the sensitivity analysis give some independent support. Sample size is small (20 participants), and no user-study data or versioned code snapshot is shipped—addressable, not fatal.\n\nBottom line: the system is a useful contribution to HCI for generative tools, and the qualitative work alone is worth a reader's time. But the central quantitative claim currently rests on an incompletely specified evaluation. I would send this to peer review, and I would ask for a revision that discloses and symmetrizes the augmentation, reports the unaugmented image sets, and releases the data.\n\nWho this is for: researchers building user-driven generative tools, especially for creative ideation or fairness-oriented diversification. It deserves a serious referee.","headline":"A genuinely useful integrated loop for user-driven image diversity, with a real evaluation gap around how the RQ1 diversity metric was computed.","tokens_in":26426,"tokens_out":2252,"would_cite":true,"duration_ms":25682,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that letting users define and adjust attribute distributions in an image-generation loop yields measurably more diverse image sets than prompt-only generation, and aligns better with explicit diversity targets.","keywords":["image generation","diversity","user-driven control","human-AI interaction","text-to-image","CLIP","probabilistic prompting","creative ideation"],"falsifier":"Take Varif.ai and a prompt-only baseline, generate image sets for attributes where CLIP is known to misclassify (for instance, object-specific colors such as frog color versus background color), and have users or manual annotators judge diversity. If manual-label diversity shows no advantage for Varif.ai over prompt-only when CLIP accuracy is below roughly 0.6, the user-driven diversity claim fails. A simpler ablation would remove the verification histograms and let users vary distributions blindly; if the diversity span advantage persists, the verify step is not load-bearing.","tokens_in":25451,"feed_emoji":"🎨","tokens_out":4215,"duration_ms":40277,"temperature":0.7,"pith_summary":"The paper claims that image-generation diversity is not something a model can decide alone: different users want different things to vary, so diversity should be user-driven. It introduces Varif.ai, which turns diversity control into a three-step loop: generate a set of images, verify how well user-chosen attributes are covered using CLIP-based classification shown as histograms, and vary by sampling attribute labels probabilistically into the prompt. In a controlled study with 20 participants, Varif.ai produced image sets with significantly higher diversity span than plain prompt engineering (span 0.65 vs 0.45, p<.0001) and matched precise target label distributions more closely than automatic diversification baselines. If true, this shifts the burden from the generator to the interaction: users can get the diversity they want without retraining models.","feed_headline":"Users who dial in image attributes get 44 percent more diversity","feed_subtitle":"Varif.ai's verify-and-vary loop beat prompt-only in a 20-person study and hit target label mixes better than automatic baselines.","key_machinery":"The central mechanism is probabilistic prompt generation driven by user-editable attribute histograms. For each attribute, a large language model proposes labels; the user adjusts each label's weight; Varif.ai samples labels according to those weights and appends them to the base prompt; the diffusion model then generates images whose attribute distribution approximates the histogram. Verification is done by classifying each image with CLIP against the attribute labels and counting matches into the histogram, closing the loop. The work this mechanism does is to make diversity a measurable, steerable quantity rather than a side effect of prompt wording.","core_discovery":"The central claim is that an interactive verify-and-vary loop, not the underlying diffusion model, is what lets users achieve their desired image diversity. On the paper's own terms: Varif.ai enables users to define attributes, see their current label distribution as a histogram, adjust sliders to set target proportions, and generate a new batch from probabilistically sampled prompts; this produces image sets with higher CLIP diversity span than prompt-only generation (M=0.65 vs 0.45, p<.0001) and higher diversity alignment when a precise target distribution is given (alignment 0.79 vs 0.73 open-ended, p<.001). The authors argue this shows user-driven diversity control is both feasible and more aligned with user aims than automatic diversification.","pith_inferences":["One implication the authors leave implicit is that the tool's power is bounded by the classifier: if CLIP's notion of an attribute diverges from the user's, the histograms can mislead, so a natural extension is to let users confirm or correct labels rather than trusting CLIP silently.","The same verify-and-vary loop could transfer to other generative media such as text, code, or 3D scenes, since it only manipulates prompts and label counts; the paper hints at text generation but does not test it.","A testable extension is whether setting a label to zero percent actually suppresses the concept, since the paper notes participants wanted to blacklist labels; this could be checked by prompting with zero-weight labels and measuring occurrence in generated images.","The dependency on CLIP suggests a concrete benchmark: measure diversity gains under attributes CLIP is known to miss, such as fine-grained object-specific styles, and if the gains vanish, the interface is only as good as the classifier."],"forward_implications":["If the central claim holds, image-generation tools can offer diversity control without retraining or fine-tuning the generative model; any text-to-image model can be steered this way.","Users can satisfy fairness-style requirements, such as balanced ethnicity in doctor images, by setting target proportions, something the paper's comparisons suggest automatic diversification handles worse.","The measured engagement gain (10.2 vs 6.3 minutes on task) indicates users persist longer when they can see and manipulate attribute distributions, which may matter for creative ideation workflows.","The sensitivity analysis implies that even when CLIP classification is noisy, the generated image sets remain diverse by manual-label measures, so the vary step is robust to verification error."],"supporting_citations":[{"why":"Supplies the CLIP image-text model used to classify images into attribute labels during verification and to compute the diversity span metric in evaluation.","marker":"[55]"},{"why":"Provides the SD-XL Lightning text-to-image diffusion model that generates all image sets in Varif.ai and the user studies.","marker":"[40]"},{"why":"Supplies the open-weight LLaMA-2 model that proposes attribute labels and attribute suggestions to the user.","marker":"[66]"},{"why":"Serves as the automatic prompt-optimization baseline (Promptist) against which Varif.ai is compared for open-ended diversity.","marker":"[28]"},{"why":"Serves as the GPT-4o automatic diversified-prompt baseline compared in the summative study.","marker":"[22]"},{"why":"Establishes the diversity span metric that the paper uses to quantify image-set coverage and diversity.","marker":"[11]"},{"why":"Also cited for the span metric, connecting it to earlier directed-diversity work in ideation.","marker":"[16]"}],"fun_headline_variants":["User-set diversity targets beat auto modes in image gen","Dial-in diversity: verify-and-vary loop beats prompt-only","User-driven image diversity: 20-person study shows higher coverage","Varif.ai: let users set diversity goals, get better image sets","Interactive verify-and-vary loop yields user-aligned image diversity"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The results assume CLIP embeddings capture the diversity attributes users actually care about; if CLIP misreads or entangles attributes, the measured span and alignment gains may not reflect genuine user-perceived diversity.","fun_headline_variants_meta":{"raw":{"variants":["User-set diversity targets beat auto modes in image gen","Dial-in diversity: verify-and-vary loop beats prompt-only","User-driven image diversity: 20-person study shows higher coverage","Varif.ai: let users set diversity goals, get better image sets","Interactive verify-and-vary loop yields user-aligned image diversity"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00044,"raw_usage":{"total_tokens":2202,"prompt_tokens":882,"completion_tokens":1320,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":498,"completion_tokens_details":{"reasoning_tokens":1234}},"tokens_in":498,"tokens_out":1320,"duration_ms":9655,"temperature":1.0,"reasoning_tokens":1234,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T18:29:26.407845+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take Varif.ai and a prompt-only baseline, generate image sets for attributes where CLIP is known to misclassify (for instance, object-specific colors such as frog color versus background color), and have users or manual annotators judge diversity. If manual-label diversity shows no advantage for Varif.ai over prompt-only when CLIP accuracy is below roughly 0.6, the user-driven diversity claim fails. A simpler ablation would remove the verification histograms and let users vary distributions blindly; if the diversity span advantage persists, the verify step is not load-bearing.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes the diversity span metric that the paper uses to quantify image-set coverage and diversity."}],"review_version":1}