{"id":"082675a1-c24d-4976-b92b-c01cda04fe84","arxiv_id":"2508.06065","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A new interface lets image editors navigate themes as a plane rather than prompt text, with an exploratory six-person study suggesting creative flow but weak predictability.","lead":"ThematicPlane is an image editing interface where users steer high level concepts like mood or style through a visual thematic plane instead of writing prompts. A six person exploratory study reports high satisfaction but also unpredictable outputs and a desire for more explainable controls.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The core 'bridge' claim rests on an unvalidated assumption that embedding cosine similarity defines meaningful thematic axes; the study's low outcome-expectation scores suggest the steering may not work as intended.","rationale":"The reader identified the same weakest assumption: that cosine similarity between DINOv2 image embeddings and text-encoder embeddings of LLM-extracted keywords defines meaningful visual directions. My stress test sharpens this into a concrete, falsifiable mechanism concern. The paper's own study data—low outcome-expectation ratings and participants' reports of unclear connections—are consistent with the mapping failing, which makes the central claim of 'bridging tacit user intent and system control' unproven. No baseline comparison or controlled evaluation of the semantic axis is reported, so the paper currently offers only qualitative self-reports. However, the paper is explicitly exploratory and names limitations, so a conditional accept is appropriate: the contribution is potentially valuable, but the core mechanism needs validation. The concrete test I propose would directly settle whether moving along the plane produces monotonic changes in the intended theme. If it fails, the central claim would need to be substantially weakened. I therefore do not change the reader's verdict of CONDITIONAL, but I emphasize that the condition is not merely a larger study—it is a technical validation of the semantic steering mechanism.","tokens_in":4850,"tokens_out":3085,"duration_ms":38616,"concrete_test":"Implement the exact pipeline from §1.1 on a fixed input image and a single theme axis (e.g., 'warmer' vs 'colder'). Sample 5 intermediate positions along that axis, regenerate with identical Imagen 3 settings, and (a) compute DINOv2 embedding cosine similarity of each output to a reference text embedding for 'warm' and check for monotonic increase; (b) have blind human raters rank the outputs by perceived warmth. If either monotonicity or human ranking fails, the plane does not steer by the stated semantic concept, and the 'bridge' claim is unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that users can manipulate high-level thematic concepts to steer image generation (Abstract, §1). For this to hold, moving along the ThematicPlane axes must produce predictable, monotonic changes in the intended visual theme. Section 1.1 describes the mechanism: GPT-4o extracts thematic keywords, 12 perturbations per theme are assigned left/right 'along a semantic axis,' DINOv2 image embeddings and text-encoder embeddings are 'ranked by cosine similarity,' and ranked descriptors are injected into Imagen 3 prompts. No evidence is provided that this embedding-similarity ranking corresponds to a perceptually meaningful direction. The paper's only quantitative indication about predictability is that participants did not anticipate outputs (M=3.43, SD=1.27, §3)—which is exactly what would be observed if the axis were an artifact of embedding similarity rather than a semantically coherent direction. Participants P2, P4, and P5 explicitly reported unclear connections between the plane and initial themes. Additionally, the Circumplex Model is adopted without validation for arbitrary image themes such as mood, style, or narrative tone. Because the entire contribution is the interface's ability to bridge tacit intent and system control, this unvalidated semantic mapping is load-bearing.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"ThematicPlane is an interactive image-generation tool that lets users manipulate high-level thematic attributes (mood, style, narrative tone) on a two-dimensional 'thematic design plane' rather than writing prompts. The system pipeline (Section 1.1) uses GPT-4o to extract thematic keywords, generates 12 perturbations per theme, assigns them to left/right semantic axes, ranks them by cosine similarity between DINOv2 image embeddings and text-encoder embeddings, and injects ranked descriptors into Imagen 3 prompts. The paper reports an exploratory study with N=6 participants across three conditions, presenting two self-report means (satisfaction M=5.86, SD=0.90; surprise/anticipation M=3.43, SD=1.27) and qualitative observations about divergent/convergent creative behavior. The central claim is that the interface 'bridges the gap between tacit creative intent and system control.'","tokens_in":5150,"tokens_out":4116,"duration_ms":46180,"significance":"If the embedding-based semantic axes were validated, ThematicPlane would be a useful contribution to creativity support tools, offering a low-externalization interaction style for non-experts. The paper is honest about limitations (single axis, differing user expectations, unexplainable mappings) and describes a concrete, reproducible system architecture. The qualitative observations about users treating unexpected outputs as inspiration are valuable. However, the central 'bridging' claim currently rests on an unvalidated technical assumption about embedding cosine similarity, and the evaluation is too small and too descriptive to support the general claim. The contribution's significance is therefore conditional on additional validation.","major_comments":[{"comment":"The load-bearing assumption that cosine similarity between DINOv2 image embeddings and text-encoder embeddings of GPT-4o-extracted keywords defines perceptually meaningful thematic directions is not validated. The paper gives no evidence that moving along the 'semantic axis' produces monotonic, recognizable changes in the intended theme. The study's quantitative result (M=3.43 for anticipating outputs, §3) and P2/P4/P5's reported unclear connections between the plane and initial themes are exactly what would be observed if the axis is an embedding artifact. I recommend a dedicated technical validation: independent raters judge whether outputs at opposite ends of an axis differ in the intended theme and whether intermediate points interpolate, with agreement statistics and example image pairs.","section":"§1.1"},{"comment":"The central claim that ThematicPlane 'bridges tacit intent' is not supported by the evaluation design. The study is N=6, reports no inferential statistics, no comparison across the three conditions, and no direct measure of whether outputs matched users' tacit intent. The quantitative evidence is two self-report means without scale anchors. To support the claim, the paper needs either a controlled comparison (e.g., task success, edit distance to target, number of iterations) or a qualitative analysis with inter-rater reliability. At minimum, the claims should be restricted to 'exploratory observations' rather than a demonstrated bridging of tacit intent.","section":"§2 and §3"},{"comment":"The Circumplex Model of affect [19] is adopted without justification for arbitrary image themes. Its dimensions (valence and arousal) do not obviously cover 'style' or 'narrative tone,' and the paper assumes a single left/right axis can represent each theme. The Discussion itself notes that the interaction provides only a single axis, which may restrict exploration. If the Circumplex is only a loose inspiration, the text should say so; if it is a design commitment, the paper should test whether theme perturbations actually fall along a bipolar dimension.","section":"§1.1 and §3"}],"minor_comments":[{"comment":"The title says 'Image Generation' while the abstract and body repeatedly say 'Image Editing'; please align the terminology.","section":"Title/Abstract"},{"comment":"Participant P7 is mentioned in the 'Different perceptions' paragraph, but the study reports N=6 participants. Please correct the numbering.","section":"§3"},{"comment":"The caption says 'thematic plan' while the text uses 'thematic plane'; unify the terminology.","section":"Figure 1"},{"comment":"The Likert scales for satisfaction and surprise (M=5.86 and M=3.43) are not defined; specify the range (e.g., 1–7) and what the endpoints represent.","section":"§3"}],"recommendation":"major_revision","confidential_remarks":"This is a 3-page UIST Adjunct paper. The central idea is interesting, but the validation is thin even for a short paper. The stress-test concern about embedding similarity being an unvalidated load-bearing mapping is directly supported by the participants' reported confusion. I would encourage the editor to request a technical validation of the semantic axes (or a clear framing as a speculative system) and to require the authors to temper the 'bridges tacit intent' language. With those changes, the paper could be suitable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a plausible little adjunct paper with a genuinely new interface concept and an honest exploratory study, but the central “bridging tacit intent” claim is not backed by the evidence, and the stress-test concern is real. The semantic axes are built by cosine similarity between DINOv2 image embeddings and text-encoder embeddings of LLM-extracted keywords; the paper never shows that moving along an axis produces a perceptually coherent or predictable theme change. That is load-bearing, and the participants’ own reports of unclear connections (P2, P4, P5) are consistent with the mapping being an artifact rather than a semantic direction.\n\nWhat’s new and good: ThematicPlane recombines existing parts — GPT-4o keyword extraction, DINOv2 text-aligned embeddings, Imagen 3, the Circumplex Model — into an interface concept that looks absent from the cited prior work (PromptMap, DeepSI, Concept Sliders). For a 3-page paper, the prototype is described clearly enough to reproduce. The study is small but the qualitative observations about divergent vs. convergent modes are plausible, and the authors are unusually honest about the explainability gap and the single-axis limitation.\n\nSoft spots: N=6 convenience sample, no reported comparison across the three conditions, no inferential statistics, and only self-report means for satisfaction and surprise. The abstract overreaches when it says the system “bridges the gap between tacit creative intent and system control” — the data support “users found it fun and sometimes useful,” not “the bridge works.” The surprise score (M=3.43) is hard to interpret because the scale anchors aren’t reported, so I wouldn’t lean on it either way. The Circumplex Model is adopted without validation for arbitrary image themes; that is a real gap but a minor one for an exploratory paper. The citation pattern is fine: LM-Steer appears only as future work, no self-citation inflation. No code or data are shipped, so the paper earns no reproducibility credit beyond the quoted interview excerpts.\n\nBottom line: the core idea is worth a serious look, but the claim-to-evidence ratio needs fixing. I would send it to peer review for an adjunct-level venue, expecting the authors to soften the language and, ideally, add even a small perceptual validation of the axes or a baseline comparison. This is not a full CHI paper as it stands. Readers who work on creative tools or semantic interaction will get value; ML folks won’t learn much.","headline":"Nice interface concept, honest exploratory study, but the load-bearing semantic-mapping assumption is unvalidated and the abstract overclaims.","tokens_in":5618,"tokens_out":3084,"would_cite":false,"duration_ms":33622,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"ThematicPlane claims that a thematic design plane—an interactive space of high-level concepts such as mood, style, and narrative tone—can let non-experts steer image generation without writing detailed prompts or hunting for reference image","keywords":["creativity support tool","visual exploration","generative AI","semantic interaction","latent space","thematic plane","text-to-image generation","circumplex model"],"falsifier":"Take a set of images and define one thematic axis (e.g., warm vs cold). Have users move the plane along that axis and rate whether the generated images become monotonically warmer, or compare the plane's direction against random keyword perturbations in a blinded test. If users cannot reliably tell the steered outputs from random ones, or if the semantic direction is not preserved, the core mechanism fails.","tokens_in":4777,"feed_emoji":"🎨","tokens_out":4299,"duration_ms":44396,"temperature":0.7,"pith_summary":"The paper introduces ThematicPlane, a system that lets users edit images by moving through a two-dimensional 'thematic plane' of high-level concepts instead of typing prompts. It argues this bridges the gap between tacit creative intent and system control: a user can push a point toward 'warmer' or 'more dramatic' and the image changes accordingly. The system extracts thematic keywords from the input image, ranks them by similarity between image and text embeddings, and feeds them to a text-to-image generator. An exploratory study of six participants suggests the approach supports iterative, expressive workflows, while revealing that users' mental models of theme-to-output mappings vary and would benefit from more explainable controls.","feed_headline":"Move a point to steer AI images by theme","feed_subtitle":"ThematicPlane turns tacit ideas like 'warmer' or 'more dramatic' into direct controls for image generation.","key_machinery":"The central object is the thematic design plane, an interactive two-dimensional space adapted from the Circumplex Model of affect. The plane maps a user's position to concrete prompt perturbations: thematic keywords ranked by cosine similarity between image embeddings and text-encoder embeddings, which then steer the text-to-image generator. This embedding-similarity ranking is the mechanism that turns a vague thematic direction into a generation command.","core_discovery":"The central claim is that a Circumplex-style thematic plane can act as a scaffolding layer between a user's tacit intent and the latent space of a generative image model. The plane is created by having a large language model extract thematic keywords from a starting image, removing object-level descriptors, and arranging twelve perturbations along an axis. Embeddings of the image and of each keyword are computed with a self-supervised vision model and its aligned text encoder, ranked by cosine similarity, and the top descriptors are injected into the prompt of a text-to-image model. Users can drag along the axis, regenerate, and set the new image as the reference for further edits. The paper","pith_inferences":["The same embedding-similarity steering mechanism could generalize to other generative modalities (video, audio, 3D) where a reference embedding and an aligned text encoder exist.","A testable extension is to replace the fixed Circumplex axes with data-driven or user-defined theme axes, which might resolve the mismatched mental models observed in the study.","The paper's assumption that moving along a semantic axis behaves linearly is probably false; editing multiple themes at once may yield nonlinear interactions, which could be measured and corrected."],"forward_implications":["If the plane works, non-experts can produce nuanced image edits by moving a point, without mastering prompt engineering or externalizing ideas.","Iterative regeneration with the new image as reference supports divergent exploration and convergent refinement in a single tool.","Because users' theme-to-output expectations varied, the paper implies that explainable controls—showing why a theme drew the image—are needed for practical deployment.","The single-axis interaction is a known limitation; the paper points toward multidimensional planes that combine mood and style simultaneously."],"supporting_citations":[{"why":"Supplies the Circumplex Model of affect, which structures the two-dimensional thematic plane.","marker":"[19]"},{"why":"Provides the self-supervised image embeddings used for similarity ranking between images and theme keywords.","marker":"[18]"},{"why":"Provides the aligned text encoder that produces embeddings for thematic keywords, enabling cross-modal similarity.","marker":"[15]"},{"why":"The text-to-image generation model that produces the edited images from prompt perturbations.","marker":"[3]"},{"why":"Grounds the semantic interaction paradigm that the thematic plane builds on, connecting tacit user intent to system controls.","marker":"[7]"},{"why":"Prior semantic interaction system that serves as a conceptual baseline for steering generative outputs through user-meaningful dimensions.","marker":"[4]"}],"fun_headline_variants":["Drag a theme to steer AI image generation","No prompts needed: a plane for implicit image intent","Click and drag on a mood map to create images","A point on a plane captures your creative vibe","ThematicPlane: guide AI art by feel, not words"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The system assumes that cosine similarity between image embeddings and text embeddings of LLM-extracted theme words really captures the direction a user means by 'warmer' or 'more dramatic,' and that the Circumplex Model's affect axes generalize to all visual themes.","fun_headline_variants_meta":{"raw":{"variants":["Drag a theme to steer AI image generation","No prompts needed: a plane for implicit image intent","Click and drag on a mood map to create images","A point on a plane captures your creative vibe","ThematicPlane: guide AI art by feel, not words"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000687,"raw_usage":{"total_tokens":2915,"prompt_tokens":671,"completion_tokens":2244,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":415,"completion_tokens_details":{"reasoning_tokens":2168}},"tokens_in":415,"tokens_out":2244,"duration_ms":15777,"temperature":1.0,"reasoning_tokens":2168,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T22:56:21.679186+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a set of images and define one thematic axis (e.g., warm vs cold). Have users move the plane along that axis and rate whether the generated images become monotonically warmer, or compare the plane's direction against random keyword perturbations in a blinded test. If users cannot reliably tell the steered outputs from random ones, or if the semantic direction is not preserved, the core mechanism fails.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the Circumplex Model of affect, which structures the two-dimensional thematic plane."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the aligned text encoder that produces embeddings for thematic keywords, enabling cross-modal similarity."}],"review_version":1}