{"id":"3d66a453-36a4-4cbb-b6a5-0e65e63e10c7","arxiv_id":"2508.07135","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Canvas3D lets users arrange objects in a 3D canvas generated from a text prompt, then feeds depth, skeleton, and lighting constraints to diffusion models to produce images that match the layout.","lead":"This paper presents Canvas3D, a system that turns a user's text prompt into a 3D scene they can rearrange, then uses that arrangement as spatial constraints for image generation. The result matters for anyone who wants generated images to respect an intended layout, not just a verbal description.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Baseline confound: slider condition also changes visual representation and depth-condition fidelity, so the significant spatial metrics do not isolate the proposed 3D-manipulation mechanism.","rationale":"The reader's stated weakest assumption — that depth conditions faithfully encode the 3D arrangement — is a genuine limitation, but it is documented and affects both Canvas3D and the baseline (both feed Uni-Con). The more load-bearing concern is the baseline confound: the paper compares a full pipeline (automatic object registration, mesh rendering, mesh-based spatial conditions, direct 3D manipulation) against a minimal slider/bounding-box baseline. Because multiple variables change simultaneously, the significant differences in Uni-Det, GPT-CLIP, and Recall cannot be assigned to the proposed interaction design itself. This does not refute the system-level claim that Canvas3D outperforms the baseline, but it does undercut the paper's central framing that direct 3D manipulation is what empowers precise spatial control. The reader's rationale already mentions baseline confounding, so the conditional verdict is appropriate; my concern sharpens the specific mechanism that needs to be isolated. No code/data release, a single target image, and uncorrected multiple comparisons further support keeping the verdict conditional rather than accept. I therefore recommend no change to the reader's conditional verdict, but I would ask for the three-arm test (or at least a bounding-box-with-mesh-depth control) before treating the interaction claim as established.","tokens_in":27469,"tokens_out":8435,"duration_ms":88401,"concrete_test":"Run a three-arm within-subject experiment using the same target image and metrics: (A) Canvas3D as shipped; (B) a slider-based interface that renders actual 3D meshes and uses Canvas3D's exact spatial-condition encoder (mesh-derived depth, skeleton, scene image); (C) the original bounding-box slider baseline. If B ≈ A > C, the advantage is driven by condition fidelity / visual representation, not direct manipulation. If A > B ≈ C, the 3D-manipulation interaction itself is supported. Release the code/interface and conditions so the three arms can be independently inspected and reproduced.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The closed-ended comparison cannot attribute the observed gains to the proposed 3D-manipulation interaction. In §5.1.1 the baseline is described as sliders for object position/orientation, and Figures 10/15 show its composition preview uses bounding boxes, whereas Canvas3D shows textured 3D meshes and renders mesh-based depth, skeleton, and scene conditions (§4.5). The two conditions therefore differ in at least three ways at once: input modality (sliders vs direct manipulation), visual feedback (boxes vs meshes), and the fidelity of the spatial condition fed to the same Uni-Con backbone. The objective metrics (GPT-CLIP, Uni-Det, Recall) could improve simply because mesh-based depth encodes object shape and occlusion better than box-based depth — indeed, §6.1.3 reports that nearby bounding boxes are 'frequently misinterpreted as a single object.' This means the study underdetermines the central HCI claim: even if Canvas3D is a better full system, the advantage may have nothing to do with direct 3D manipulation. The authors' own depth-resolution limitation in §7.3 is real but affects both systems and is not the primary threat to attribution.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents Canvas3D, an interactive system for image generation with explicit spatial control. A user enters a text prompt; the system registers 3D objects from ShapeNet/Objaverse, synthesizes an initial scene with an LLM, and maps mouse/keyboard input to object affordances in a Unity-based canvas. The user can rearrange objects, adjust camera and lighting, and configure human posture. The system encodes the resulting arrangement into spatial conditions (depth, skeleton, lighting JSON, etc.) that are passed to conditional generative models (Uni-Con, IC-Light). The paper reports a within-subject study (n=12) comparing Canvas3D to a slider-based baseline with the same generative backbone, finding significant advantages on GPT-CLIP, Uni-Det, Recall, and several NASA-TLX/Likert measures, plus an open-ended usability session with SUS score 82.22. The central claim is that direct 3D manipulation gives users significantly better spatial control than slider-based control.","tokens_in":27668,"tokens_out":7137,"duration_ms":69625,"significance":"If the result holds, Canvas3D is a valuable contribution to controllable generation: it provides an end-to-end pipeline from prompt to manipulable 3D scene to spatial conditions, and the interaction design aligns with natural manipulation. The system is implemented and documented in enough detail to be reproduced, and the use of a within-subject design with counterbalancing and qualitative interviews is appropriate. However, the current evaluation cannot uniquely attribute the observed gains to the proposed manipulation mechanism because the baseline also differs in visual representation and condition fidelity. The paper would need a controlled comparison or additional experiment to support its central claim.","major_comments":[{"comment":"The comparison in §5.1.1/Figures 10,15 does not isolate the proposed 3D-manipulation interaction. The baseline differs from Canvas3D in at least three ways: input modality (sliders vs direct manipulation), visual feedback (bounding boxes vs textured meshes), and the fidelity of the spatial condition (box-derived depth vs mesh-derived depth) fed to the same Uni-Con backbone. §6.1.3 itself attributes baseline failures to bounding boxes being 'frequently misinterpreted as a single object.' Thus the significant gains in GPT-CLIP, Uni-Det, and Recall (Table 1) could arise from the condition-encoding difference alone. To support the central HCI claim, the study needs a control that holds the condition representation fixed (e.g., sliders with mesh-based depth) or adds a third condition isolating each factor.","section":"§5.1.1, Figures 10/15"},{"comment":"The closed-ended study uses a single target image (Fig. 18) with one object set and one spatial layout; Table 3 aggregates counts/times for that stimulus. With n=12 and one stimulus, the claim that Canvas3D 'consistently outperforms' the baseline does not generalize across object categories, scene complexity, or spatial arrangements. Additional target scenes (or at least a per-stimulus analysis and a clear acknowledgment of this scope limit) are needed before drawing general conclusions about spatial controllability.","section":"§5.1.2, Figure 18"},{"comment":"Table 1 reports five objective metrics without correction for multiple comparisons, and Fig. 13 adds many subjective tests; the smallest p-values would survive Bonferroni, but the authors should report adjusted p-values or FDR and include effect sizes/confidence intervals so readers can judge magnitudes. In addition, GPT-CLIP and GPT Spatial rely on GPT-generated captions/judgments with no reported reliability (e.g., agreement with human raters or repeatability). Because both conditions are evaluated with the same LLM judge, this is not circular, but it is a source of measurement uncertainty that should be quantified.","section":"§5.1.3, Table 1"},{"comment":"Section 7.3 documents that the depth condition can lose spatial distinctions when objects are close in depth (P11 quote). This is an acknowledged limitation, but the discussion does not connect it to the closed-ended comparison. Since the two conditions use different depth encoders (mesh vs boxes), this failure mode may affect the condition-fidelity confound differently across conditions, and it also bounds the central 'precise spatial control' claim. Please discuss how this limitation interacts with the objective metrics and whether the system-level advantages persist when depth resolution is the bottleneck.","section":"§7.3"}],"minor_comments":[{"comment":"Figure 2 contains an untranslated editing note ('放citation'), and Figures 4 and 9 contain Chinese annotation text ('字加大', '字体加粗加大'). These are leftover author annotations and must be removed.","section":"Figures 2/4/9"},{"comment":"Sections 6.2.1 and 6.2.2 are both titled 'System Usability Questionnaire'; the second appears to be the System Feature Questionnaire. There are also typos: 'Metrice' (§5.1.3), 'Geneartive' (§2.2 heading), 'faciliate', 'perprndicular', and 'wildly'.","section":"§6.2.1/6.2.2"},{"comment":"Table 3 reports time-to-first-liked and liked-ratio rows without p-values or confidence intervals. If these are exploratory, say so explicitly; otherwise provide the corresponding tests.","section":"Table 3"},{"comment":"The Uni-Det score is defined by listing five spatial relationships, but the exact formula for comparing positions/depths of detected boxes is not given (thresholds, normalization, per-relationship scoring). As written, the metric is not fully reproducible.","section":"Appendix A.6"}],"recommendation":"major_revision","confidential_remarks":"I agree with the stress-test concern: the baseline confound is the key issue, and it is load-bearing for the central HCI claim. The manuscript is otherwise technically sound and the system is a useful contribution, but the current evaluation underdetermines the mechanism. The leftover Chinese editing annotations in figures suggest the manuscript is not fully submission-ready. I recommend major revision with a controlled follow-up experiment or a careful reanalysis that separates input modality from condition fidelity."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a real systems paper, not a repackaging. The Canvas3D pipeline—prompt-to-3D scene, direct mesh manipulation, automatic encoding into depth/skeleton/lighting conditions for a common diffusion backbone—is a genuine integration I haven't seen as a single workflow. The authors also did real user work: 12 participants, within-subject, counterbalanced, plus an open-ended session, and several significant differences favoring their system.\n\nWhat's best: the system is sensibly decomposed (object registration, scene synthesis via LLM, interaction decision tree, condition encoding) and the UI is described concretely. The qualitative quotes are plausible and support the usability story.\n\nWhere it's soft: the main problem is the stress-test concern, and I think it lands. The baseline changes three things at once—sliders vs direct manipulation, bounding boxes vs textured meshes, and likely the fidelity of the depth condition fed to the same backbone. The authors even observe that nearby bounding boxes are misinterpreted as a single object (§6.1.3), which means part of the objective advantage may be due to better depth encoding, not to the interaction style. So the paper's core HCI claim—that direct 3D manipulation is what helps—is not isolated. The paper would still work as an 'integrated system beats an ablated slider version' claim, but the title and contributions overreach.\n\nOther issues: only one target image in the closed-ended study, no multiple-comparison correction, GPT-based metrics without human validation, and no code/data/demo release. There are also unfinished template placeholders (conference acronym, 'Make sure to enter the correct conference title') that suggest this is a draft, not camera-ready. None of these are fatal; they're addressable. The depth-resolution limitation in §7.3 is real but affects both systems and isn't the main threat.\n\nBottom line: I'd send this to peer review. It's a serious contribution with a coherent system and an honest (if underpowered) evaluation. I'd ask for a cleaner baseline, more target images, corrected statistics, and release of at least the pipeline code. For a reading group, it's a good trigger for discussing what counts as a fair baseline in HCI+GenAI studies.","headline":"A well-built HCI system with a genuine integration, but the evaluation does not isolate the direct-manipulation claim.","tokens_in":28201,"tokens_out":2268,"would_cite":true,"duration_ms":23299,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Canvas3D argues that direct 3D object manipulation gives users precise spatial control over generated images, and reports it beating slider-based controls on every measured metric.","keywords":["Controllable Image Generation","Spatial Control","Conditional Generative Models","3D Interaction","Interactive Canvas","Depth Conditioning","Text-to-Image Generation","User Study"],"falsifier":"Place two objects very close together in depth on the Canvas3D canvas under the same prompt, and check whether the generated image keeps them as distinct objects; the paper reports this failing for a car near a house. More broadly, rerunning the closed-ended comparison with a larger sample and more scenes would settle whether the Uni-Det and Recall advantages persist beyond the 12 participants.","tokens_in":27320,"feed_emoji":"🎨","tokens_out":8437,"duration_ms":81390,"temperature":0.7,"pith_summary":"The paper tries to establish that the bottleneck in controllable image generation is not the generative model but the input: text, sketches, and sliders are clumsy ways to express where objects should sit in 3D. Canvas3D replaces those with a virtual 3D canvas that a user prompt builds automatically, letting people drag, rotate, pose, and re-light objects as if arranging a real scene. The system then encodes the arrangement into depth images, skeleton maps, and lighting metadata that a conditional generative model can consume. In a within-subject comparison against a slider-based system using the same generation backbone, the paper reports Canvas3D winning on every objective spatial-alignment metric and on user-rated interactivity, control, effort, and frustration. A second open-ended study reports that the system works with everyday prompts and earns a strong usability score.","feed_headline":"Direct 3D manipulation beats sliders for spatial control of AI images","feed_subtitle":"Users who arranged objects in a virtual scene matched target layouts more accurately and with less effort.","key_machinery":"The load-bearing mechanism is the spatial-condition encoding pipeline: an automatically constructed 3D canvas plus a function library that exports the user's arrangement as depth images, scene screenshots, OpenPose-format skeletons, and lighting JSON, along with native mesh data. This encoding sits between the interaction layer and the generative model; it is what makes the user's mouse movements into constraints a conditional diffusion model can actually obey. The automatic object registration and scene synthesis from the prompt matter too, because they remove the setup burden and keep the comparison about spatial control rather than 3D modeling skill.","core_discovery":"Canvas3D's central claim is that direct 3D manipulation gives users genuinely precise spatial control over generated images, and that this precision survives the trip from user intent to final image. The authors argue that a 3D engine captures spatial intent intuitively because users manipulate actual objects rather than sliders or bounding boxes. The closed-ended study gave participants a target image and asked them to reproduce its spatial composition with Canvas3D or with a slider-based baseline; both used the same conditional generative backbone. Canvas3D outperformed the baseline on all five metrics, with significant advantages on GPT-CLIP (p=0.0024), Uni-Det (p=0.0034), and Recall (p=0","pith_inferences":["The documented near-depth failure suggests depth-only encoding is the weak link; combining depth with instance segmentation or an explicit relation graph would likely resolve cases where two objects merge, and this is a cheap test the paper does not run.","The comparison's outcome is tied to the slider baseline's interaction design; a head-to-head against sketch-based or drag-based controllers would clarify whether the advantage comes from 3D direct manipulation or simply from not using sliders.","The same prompt-to-canvas workflow could plausibly steer non-image generative tasks, such as 3D model generation or embodied-agent instructions, since the system already exports native 3D meshes and scene metadata; the paper only sketches those uses.","The reported effect sizes come from 12 participants; a larger replication varying scenes and user backgrounds would show how far the advantage generalizes."],"forward_implications":["Users can specify object placement, orientation, human posture, camera viewpoint, and lighting by arranging a 3D scene, then regenerate while keeping the same spatial constraints.","Spatial-alignment metrics and user ratings both improve relative to slider-based control when the generation backbone is held fixed.","The system lowers the skill barrier: no sketching ability or slider calibration is needed, since the canvas is created automatically from a text prompt.","The same encoded conditions (depth, skeleton, lighting) can be retargeted to other conditional generative models through the extensible encoder library.","Because the canvas enforces physical constraints, common scene violations such as floating or intersecting objects are reduced before generation."],"supporting_citations":[{"why":"Provides the conditional generation backbone used by both Canvas3D and the baseline, so the comparison isolates interaction style.","marker":"[51]"},{"why":"Represents the slider-based depth-conditioning approach that motivates the baseline interaction design.","marker":"[5]"},{"why":"The other slider-based interactive layout system the baseline is modeled on.","marker":"[23]"},{"why":"Supplies 3D models for object registration in the interactive canvas.","marker":"[17]"},{"why":"Supplies 3D models across most object categories in the curated dataset.","marker":"[10]"},{"why":"Computes the semantic similarity used to retrieve the best 3D model per category.","marker":"[78]"},{"why":"The reasoning LLM that infers object categories, counts, and implicit dependencies from the user prompt.","marker":"[68]"},{"why":"Establishes depth-map conditioning as an effective spatial constraint for diffusion models.","marker":"[108]"},{"why":"Provides the detector used to compute the Uni-Det spatial-relationship metric.","marker":"[115]"},{"why":"Supplies the illumination-conditioned model used for lighting control in the image generation layer.","marker":"[109]"}],"fun_headline_variants":["3D scene tweaks beat sliders for precise AI image layouts","Canvas3D: direct 3D control sharpens spatial image composition","Users match target layouts better with 3D canvas than sliders","Drag objects in 3D, not sliders, for accurate AI image scenes"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The pipeline assumes the encoded spatial conditions, especially the depth image, preserve the user's 3D arrangement intact; the paper's own study shows that objects close in depth can be conflated, and that chaining separate models for pose and lighting adds style inconsistency.","fun_headline_variants_meta":{"raw":{"variants":["3D scene tweaks beat sliders for precise AI image layouts","Canvas3D: direct 3D control sharpens spatial image composition","Users match target layouts better with 3D canvas than sliders","Drag objects in 3D, not sliders, for accurate AI image scenes"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000103,"raw_usage":{"total_tokens":834,"prompt_tokens":681,"completion_tokens":153,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":425,"completion_tokens_details":{"reasoning_tokens":74}},"tokens_in":425,"tokens_out":153,"duration_ms":2272,"temperature":1.0,"reasoning_tokens":74,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T22:18:14.443918+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Place two objects very close together in depth on the Canvas3D canvas under the same prompt, and check whether the generated image keeps them as distinct objects; the paper reports this failing for a car near a house. More broadly, rerunning the closed-ended comparison with a larger sample and more scenes would settle whether the Uni-Det and Recall advantages persist beyond the 12 participants.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The reasoning LLM that infers object categories, counts, and implicit dependencies from the user prompt."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes depth-map conditioning as an effective spatial constraint for diffusion models."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the detector used to compute the Uni-Det spatial-relationship metric."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the illumination-conditioned model used for lighting control in the image generation layer."}],"review_version":1}