{"id":"784cdfde-592e-4828-8604-7b4eb8d7e478","arxiv_id":"2412.03889","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"A mesh deformation system that jointly optimizes semantic alignment with text or image prompts and body fit, producing body-aware 3D objects.","lead":"The authors present an optimization method that deforms a starting 3D mesh so the result matches a text or image description while fitting a particular body shape. The system is a tool for designing personalized accessories such as glasses, hats, rings, and slippers.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Text-guidance results are the load-bearing gap: Fig. 8 admits text guidance 'exhibits limited deformation on most examples,' yet Table 1 reports no quantitative text-conditioned variant, while contact vertices Vc in Eq. (8) are left unspecified.","rationale":"Reading in good faith, the paper has a plausible pipeline: Jacobian-based deformation with CLIP semantics and body losses, plus fabricated objects and image-guidance comparisons that give some real support. The reader's weakest assumption about contact vertices is legitimate: Vc is an input to Eq. (8) but its selection procedure is absent, so body-fit results are not reproducible. However, I would elevate the text-guidance gap as the more load-bearing weakness: the central claim explicitly promises text or image guidance, and §4.2 admits text guidance shows limited deformation on most examples, while Table 1 omits the text-conditioned row entirely. Because the reported CLIP score is the very objective being optimized, it is not an independent measure of semantic alignment; the user study (N=9) is too small to carry that weight. The verdict should remain conditional rather than reject, since the image-guided pipeline and physical fabrication demonstrate partial validity. The added condition should be a quantitative evaluation of the text-conditioned variant and a specified/released protocol for choosing Vc.","tokens_in":11469,"tokens_out":7386,"duration_ms":70865,"concrete_test":"Re-run the Table 1 protocol on the text-conditioned variant for the same prompts as Fig. 8 (e.g., 'star ring', 'cat mask') using the template mesh, reporting CLIP score, Dp, and Dc; if the text variant is not significantly better than the template baseline on most prompts, the abstract's 'text or image' claim is not supported for text guidance.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"To support the central claim, both text and image guidance must reliably produce body-aware, semantically aligned objects. The paper's own §4.2 and Fig. 8 state that text guidance 'exhibits limited deformation on most examples (e.g., \"star ring\")' and that image guidance provides a stronger deformation signal. Despite this, Table 1 contains no Template+Text row; the only quantitative semantic/body metrics are for template, guidance-mesh, and image-guided variants. The text modality, which is prominent in the abstract and Figs. 2–3, is therefore supported only by selected qualitative examples. In addition, the body-aware loss in Eq. (8) requires a set of contact vertices Vc, but the paper never explains how Vc is obtained for a new object-body pair or whether it is manually placed per example; the claimed body-fit results are not reproducible without this unstated input. The user study (N=9) is too small and incompletely described to substitute for a quantitative text-conditioned comparison. These two gaps jointly undermine the claim that the method synthesizes body-aware objects from text or image without manual intervention.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents ShapeCraft, a mesh-deformation method that optimizes per-triangle Jacobian fields to deform a template object mesh into a body-aware, semantically aligned 3D object, guided by text, image, or sketch. The optimization combines CLIP-based semantic alignment losses (Eqs. 3 and 5) with contact and penetration losses (Eqs. 8–10) and a Jacobian regularizer. The authors show qualitative results across object categories and body shapes, a small quantitative comparison in Table 1, a user study in Table 2, and applications including 3D printing and sketch-guided design.","tokens_in":11737,"tokens_out":4474,"duration_ms":39230,"significance":"The core idea—jointly optimizing semantic alignment and body contact/penetration in a differentiable mesh-deformation framework—is timely and potentially useful for personalized object design. Strengths include the breadth of qualitative demonstrations, the use of a Jacobian-field parameterization, and the fabrication and sketch applications. However, the paper's claims are broader than the evidence: the main quantitative metrics coincide with the optimized losses, and the text-guidance modality is not quantitatively evaluated despite being a headline contribution.","major_comments":[{"comment":"The reported metrics are the very objectives being optimized: CLIP cosine similarity is minimized in Eq. (3) for text guidance and in Eq. (5) for image guidance, Dp is derived from the signed distances penalized in Eq. (9), and Dc is a Chamfer-like distance over the same contact vertices used in Eq. (8). Consequently, the table cannot establish that ShapeCraft outperforms baselines on independent notions of semantic alignment or body fit. Please add independent metrics not used in the optimization, report variance or error bars, and give per-example counts.","section":"§4.3, Table 1"},{"comment":"Fig. 8 states that text guidance \"exhibits limited deformation on most examples (e.g., 'star ring')\", yet Table 1 contains no Template+Text row and the user study in Table 2 does not separate text-conditioned from image-conditioned outputs. Because the abstract and introduction claim generation \"from text, image, or sketch\", the text modality needs its own quantitative evaluation, including failure cases; without it, the central claim is supported only by selected qualitative examples.","section":"§4.2 and Table 1"},{"comment":"The contact loss requires a set of contact vertices Vc, described in Fig. 4 as an input, but the paper never explains how Vc is obtained—whether manually annotated per example, derived automatically from the body mesh, or optimized. This matters because incorrect contact vertices will pull the object to wrong locations, and the claim of \"no manual artist intervention\" depends on this step being automatic. Please specify the procedure and, ideally, ablate sensitivity to Vc.","section":"§3.2, Eq. (8)"},{"comment":"The user study uses only N=9 participants and reports means without variance, confidence intervals, or statistical tests. For a subjective comparison over methods, this is insufficient to support the claim that ShapeCraft is \"most prompt aligned and aesthetic\" while maintaining comfort. At minimum, report per-participant distributions and a paired significance test, or treat the user study as a pilot.","section":"§4.3, Table 2"}],"minor_comments":[{"comment":"The weight λc appears both inside Lc in Eq. (8) and as a multiplier in Eq. (10), so the effective contact weight is λc^2; please clarify which definition is intended.","section":"§3.2, Eqs. (8) and (10)"},{"comment":"The caption of Table 1 does not define Dp and Dc; in particular Dc is called \"chamfer distance of the contact points\" but no formula or sampling details are given.","section":"§4.3, Table 1"},{"comment":"The caption and the first-row label (\"NoWith\") appear truncated; the penetration-map color scale is described but the displayed axes and legend are missing.","section":"Figure 10"},{"comment":"Reference [61] has garbled text (\"Ergoboss: onomic ptimization of dy-upporting urfaces\"); please fix the title and author formatting.","section":"References"},{"comment":"The statement that different body shapes affect \"creativity\" is subjective; the supporting discussion is qualitative and would benefit from a quantitative measure of deformation or prompt alignment per body shape.","section":"§4.1"}],"recommendation":"major_revision","confidential_remarks":"This is a demo-oriented systems paper; the core pipeline is plausible and many qualitative results are compelling, but the evaluation is not yet at the level needed to support the strong claims. The main gaps are addressable within a revision: independent metrics, text-conditioned quantitative results, and a clear specification of Vc. The text-guidance limitation is the most serious; if the method genuinely fails on most text prompts, the paper's framing should be revised accordingly."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague —\n\nThis is a system paper for generating body-aware 3D objects from a template mesh plus text or image guidance. The genuinely new piece is the joint optimization: it combines CLIP-based semantic losses, a Jacobian-field deformation parameterization, and contact/penetration losses into one objective, then shows it working across object categories (glasses, rings, masks, slippers) and body shapes, including 3D-printed results. The image-guidance pipeline, with a guidance mesh from an off-the-shelf image-to-3D model plus Chamfer and L1 losses, is a sensible reuse of existing pieces. The authors also directly compare against two-stage body refinement, which is the right baseline, and they show penetration maps rather than only final renders. I found the qualitative evidence moderately convincing for image guidance and for the effect of the body-aware losses.\n\nThe soft spots are real and concentrate in the evaluation. Table 1 reports exactly the objectives being optimized: CLIP cosine similarity, penetration distance, and contact chamfer distance. There are no error bars, no per-example counts, and no independent metric. So the quantitative claim that the method \"achieves the best results\" rests on circular measurements and a user study of nine participants. The bigger issue is text guidance. Figure 8 admits text guidance \"exhibits limited deformation on most examples,\" yet Table 1 has no Template+Text row at all. The abstract and teaser lead with text input; without a quantitative text-conditioned comparison, the central claim is half-supported. The paper also never says where the contact vertices Vc in Eq. (8) come from. Since the body-aware loss is defined on them, this is a load-bearing unspecified input. If they are hand-placed per example, the method is less autonomous than claimed; if they are derived automatically, that derivation should be described. I would also ask for code and implementation hyperparameters; the current submission lacks reproducibility details.\n\nI disagree with the stress-test on one minor point: it suggests text guidance is the load-bearing gap, which is right, but the paper itself is honest about it in §4.2. That honesty does not fix the missing table row, though.\n\nWho should read this: researchers in 3D content creation and human-object interaction, especially those building on TextDeformer or scene-synthesis contact losses. It is a reasonable workshop-to-conference system paper, not a paradigm shift. I would send it to peer review, but only with major revisions: add text-conditioned quantitative results, non-circular evaluation, specify Vc, and release code. With those, it could be a solid conference contribution.\n\nRecommendation: give it a serious referee; expect heavy revision.","headline":"Useful joint-optimization system paper, but the text-guidance headline is under-supported by circular metrics and an unspecified contact-point input.","tokens_in":12244,"tokens_out":2955,"would_cite":false,"duration_ms":27705,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"ShapeCraft deforms a template mesh into a body-fitting, semantically aligned 3D object from text, image, or sketch.","keywords":["body-aware 3D generation","mesh deformation","CLIP guidance","contact and penetration optimization","text-to-3D","image-to-3D","Jacobian fields","3D printing"],"falsifier":"Run the optimizer with two different contact-vertex sets for the same object-body pair, one placed on the body region where the object should sit and one placed far away; if the final objects do not differ meaningfully or both fail to fit, the contact loss is not doing the claimed work. Alternatively, swap in randomly chosen contact vertices and measure contact distance and penetration against the paper's reported values.","tokens_in":1397,"feed_emoji":"🖨️","tokens_out":2120,"duration_ms":62461,"temperature":0.7,"pith_summary":"The paper claims that a single joint optimization over per-triangle Jacobian deformations can make everyday 3D objects simultaneously match a text, image, or sketch prompt and fit a given body shape. It treats semantic alignment as CLIP embedding similarity of rendered views and body fit as contact plus penetration losses. If true, this gives a data-free design tool for glasses, hats, rings, shoes, and other wearables, adjustable to different bodies and fabricable by 3D printing. The method is demonstrated across categories and body shapes with both objective metrics and a user study.","feed_headline":"One optimizer turns template meshes into body-fitting 3D designs","feed_subtitle":"ShapeCraft deforms a base mesh to match text, image, or sketch while keeping contact and penetration losses low.","key_machinery":"The central object is the per-triangle Jacobian field $J_i$ of the deformed mesh, optimized through a Poisson solve rather than raw vertex positions, following Neural Jacobian Fields. Semantic alignment is driven by CLIP cosine similarity between differentiable renders and the prompt or guidance-image embeddings, plus a patch-level feature regularizer for multi-view consistency. Body fit is driven by the contact loss $L_c(V, V_c) = \\lambda_c \\frac{1}{|V_c|} \\sum_{v_c \\in V_c} \\min_{v \\in V} \\|v_c - v\\|_2^2$ and the penetration loss $L_p = \\sum_{d_i < D} d_i^2$, where $d_i$ are signed distances between the object and body mesh.","core_discovery":"ShapeCraft's central claim is that optimizing per-face Jacobian matrices of a template mesh against a weighted sum of semantic and body-aware losses suffices to produce creative, functional objects rigged to a given body, without per-object datasets or manual artist intervention. Starting from one mesh per category, the optimizer deforms it toward text, image, or sketch guidance while holding selected body contact vertices close and keeping penetration below a threshold. The resulting meshes are watertight enough to simulate on virtual characters and to fabricate in the real world.","pith_inferences":["A natural extension the paper leaves implicit is to predict the contact-vertex set automatically from the prompt and body shape, since the current pipeline delegates that choice to the user.","If the contact weight $\\lambda_c$ is swept upward, semantic creativity should shrink as the object is pulled harder onto the body; this trade-off curve is a testable extension of the paper's results.","Because the method inherits CLIP's text-image associations, designs may skew toward stereotyped renderings of prompts, and the sketch-guided path via image guidance is one way around that limitation."],"forward_implications":["A single template mesh per object category can be reshaped into many prompt-specified designs, such as a star ring, cat mask, or cow hat, rather than requiring a separate generative model per object.","Image guidance gives stronger semantic control than text guidance alone, while text guidance keeps the deformation fluid and requires no external 3D reference.","Joint optimization beats two-stage pipelines: refining a guidance mesh after the fact cannot fix wrong topology or overly thin structures, whereas starting from a template and optimizing both objectives can.","Generated objects can be 3D printed and worn or simulated on virtual characters, and the same prompt adapts to different body shapes, from adults to cartoon characters.","In the user study, the joint method scores highest on prompt alignment and aesthetics while maintaining comparable comfort to the untouched template mesh."],"supporting_citations":[{"why":"Supplies the per-triangle Jacobian parameterization and Poisson solve used to deform the mesh while avoiding vertex-level artifacts.","marker":"[2]"},{"why":"Supplies the CLIP-based text-guided mesh deformation approach and the patch-level feature regularizer for multi-view consistency.","marker":"[15]"},{"why":"Supplies the CLIP embedding space used for both text and image semantic alignment losses.","marker":"[40]"},{"why":"Lifts input images to 3D guidance meshes for the image-guided deformation pipeline.","marker":"[1]"},{"why":"Motivates the contact loss formulation for body-aware optimization in human-scene interaction.","marker":"[54]"},{"why":"Provides the differentiable renderer that produces multi-view images of the mesh for the CLIP losses.","marker":"[23]"},{"why":"Converts user sketches to 2D images in the sketch-guided design application.","marker":"[60]"}],"fun_headline_variants":["One optimizer turns base meshes into body-aware 3D designs","Deform a single mesh to fit any body and prompt without artist effort","ShapeCraft: body-aware 3D objects from text, image, or sketch via mesh optimization","Optimizing mesh Jacobians for functional and body-fit 3D creation"],"cache_read_input_tokens":14464,"weakest_assumption_plain":"The load-bearing premise is that someone already knows which body vertices the object should touch; the paper takes the contact-vertex set as an input and never explains how to pick it, so a wrong choice would pull the object to the wrong spot and the claimed fit would fail.","fun_headline_variants_meta":{"raw":{"variants":["One optimizer turns base meshes into body-aware 3D designs","Deform a single mesh to fit any body and prompt without artist effort","ShapeCraft: body-aware 3D objects from text, image, or sketch via mesh optimization","Optimizing mesh Jacobians for functional and body-fit 3D creation"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000863,"raw_usage":{"total_tokens":3663,"prompt_tokens":786,"completion_tokens":2877,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":402,"completion_tokens_details":{"reasoning_tokens":2793}},"tokens_in":402,"tokens_out":2877,"duration_ms":22003,"temperature":1.0,"reasoning_tokens":2793,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T21:57:59.831838+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the optimizer with two different contact-vertex sets for the same object-body pair, one placed on the body region where the object should sit and one placed far away; if the final objects do not differ meaningfully or both fail to fit, the contact loss is not doing the claimed work. Alternatively, swap in randomly chosen contact vertices and measure contact distance and penetration against the paper's reported values.","supporting_citations":[{"cited_title":"Kim, Siddhartha Chaudhuri, Jun Saito, and Thibault Groueix","cited_arxiv_id":null,"evidence_quote":"Supplies the per-triangle Jacobian parameterization and Poisson solve used to deform the mesh while avoiding vertex-level artifacts."},{"cited_title":"Textdeformer: Geometry manipu- lation using text guidance","cited_arxiv_id":null,"evidence_quote":"Supplies the CLIP-based text-guided mesh deformation approach and the patch-level feature regularizer for multi-view consistency."},{"cited_title":"Tripo ai","cited_arxiv_id":null,"evidence_quote":"Lifts input images to 3D guidance meshes for the image-guided deformation pipeline."},{"cited_title":"Scene synthesis from hu- man motion","cited_arxiv_id":null,"evidence_quote":"Motivates the contact loss formulation for body-aware optimization in human-scene interaction."},{"cited_title":"Adding conditional control to text-to-image diffusion models","cited_arxiv_id":null,"evidence_quote":"Converts user sketches to 2D images in the sketch-guided design application."}],"review_version":1}