{"id":"553be21a-168a-49f2-9bc5-0ae9f3a04f75","arxiv_id":"2501.01998","paper_version":2,"verdict":"REJECT","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":4,"one_line_summary":"SmartSpatial combines depth injection and attention guidance in Stable Diffusion with a new VLM-based spatial metric, but reported improvements are not statistically significant per the paper's own p-value statement.","lead":"SmartSpatial steers Stable Diffusion with depth maps and cross-attention guidance to place objects in 3D order, and adds a VLM-based scoring system for spatial accuracy. The paper claims strong gains, but its own significance tests say the differences are not statistically significant.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paper's own significance test, reported in §5.2 after Table 1, finds p>0.05 for all datasets, directly contradicting the Abstract's 'significantly outperforms'; the central claim lacks statistical support.","rationale":"The reader's weakest assumption focused on SmartSpatialEval's reliance on ChatGPT-4o descriptions and hand-assigned spatial-sphere coordinates. That is a genuine construct-validity threat, but the most load-bearing problem is more direct: the manuscript itself reports that the performance differences are not statistically significant across all datasets (p>0.05). This is an internal inconsistency with the Abstract's central claim, not a disagreement with external consensus, and it does not depend on accepting or rejecting the evaluator. The reader mentioned this significance contradiction in the rationale but selected the evaluator as the weakest assumption, so my agreement is only partial. My recommended verdict remains REJECT, matching the reader's verdict, because the significance statement removes the inferential basis for 'significantly outperforms.' The concrete paired re-analysis with confidence intervals would settle whether the manuscript's p>0.05 sentence is a reporting error or an accurate account of the data; until that is resolved, the headline claim is unsupported.","tokens_in":9561,"tokens_out":4374,"duration_ms":43749,"concrete_test":"Obtain per-prompt scores from the released code for Table 1 (SpatialPrompts N=120, COCO-derived N=1000, VISOR-derived N=1000) for SmartSpatial and at least SD+AG and SD+ControlNet; compute paired differences for OP, SR, OR, mAP, and IoU and report bootstrap 95% confidence intervals and p-values. If for the headline spatial metrics the 95% CI crosses zero or p>0.05, the 'significantly outperforms' claim is not supported. A useful auxiliary check is to re-score a random subset with human annotators to rule out VLM/coordinate-mapping bias, but the paired significance rerun is decisive.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 5.2 (immediately after Table 1) states: 'Statistical significance tests confirmed that these performance differences are not significant across all datasets (p >0.05).' This is an explicit in-paper limitation statement that undercuts the central claim. The Abstract claims SmartSpatial 'significantly outperforms existing methods' and delivers 'notable improvements'; Section 1 repeats that it 'significantly enhances spatial accuracy'; Section 7 says it 'improves spatial precision.' If the only reported significance test yields p>0.05 across every dataset, then the positive Table 1 margins (e.g., SmartSpatial OP 0.433 vs SD+AG 0.380 on SpatialPrompts; SR 0.358 vs 0.300) are descriptive, not inferential, and could arise from noise, from the evaluation protocol, or from the unvalidated VLM-based scoring. No effect sizes, confidence intervals, or per-sample distributions are reported to support the 'significant' language. Thus the load-bearing assumption of the paper—measurable, statistically reliable gains over SD+AG and SD+ControlNet—is contradicted by the authors' own test. The evaluator-validity concern (ChatGPT-4o perception and hand-coded coordinate mapping) is real, but the significance statement alone is sufficient to reject the headline claim: even a perfectly valid evaluator would not rescue a claim that the reported data do not support.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript proposes SmartSpatial, a training-free enhancement to Stable Diffusion for 3D spatial arrangement. The method injects a depth map from a reference image into ControlNet and applies cross-attention guidance with bounding boxes, optimizing a latent-space loss that combines UNet and ControlNet terms. It also proposes SmartSpatialEval, a VLM-based evaluation framework with three metrics (OR, OP, SR) built on a 'spatial sphere' representation of spatial relations in text and images. Experiments on SpatialPrompts, COCO2017, and VISOR compare against MultiDiff, eDiff-I, BoxDiff, SD, SD+AG, and SD+ControlNet. The paper reports that the proposed method yields higher OP/SR/mAP/IoU scores in Table 1, while CLIP scores remain competitive, and provides an ablation study on VISOR.","tokens_in":9957,"tokens_out":5148,"duration_ms":46735,"significance":"The combination of training-free spatial control and a graph- and VLM-based spatial metric would be genuinely useful to the text-to-image community, and releasing datasets and code is commendable. However, the central claim that SmartSpatial significantly outperforms existing methods is directly contradicted by the authors' own significance test in Section 5.2 (p>0.05 for all datasets), and the proposed evaluation metric is not validated. As reported, the paper establishes at most descriptive improvements on an unvalidated evaluator, not a statistically reliable or measurable spatial-fidelity gain.","major_comments":[{"comment":"The text immediately after Table 1 states: 'Statistical significance tests confirmed that these performance differences are not significant across all datasets (p >0.05).' This directly contradicts the Abstract, Section 1, and Section 7, which claim that SmartSpatial 'significantly outperforms' existing methods. Since every dataset fails to reach significance, the positive margins in Table 1 (e.g., OP 0.433 vs 0.380 and SR 0.358 vs 0.300 on SpatialPrompts) are descriptive only and do not support the headline claim. The manuscript needs either a properly powered significance test with effect sizes and confidence intervals, or a revision of all significance claims to descriptive language.","section":"Section 5.2, Table 1"},{"comment":"The OP and SR metrics depend on an unvalidated 'spatial sphere' coordinate assignment (e.g., left=(-1,0,0), on=(0,1,0)) and on ChatGPT-4o's text descriptions of generated images. No justification for the coordinate mapping, no sensitivity analysis, and no comparison against human spatial judgments or existing spatial benchmarks is provided. Because a single proprietary VLM mediates the mapping from image to coordinates, and the coordinate mapping is hand-assigned, OP and SR cannot be interpreted as measuring true 3D spatial fidelity without external validation.","section":"Section 4.1, Eq. (5)"},{"comment":"The method's central assumption is that a depth map extracted from one object pair (e.g., 'ball behind box') transfers to a different object pair (e.g., 'vase behind orange'). This assumption is not tested. The paper should vary the reference depth map across object geometries, aspect ratios, and spatial scales, and report whether the target spatial relation and object placement remain correct. Without such experiments, the improved layout metrics could reflect ControlNet copying the reference layout rather than generalizing the intended spatial relation.","section":"Sections 3.2-3.4"},{"comment":"The experimental evaluation lacks any measure of variance or reproducibility. Only one seed (42) is reported, and no error bars, confidence intervals, or per-sample distributions are shown for any metric. Combined with the non-significant difference reported in Section 5.2, the numerical gains in Table 1 cannot be distinguished from experimental noise, and the claim of consistent superiority is therefore not supported by the data as presented.","section":"Section 5.1, Table 1"}],"minor_comments":[{"comment":"The dataset description is internally inconsistent: it opens with '1,000 samples derived from VISOR' but then says 'we randomly selected 336 instances and replaced their spatial terms' without explaining how these numbers relate. Please clarify the sampling procedure.","section":"Section 5.1, VISOR paragraph"},{"comment":"The eDiff-I baseline is cited as [Zhang et al., 2023a], whose title is 'A tale of two features: Stable diffusion complements dino for zero-shot semantic correspondence'; eDiff-I is by Balaji et al., 2023. The citation should be corrected.","section":"Section 5.1, comparison models"},{"comment":"The ablation text says 'all three components (AG, CN, and CNAG) are employed,' but CNAG already denotes the combination of cross-attention guidance with ControlNet, making the naming confusing. Please clarify the component notation.","section":"Section 5.3, Table 2"},{"comment":"Equation (2) is an update rule for the latent variable with momentum, not a loss. The sentence 'The calculations for Lunet and Lcontrol are consistent with those in Eq. 2' is therefore confusing; please distinguish the loss definition from the optimization update.","section":"Section 3.4, Eq. (2)"}],"recommendation":"reject","confidential_remarks":"For the editor: the citation mismatch for eDiff-I and the dataset-size inconsistency in Section 5.1 suggest that the experimental section needs careful re-verification. The deciding issue, however, is the contradiction between the claimed statistical significance and the reported p>0.05 across all datasets, together with the unvalidated VLM/coordinate-mapping evaluator. Unless the significance tests can be redone with adequate power, the evaluator can be validated, and the claims recalibrated, this manuscript is not publishable in its current form."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The headline is the contradiction: the abstract says SmartSpatial \"significantly outperforms\" baselines, but Section 5.2 reports that significance tests found p>0.05 across all datasets. That is not a minor slip. It means the main numerical margins in Table 1 (e.g., OP 0.433 vs 0.380, SR 0.358 vs 0.300) are descriptive only. There are no error bars, confidence intervals, or effect sizes to support the \"significant\" language. The stress-test note is right: even a perfectly valid evaluator would not rescue a claim the reported data do not support.\n\nNow the credit. The method is a clean, training-free combination of known pieces: reference-image depth maps injected through ControlNet, plus cross-attention guidance from Chen et al. The ablation study in Table 2 is genuinely informative and shows each component contributes. The code and SpatialPrompts dataset are released, which is valuable. SmartSpatialEval is also a creative construct: building a \"spatial sphere\" from a VLM description and comparing it to a prompt-derived sphere directly targets a real gap in 3D spatial evaluation.\n\nThe soft spots, in proportion. The evaluator-validity concern is real: the coordinates (left = (-1,0,0), on = (0,1,0)) are hand-assigned, and the whole pipeline depends on ChatGPT-4o's perception, which is proprietary and unvalidated against human judgment or established benchmarks. The OR/OP/SR scores could be measuring the VLM's spatial reasoning rather than the image's actual layout. This is not a cheap shot; it is a missing validation step that the authors should have included. Second, the method requires a reference image with the desired spatial relation, which is an extra input that baselines like SD+ControlNet do not get; the comparison is not quite apples-to-apples. Third, the claim that SmartSpatialEval can serve as an RL reward signal is speculative and untested.\n\nWho should read this? People working on layout control and spatial evaluation in diffusion models. It is not a \"must-cite\" for me, but the evaluation framework, once validated, could become a useful tool. As submitted, the central claim is unsupported by the authors' own statistical test, and the evaluator needs rigorous validation. That said, the paper deserves a serious referee: the method is plausible, the ablation is honest, and the evaluation gap it targets is real. My recommendation: send it to peer review, but expect major revision. The authors need to either soften the significance language to match the data or run a properly powered experiment with confidence intervals, and they need to validate SmartSpatialEval against human spatial judgments and against a non-proprietary VLM.","headline":"The paper's own significance test (p>0.05 for all datasets) directly contradicts the abstract's 'significantly outperforms,' and the evaluator is unvalidated; that combination sinks the central claim as written.","tokens_in":10383,"tokens_out":1787,"would_cite":false,"duration_ms":18135,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"SmartSpatial uses depth maps and attention guidance to make Stable Diffusion place objects where prompts specify.","keywords":["3D spatial arrangement","text-to-image generation","Stable Diffusion","cross-attention guidance","depth map conditioning","layout control","vision-language evaluation","spatial relationship metrics"],"falsifier":"Run SmartSpatialEval on the same generated images with human-assigned 3D coordinates or with a different vision-language model; if method rankings change, for example SD+AG beating SmartSpatial, the reported improvement is an artifact of the evaluator rather than a spatial gain. A complementary check is to measure depth ordering directly from the generated images, verifying that pixels for the object described as 'behind' are consistently farther than the 'front' object.","tokens_in":9411,"feed_emoji":"🧭","tokens_out":10871,"duration_ms":87460,"temperature":0.7,"pith_summary":"The paper attempts to establish that Stable Diffusion's weak handling of 3D spatial language can be corrected without retraining, by feeding the model a depth map of a reference scene and steering its cross-attention maps with a weighted loss. The authors pair this with SmartSpatialEval, an evaluation framework that turns a prompt and a generated image into 'spatial spheres' with coordinates, so object presence, proximity, and relational order can be scored. If the claim holds, text-to-image models gain a practical way to respect relations like 'in front of,' 'behind,' 'on,' and 'under' while keeping image quality, and researchers get a quantitative metric aimed at 3D layout rather than image-text similarity alone.","feed_headline":"Depth guidance improves 3D placement in Stable Diffusion","feed_subtitle":"A depth-plus-attention method beats layout baselines on front/behind/on relations, with a vision-language 3D evaluator.","key_machinery":"Generation side: depth-information injection, where a depth map from a reference image is processed by a ControlNet depth extractor and inserted into the upsampling blocks of the denoising UNet, combined with cross-attention guidance, which extracts attention maps from selected mid and up-sampling cross-attention blocks and applies a loss that concentrates each token's attention inside its bounding box. Evaluation side: the spatial sphere, a graph-based coordinate model in which the center object sits at the origin and every other object is placed at one of eight hand-assigned 3D positions (for example, left = $(-1,0,0)$, on = $(0,1,0)$); shortest paths from the center in the prompt's sphere and the image's sphere give the OP and SR scores.","core_discovery":"SmartSpatial's central claim is that combining depth-conditioning with cross-attention guidance removes Stable Diffusion's spatial-arrangement failures. A depth estimator turns an arbitrary reference image, for example 'a ball is behind a box,' into a depth map; a ControlNet depth extractor injects that map into the denoising UNet; and a momentum-based update nudges the latent so that each prompt token's attention mass falls inside its designated bounding box. The final loss is a weighted sum of the UNet and ControlNet guidance terms. In comparisons on SpatialPrompts, COCO2017-derived prompts, and VISOR-derived prompts, the paper reports that SmartSpatial beats layout baselines such as SD+AG and SD+ControlNet on object proximity, spatial relationship, object recognition, IoU, and mAP, with only a minor CLIPScore dip. The companion evaluator, SmartSpatialEval, uses a vision-language model (ChatGPT-4o), dependency parsing, and a spatial-sphere graph to produce the OR, OP, and SR metrics.","pith_inferences":["The paper fixes a set of eight spatial relations; a natural next test is whether the same sphere coordinates extend to graded terms like 'near' or compound relations such as 'between' and 'in the corner.'","Because the evaluator relies on a vision-language model's description of the image, replacing that model or comparing rankings across models would show whether the OR/OP/SR benchmark is stable.","The method's ceiling depends on the reference depth map: if depth estimation fails on stylized or abstract reference scenes, the guidance should degrade, which is a testable prediction."],"forward_implications":["Spatial control becomes available without extra training, so a user can steer Stable Diffusion v1.5 with one reference image and bounding-box constraints.","The OR, OP, and SR metrics give a quantitative target for 3D layout that CLIP and IoU miss, making spatial-fidelity improvements directly measurable.","SmartSpatialEval can also serve as a reward signal for reinforcement-learning fine-tuning of diffusion models, so spatial reasoning could be optimized during training.","Using SmartSpatial to build image-text pairs could supply spatial training data for vision-language models, which currently lack such examples.","The method keeps Stable Diffusion's visual quality while improving layout, with only a small CLIPScore trade-off."],"supporting_citations":[{"why":"Supplies the Stable Diffusion v1.5 backbone that SmartSpatial enhances and the 'SD' baseline.","marker":"[Rombach et al., 2021]"},{"why":"Provides the cross-attention guidance loss and the block-selection choice that SmartSpatial adapts.","marker":"[Chen et al., 2023]"},{"why":"Supplies ControlNet and its depth extractor, which injects the reference depth map into the denoising UNet.","marker":"[Zhang et al., 2023b]"},{"why":"Supplies the dependency parser used to turn prompts and VLM descriptions into spatial relationship graphs.","marker":"[Honnibal et al., 2020]"},{"why":"Provides the VISOR benchmark, one of the three evaluation datasets and the 2D spatial benchmark extended to 3D here.","marker":"[Gokhale et al., 2023]"},{"why":"Supplies the COCO2017 dataset from which 1,000 evaluation prompts are sampled.","marker":"[Lin et al., 2015]"},{"why":"Identifies the DDPO reinforcement-learning method that SmartSpatialEval could reward, connecting the evaluator to training.","marker":"[Black et al., 2024]"}],"fun_headline_variants":["Depth-plus-attention boosts Stable Diffusion's 3D placement","Stable Diffusion gets better at 3D layouts with depth guidance","SmartSpatial: AI art with accurate front-behind-on relations","New framework evaluates 3D spatial accuracy in generated images","Combining depth and attention fixes SD's spatial arrangement"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The reported gains rest on an evaluation that assumes the vision-language model's description of a generated image, together with the hand-assigned sphere coordinates for words like 'left' and 'on', faithfully captures true 3D spatial relations; if either is biased, the OP and SR scores do not measure spatial fidelity.","fun_headline_variants_meta":{"raw":{"variants":["Depth-plus-attention boosts Stable Diffusion's 3D placement","Stable Diffusion gets better at 3D layouts with depth guidance","SmartSpatial: AI art with accurate front-behind-on relations","New framework evaluates 3D spatial accuracy in generated images","Combining depth and attention fixes SD's spatial arrangement"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000635,"raw_usage":{"total_tokens":2917,"prompt_tokens":922,"completion_tokens":1995,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":538,"completion_tokens_details":{"reasoning_tokens":1909}},"tokens_in":538,"tokens_out":1995,"duration_ms":13739,"temperature":1.0,"reasoning_tokens":1909,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T22:43:09.029902+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run SmartSpatialEval on the same generated images with human-assigned 3D coordinates or with a different vision-language model; if method rankings change, for example SD+AG beating SmartSpatial, the reported improvement is an artifact of the evaluator rather than a spatial gain. A complementary check is to measure depth ordering directly from the generated images, verifying that pixels for the object described as 'behind' are consistently farther than the 'front' object.","supporting_citations":[{"cited_title":"spaCy: Industrial-strength Natural Language Processing in Python","cited_arxiv_id":null,"evidence_quote":"Supplies the dependency parser used to turn prompts and VLM descriptions into spatial relationship graphs."}],"review_version":1}