{"id":"f95f4dad-8883-4b2d-bfd1-937c72ea8758","arxiv_id":"2608.01973","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"Roomer repairs local violations in generated furniture layouts through object-grounded diagnosis, vision-language-model-proposed local edits, and verification-gated commit, improving physical validity and practical usability across generators.","lead":"This paper introduces Roomer, a repair system that finds mistakes in AI-generated 3D room layouts, identifies the furniture objects responsible, and applies small verified edits to fix them. It also contributes a training dataset and an evaluation protocol for physical validity and everyday usability.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Practical-usability claim rests on a metric that is also the repair objective; professional validation samples only baseline layouts, so it does not certify Roomer's repaired outputs.","rationale":"I read the paper as an empirical systems claim: a verification-gated local repair loop with object-grounded attribution improves physical validity and practical usability over initial and external-generator layouts. That claim is plausible and supported by internal controls: frozen common-1100 cohort, byte-identical initialization for ablations, rollback, deterministic verification, and detailed metric definitions. The strongest independent support is the professional-validation monotonic trend (Z=4.67), which anchors Practical as a meaningful proxy. The load-bearing weakness is that the proxy is also the optimization target and the validation set does not include Roomer outputs. This is the same point the reader made; I agree. The concern does not invalidate the paper but makes acceptance conditional on closing the gap. A direct blinded evaluation of Roomer-final layouts at matched Practical levels is the natural test. Secondary reproducibility concerns (no code/data release, missing error bars in main tables) do not change the conditional verdict beyond what the reader stated.","tokens_in":31569,"tokens_out":5316,"duration_ms":56552,"concrete_test":"Run a blinded professional study on 90 Roomer-final layouts (plus Roomer-repaired external-generator outputs) stratified by the same Practical levels, balanced by room type and disjoint from the repair cohort, using the same ten design-experienced evaluators and approval threshold. Compare monotonic approval across levels and, critically, the approval rate of Roomer High-Practical layouts against the baseline High-Practical layouts from Appendix M. If Roomer High layouts are approved at comparable or higher rates, the metric is not overfit to optimization; if they are approved at lower rates, the Practical-based usability claim is not established for Roomer.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central usability claim ('improves physical validity and practical usability') is measured by the same five Practical rule families that Roomer's verification gate optimizes (Eq. 3 with V; Eq. 4 micro-average). The paper recomputes final scores and uses frozen definitions, so the measurement itself is internally consistent; the gap is external. The professional validation in Appendix M samples 90 layouts from baseline generators only, stratified by Practical level, and shows monotonic approval. That anchors the metric on the baseline distribution; it does not show that Roomer's repaired layouts, which are optimized against those rules and may reach High Practical by unusual local edits, receive comparable professional approval. If optimizing the frozen rules exploits proxy loopholes or displaces failures to rule families not in the five (e.g., aesthetic or ergonomic issues outside the rule set), the reported Practical gains would not establish the abstract's usability claim. The paper's own limitation statement (conclusions) concedes coverage is limited to its predefined residential rule set.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes Roomer, a post-generation repair module for 3D indoor layouts. It converts a complete layout into an object-addressable RoState, detects rule-based violations, constructs an object-grounded RoReview, and uses a geometry-conditioned VLM planner to propose a StatePatch action. A deterministic solver instantiates candidate edits, and each candidate is committed only if a full-scene verification gate confirms target resolution, no new hard violations, preservation of protected relations, structural validity, and global decrease of a family-balanced residual. The planner is trained on Roomer-CC, a controlled-corruption dataset built from 3D-FRONT layouts. The paper also introduces Roomer-Eval, which adds five rule families for practical spatial usability to standard distributional and physical metrics. Experiments on a frozen 1,100-room cohort and on outputs of four external generators report consistent gains in physical validity and Practical, with ablations showing the value of RoReview, geometry conditioning, deterministic search, and verification-gated commit.","tokens_in":31794,"tokens_out":7046,"duration_ms":64770,"significance":"Roomer addresses a real and under-served problem: generated indoor layouts often contain local, object-level violations that full-scene regeneration disrupts. The empirical design is a clear strength: the frozen common-1100 cohort, density-controlled baseline selection that avoids evaluation metrics, commit-prefix pinning and seed records for baselines, byte-identical initialization for ablations, and rollback semantics are all described at a level that supports the internal claims. If the practical-usability claim can be externally anchored, Roomer would be a useful generator-agnostic post-processing module. However, the current manuscript does not yet provide that external anchor, and the main quantitative tables lack uncertainty estimates. No code, data, or checkpoints are released, which limits reproducibility.","major_comments":[{"comment":"The abstract and Experiments section claim that Roomer improves practical usability, but the Practical metric used to measure usability is also the objective that the repair loop optimizes. The verification gate's aggregate residual V in Eq. (3) is family-balanced over hard, content, relational, and practical rules (Appendix F, Eq. (A21)), and the practical families are exactly the five rule families in Roomer-Eval (Appendix J). Consequently the 10.48-point Practical gain in Table 1 and the before/after gains in Table 2 partly measure the system's success at optimizing the evaluation instrument. The Conclusions concede that the system is limited to its predefined residential rule set. To support the usability claim, the paper needs an external anchor that is not used in repair: for example, professional judgments on Roomer-repaired outputs, or a held-out rule family never seen by the verifier.","section":"Verification gate (Eq. 3) and Roomer-Eval (Eq. 4)"},{"comment":"The professional validation samples 90 layouts from baseline generators only and is disjoint from the Roomer repair cohort. It establishes a monotonic relationship between Practical strata and professional approval on the baseline distribution, but it does not certify that Roomer's repaired outputs receive comparable approval. A repaired layout may reach High Practical by local edits that satisfy the frozen rules while displacing a failure to an aspect outside the five families, such as ergonomics or aesthetics. The sentence in the Experiments section — 'the higher Practical scores achieved by Roomer reflect improvements that are aligned with professional usability judgments' — is therefore not directly supported. I request a professional validation set that includes Roomer-repaired outputs stratified by Practical level, evaluated with the same blinding and the same Cochran–Armitage test.","section":"Appendix M, Professional Validation"},{"comment":"Main quantitative claims are reported as single point estimates without confidence intervals or standard errors for OOB, COL, Practical, Target Resolution, and Strict Safe Repair. For example, Table 2 reports DiffuScene-RS Practical improving from 45.09% to 47.73%, and Table 3 reports several ablation differences of two to five points; without uncertainty estimates it is impossible to tell whether these differences are meaningful. Appendix I reports bootstrap variance for FID and KID only, and the appendix states that OOB surface sampling has no fixed seed, tying the reported values to a single execution. Please provide bootstrap confidence intervals or standard errors for all central metrics in the main tables, and report variance over repeated OOB evaluations.","section":"Tables 1–3; Appendix I"}],"minor_comments":[{"comment":"Several tables have formatting problems in the supplied text: Table A9 shows concatenated numerical values such as '57.425.1021.73', and the caption of Table A9 and Table 1 use 'oncommon-1100' without a space. Please ensure the camera-ready version renders all columns and numbers cleanly.","section":"Tables and typesetting"},{"comment":"No code, checkpoints, or data are released. Given the complexity of Roomer-CC and the Roomer-Eval evaluator, an artifact release would materially support reproducibility and would help readers verify the unusual care taken in the experimental protocol.","section":"Reproducibility artifacts"},{"comment":"Krippendorff's alpha is reported as 0.65 with confidence interval [0.54, 0.75], which is moderate agreement. Please state whether the Cochran–Armitage result is robust when any single evaluator is excluded, and justify the 'at least seven Yes votes' threshold.","section":"Appendix M, inter-rater reliability"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is within the journal's scope and the empirical work is unusually careful. The main risk is the circularity around Practical: the repair objective and the evaluation metric share the same rule families, and the professional validation does not include Roomer-repaired outputs. If the authors add professional validation on Roomer outputs and report uncertainty estimates for the central metrics, I would support acceptance. No concerns about citation patterns or novelty disclosure."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here is the thing to know: this is a serious, unusually careful empirical systems paper. The central idea — formulate residual layout violations as sparse, object-grounded repair tasks, then use a VLM planner plus a deterministic solver and only commit edits that pass full-scene verification — is new in combination and practically useful. It is a drop-in repair module for any layout generator, and the experiments back that up better than most work in this area. What is genuinely new: Roomer-CC, the controlled-corruption dataset with 67k paired repair examples; the geometry-conditioned planner with repair queries; the deterministic fallback search; and Roomer-Eval, which adds a practical-usability dimension that the field lacks. The ablations are well designed: frozen cohorts, byte-identical initial states, commit-prefix pinning for baselines, and clear separation of repair-time detection from final evaluation. The transfer results across four external generators are credible, and the physical-validity gains (OOB, COL) are measured with independent assembled-mesh metrics, so those results are solid. The soft spot is exactly what the stress-test note says. The five Practical rule families are both what Roomer optimizes (through the verification gate) and what Roomer-Eval measures, so the reported practical-usability gains partly reflect optimizing the evaluation instrument. The paper does the right internal things — frozen definitions, recomputed final scores — but the professional validation in Appendix M only samples baseline layouts. It shows the metric correlates with professional judgment on the baseline distribution; it does not show that Roomer's repaired layouts, which are optimized against those same rules, get comparable professional approval. That gap is real but not fatal. It is also bounded by the paper's own limitation statement: coverage is limited to the predefined residential rule set, so the claim is about those five families, not general usability. The other issues are relatively minor: no released code or data, no error bars in main tables. The FID/KID numbers also rely on the Qwen-Image upstream generator, which is fine, but it means the claimed FID/KID improvements are not the main selling point. Who is this for: anyone working on indoor layout generation, embodied AI simulation, or scene editing. It gives the field a useful repair baseline and a more complete evaluation protocol. It deserves a serious referee. The main requests should be code/data release, error bars, and a professional validation that includes repaired Roomer outputs alongside baselines. I would not desk-reject it; I would send it to review and expect a conditional accept after those additions.","headline":"A genuinely careful empirical systems paper on verification-gated local repair for 3D indoor layouts, with one real weakness: the practical-usability metric is also the repair objective, and the professional validation doesn't cover Roomer's own outputs.","tokens_in":792,"tokens_out":1640,"would_cite":true,"duration_ms":25530,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Roomer claims that residual local violations in generated indoor layouts can be repaired by verification-gated local edits that preserve already-valid regions.","keywords":["3D indoor layout synthesis","local layout repair","object-grounded attribution","verification-gated commit","controlled-corruption dataset","vision-language planner","practical spatial usability","layout evaluation"],"falsifier":"Take a held-out set of Roomer-repaired outputs and unrepaired baselines, have fresh interior-design experts rate them blind, and check whether approval rises with Practical on Roomer's own outputs; if it does not, the usability claim fails. Alternatively, add a sixth common usability rule family such as kitchen work-triangle clearance or bath-fixture spacing to the evaluator and rerun the repair loop; if the added rule reverses the before/after improvement, the five frozen rules missed a dominant failure class and the claim that Roomer improves practical usability generically fails.","tokens_in":31391,"feed_emoji":"🛋️","tokens_out":7270,"duration_ms":62334,"temperature":0.7,"pith_summary":"Roomer's central claim is that the residual failures left by indoor-layout generators are sparse, local, and attributable to a small set of furniture objects, so they can be repaired by targeted local edits instead of full-scene regeneration. The paper presents a reflective repair loop: measured violations are bound to implicated objects, a vision-language planner proposes a structured local edit, a deterministic solver generates verified candidates, and a patch is committed only if full-scene re-verification shows the target resolved, no new hard violations, protected relations preserved, and overall residual reduced. Training on a controlled-corruption dataset of 67,550 pairs and evaluating with a new protocol that adds practical-usability rules, the paper reports that Roomer lowers out-of-bounds and collision rates and improves usability on its own layouts and on outputs of four external generators. If true, this makes post-generation local repair a reliable and generator-agnostic complement to layout synthesis.","feed_headline":"Roomer repairs layout flaws by editing only the offending objects","feed_subtitle":"Verification-gated local edits cut collisions and out-of-bounds furniture and lift usability across five generators.","key_machinery":"The carrying mechanism is the Roomer loop: RoState (an object-addressable canonical layout), RoReview (instance-grounded violation evidence that attributes each measured failure to concrete objects and roles), and a schema-constrained StatePatch whose action, target, and parameter seed are proposed by the planner and instantiated by a deterministic solver. The load-bearing identity is the verification gate: a candidate patch is accepted only if it resolves the target issue, introduces no new hard-violation keys, preserves all protected satisfied relations, keeps structural validity, and strictly decreases the family-balanced residual. This gate is what converts the planner's suggestion into a safe commit and supplies the 'preserve valid regions' part of the claim.","core_discovery":"The paper's discovery claim is that layout repair is better modeled as verification-gated local state repair than as regeneration: most violations are caused by a few objects, and a deterministic full-scene check can decide which local edit is safe. Roomer operationalizes this with RoState, RoReview, and StatePatch: the first is an object-addressable layout encoding; the second converts each measured rule failure into a tuple of violation type, implicated entities, relational roles, and geometric measurements; the third is the structured edit (MOVE, ROTATE, SCALE, INSERT, DELETE, REPLACE) proposed by a geometry-conditioned vision-language planner. The deterministic solver instantiates a finite candidate set and commits the first patch passing the five-clause verification gate. On a frozen cohort of 1,100 held-out rooms, the method improves hard validity, target resolution, strict safe repair, and non-target preservation, while raising the Practical usability score from 72.50% to 82.98% and transferring to four external generators.","pith_inferences":["The verification-gated local-repair pattern is not tied to residential furniture; it could transfer to other structured generation domains where errors are sparse and measurable, such as circuit-board placement or warehouse layout.","Roomer's ceiling is set by its rule library and by the action–issue coverage of its controlled-corruption data; adding a new usability rule family after training would likely require new paired supervision because the planner learns specific action–issue mappings.","The professional validation anchors the Practical metric on baseline layouts rather than on Roomer's own repaired outputs, so a direct blind expert study on Roomer-repaired scenes is the natural missing test of the usability claim.","A stricter test of the framework would be to treat the five Practical rule families as an open, versioned library and check whether repair still improves a newly added rule family without retraining the planner."],"forward_implications":["Indoor layout generators no longer need to be globally perfect; a verification-gated post-generation repair stage can absorb their residual errors.","Practical usability should be measured separately from distributional quality and collision or out-of-bounds checks, because the three can move independently.","Deterministic candidate search plus full-scene verification can compensate for imperfect vision-language proposals, making repair robust to planner noise.","Roomer is generator-agnostic: any layout that can be converted into RoState can be run through the same repair loop, as demonstrated on four external generators.","Because rejected candidates are rolled back, repair is safe by construction: a failed edit leaves the committed layout unchanged."],"supporting_citations":[{"why":"Supplies the valid 3D-FRONT layouts from which Roomer-CC corruptions are derived and the held-out rooms in common-1100.","marker":"Fu et al. 2021a"},{"why":"Supplies the 3D furniture shapes used in assembly and same-category asset retrieval.","marker":"Fu et al. 2021b"},{"why":"The image-generation model fine-tuned as Roomer's upstream generator for the initial layouts.","marker":"Wu et al. 2025"},{"why":"The vision-language backbone for the repair planner, kept frozen with LoRA adapters.","marker":"Bai et al. 2025"},{"why":"DiffuScene, an external diffusion generator whose frozen outputs Roomer repairs in the transfer experiments.","marker":"Tang et al. 2024"},{"why":"InstructScene, an instruction-driven external generator used as a transfer-test baseline.","marker":"Lin and Mu 2024"},{"why":"SemLayoutDiff, a diffusion-based external generator used as a transfer-test baseline.","marker":"Sun, Goel, and Chang 2026"},{"why":"ReSpace, an autoregressive external generator used as a transfer-test baseline.","marker":"Bucher and Armeni 2026"},{"why":"SceneEval, whose assembled-mesh evaluator supplies the OOB and COL metrics in Roomer-Eval.","marker":"Tam et al. 2026"}],"fun_headline_variants":["Roomer repairs layouts by editing only the problematic objects","Surgical layout repair: Roomer fixes rooms object by object","Roomer: verification-gated edits fix layout violations locally","Repair rooms by patching the objects that cause flaws","Roomer: local object edits, verified globally, fix layouts"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The entire practical-usability conclusion rests on the five author-defined Practical rule families (bedside clearance, dining-table clearance, living-room functional angle, door-swing proxy, and walkable connectivity) being a fair stand-in for what makes a layout usable, and the expert validation that anchors these rules was run on baseline layouts rather than on Roomer's repaired outputs.","fun_headline_variants_meta":{"raw":{"variants":["Roomer repairs layouts by editing only the problematic objects","Surgical layout repair: Roomer fixes rooms object by object","Roomer: verification-gated edits fix layout violations locally","Repair rooms by patching the objects that cause flaws","Roomer: local object edits, verified globally, fix layouts"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000646,"raw_usage":{"total_tokens":2986,"prompt_tokens":978,"completion_tokens":2008,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":594,"completion_tokens_details":{"reasoning_tokens":1926}},"tokens_in":594,"tokens_out":2008,"duration_ms":13329,"temperature":1.0,"reasoning_tokens":1926,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T15:02:45.633424+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a held-out set of Roomer-repaired outputs and unrepaired baselines, have fresh interior-design experts rate them blind, and check whether approval rises with Practical on Roomer's own outputs; if it does not, the usability claim fails. Alternatively, add a sixth common usability rule family such as kitchen work-triangle clearance or bath-fixture spacing to the evaluator and rerun the repair loop; if the added rule reverses the before/after improvement, the five frozen rules missed a dominant failure class and the claim that Roomer improves practical usability generically fails.","supporting_citations":[{"cited_title":"The Twelfth International Conference on Learning Representations,","cited_arxiv_id":null,"evidence_quote":"InstructScene, an instruction-driven external generator used as a transfer-test baseline."},{"cited_title":"Wang and Angel X","cited_arxiv_id":null,"evidence_quote":"SceneEval, whose assembled-mesh evaluator supplies the OOB and COL metrics in Roomer-Eval."}],"review_version":2}