{"id":"eac2935f-3a03-40e9-a343-ef3a1bb3f7de","arxiv_id":"2608.09597","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A two-stage pipeline that learns where to spend a fixed voxel budget for best appearance and then assembles the resulting grid into bricks with zero floating or unstable components, beating prior systems on both fidelity and stability at the compared resolution.","lead":"ResemBrick turns a few casual photos of an object into a colored brick model that both resembles the object and can actually be assembled by hand. It treats the choice of which voxels to fill as a budget-allocation problem, then welds that to a builder that guarantees no floating bricks.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Full-pipeline perceptual target unspecified: if MS-SSIM/LPIPS compare against the reconstructed mesh M rather than the ground-truth object, the photo-to-brick claim is not actually tested.","rationale":"After reading the paper, the central claim is conditional on the metric reference in the full-pipeline comparison. The paper is meticulous about many things (matched budgets, exact-solver certification, explicit scope in H.1), which makes the omission of the target definition conspicuous. The completion-stage claims are internally consistent: the network is trained against a repaired oracle on the mesh M, and the matched-budget comparisons against geometric selectors are fair by construction. The building-stage claims are also carefully scoped: zero floating is by construction, static stability is empirically certified, and H.1 separates these. The only place where the argument can silently support a stronger conclusion than demonstrated is the full-pipeline perceptual evaluation. If the reference is the reconstructed mesh M, the 'from photographs' claim in the abstract and title is not tested; if the reference is ground truth, the paper should say so and would then be measuring reconstruction error as well. The proposed test resolves this by rerunning the same metric pipeline with GT references. I agree with the reader's weakest assumption and recommend keeping the CONDITIONAL verdict, since the method appears sound but the headline claim needs this clarification before acceptance.","tokens_in":30166,"tokens_out":6473,"duration_ms":59127,"concrete_test":"Recompute Table 1 and Table 12 using renderings of the ground-truth OmniObject3D mesh as the reference instead of (or in addition to) the current unspecified target. Concretely: for each of the 198 held-out objects, render the GT mesh and the final ResemBrick assembly (and each baseline's output) from the same 12 fixed viewpoints used in Sec. G.3, at 256^2, and recompute MS-SSIM, LPIPS, PSNR, SSIM, and DISTS. If Ours still beats BrickGPT, Legolization, and BrickLink Studio with the same margins, the central claim is robust; if the ordering or magnitude changes, the claim must be explicitly rescoped to mesh-to-brick conversion and the photo-to-brick promise separated from the front-end's performance. Since OmniObject3D ships GT meshes, this is a pure metric-computation rerun with no retraining.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central full-pipeline claim (abstract, Sec. 4.2, Table 1) is that 'as a complete pipeline, it attains the best perceptual fidelity among prior brick-construction systems' from casual photographs. The truth of this claim depends on what the rendered-view reference is. Sec. 4.1 defines metrics only as 'MS-SSIM and LPIPS ... on rendered views' and never identifies the target; Appendix E.1 says only 'rendered views of the predicted output and the target'; the Table 1 caption is silent. Figure 4 labels the starting point 'input mesh' and the method itself treats the reconstructed mesh M as the object to reproduce (Sec. 3.2), while all baselines consume the same M. The natural reading is therefore that the target is the FreeSplatter-reconstructed mesh M, not the ground-truth scanned object. If that is the case, the perceptual numbers measure mesh-to-brick fidelity only; the reconstruction front end, which the paper itself identifies as a source of propagated errors (Sec. I), is never evaluated. The abstract's 'from photographs' framing and the title's 'from Photographs' would then overstate what is shown: the experiments would not establish that the complete photo-to-brick pipeline resembles the real object, only that the brick stage resembles an intermediate reconstruction that may itself be inaccurate. This directly weakens the strongest_claim, because 'best perceptual fidelity among prior brick-construction systems' would need the qualifier 'given a shared reconstructed mesh.' The reader's concern is exactly this ambiguity, and I find no internal contradiction elsewhere that would displace it as the most load-bearing issue.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes ResemBrick, a two-stage pipeline that converts a handful of casual photographs into a hand-buildable colored brick model. A pose-free reconstructor (FreeSplatter) first produces a watertight textured mesh; a resolution- and budget-conditioned 3D U-Net then selects which surface voxels to occupy under a user-specified occupied-voxel count; a greedy builder with a two-tier, provably terminating repair stage partitions the resulting grid into library bricks and guarantees grounded connectivity. The paper reports that, under a matched budget, ResemBrick's completion network beats geometric voxel selectors in perceptual fidelity, and that the full pipeline attains the best perceptual fidelity among prior brick-construction systems while uniquely reaching zero floating and zero statically unstable bricks on the held-out OmniObject3D set. Experiments are conducted at R=24 for full-pipeline comparisons, with voxelization-stage comparisons at R in {24,32,48} and a resolution-generalization study across 13 resolutions.","tokens_in":30574,"tokens_out":8569,"duration_ms":74411,"significance":"If the claims hold, the paper makes a useful contribution by reframing coarse voxelization as budget allocation and showing that completion and assembly can be co-designed rather than treated as independent stages. The experimental protocol is unusually careful: matched occupancy budgets are enforced by construction, all voxelization baselines share the same frozen SDF, camera poses, and mesh, the stability numbers are certified by an exact force-balance solver, a separate sweep-validation set is used for building-stage hyperparameters, and the supplement (Sec. H.1) explicitly separates what is guaranteed by construction (grounded connectivity) from what is empirical (static stability and physical hand-buildability). The honest treatment of the lazy-greedy oracle's non-submodularity is also a strength. However, the central full-pipeline claim is currently under-specified because the perceptual evaluation never identifies its reference target, and there is an apparent inconsistency in the reported stability of the colored system across tables. These issues are load-bearing for the abstract's 'from photographs' and 'zero unstable bricks' claims and need to be resolved.","major_comments":[{"comment":"The rendered-view perceptual metrics never specify what the reference 'target' is. The metric definition in Sec. 4.1 says only 'MS-SSIM and LPIPS ... on rendered views', and Appendix E.1 says 'rendered views of the predicted output and the target' without defining 'target'. If the target is the FreeSplatter-reconstructed mesh M rather than the ground-truth scanned object, then the Table 1 full-pipeline numbers measure mesh-to-brick fidelity only, and the title/abstract claim of reconstruction 'from photographs' is not tested; Section I itself identifies the reconstruction front end as a source of propagated errors, so the distinction matters. Please state the reference object explicitly in Sec. 4.1 and in the Table 1 caption. If it is M, please add an evaluation against the ground-truth object or qualify the abstract and title claims.","section":"Sec. 4.1, Appendix E.1, Table 1"},{"comment":"The stability numbers for the deployed system appear inconsistent. Table 5 reports 'Ours (col.)' on the 198-object held-out omni set with n_u=0.45 and Stab%=98.0, while Table 1 reports 'Ours' with Stab%=100 on the same set at the same R=24. If the full pipeline includes per-brick color, these two entries cannot both describe the same configuration; the abstract's 'uniquely reaching zero floating and zero unstable bricks' claim is therefore at risk. Please define exactly what 'Ours' and 'Ours (col.)' denote relative to the full pipeline, and report the stability of the exact configuration used in Table 1.","section":"Table 5 vs. Table 1"}],"minor_comments":[{"comment":"The phrase 'one weight set spanning 13 resolutions' could be read as a training claim; since only three resolutions are trained, please phrase this as 'evaluated on 13 resolutions' or otherwise clarify.","section":"Abstract and Sec. 4.3"},{"comment":"The term 'unfiltered held-out objects' should be defined explicitly (e.g., no post-hoc exclusion by reconstruction quality, stability, or category) so that readers understand the scope of the zero-floating/zero-unstable claim.","section":"Sec. 4.1"},{"comment":"The oracle loss uses LPIPS on colorless depth renderings while the final perceptual evaluation uses color renderings; please state whether this mismatch is intentional and discuss its effect on the completion network's training signal.","section":"Sec. 3.2, Eq. (1) and Appendix E.1"},{"comment":"The names 'uniform' and 'Top-K' are used for the identical selector; please state this equivalence in the main text or in the table captions to avoid confusion.","section":"Appendix E.1"},{"comment":"Please report the number of objects and the exact bootstrap resampling procedure used for the shaded 95% confidence bands.","section":"Figure 6"},{"comment":"The loop 'for all 4-connected components R over all layers' should specify that 4-connectivity is computed within each layer, since walls are connected only through vertical overlap in the subsequent greedy placement.","section":"Algorithm 1, line 2"}],"recommendation":"major_revision","confidential_remarks":"The central idea is sound and the experimental protocol is more careful than is typical for this area. The two major issues are fixable in revision: one requires an explicit statement (and possibly a new experiment) about the perceptual target, and the other requires reconciling the reported stability of the colored full pipeline. I do not see grounds for rejection if these are resolved. The manuscript is within the scope of the journal."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. ResemBrick is a genuinely careful systems paper: it reframes coarse voxelization as budgeted allocation, couples a resolution-conditioned 3D U-Net with a buildability-aware greedy builder, and backs its claims with matched-budget comparisons, exact-solver certification, and unusually explicit scope statements. The co-design principle is real. The second thing is the problem: the headline 'from photographs' claim is under-specified. The full-pipeline perceptual metrics (Table 1) never say whether the rendered target is the ground-truth object or the FreeSplatter-reconstructed mesh. On the natural reading it is the reconstructed mesh, since all full-pipeline baselines consume that same mesh, the method treats that mesh as the object to reproduce, and the authors' own limitation section says the contribution is downstream of the recovered mesh. If that is right, the experiments test mesh-to-brick fidelity, not photo-to-brick fidelity, and the abstract's 'from casual photographs' framing overstates what is shown. That is the load-bearing weakness. The fix is a sentence in the caption and a qualifier in the abstract.\n\nCredit where due. The allocation reframing is genuinely new, and the single 0.10M-parameter network that spans 13 resolutions without retraining is a solid result. The greedy builder with the provably terminating repair, plus the exact-Gurobi certification of all stability numbers, is a model of how to report physical-buildability claims. The separate sweep-validation set, the frozen SDF/poses/mesh for all baselines, and the explicit scope discussion in Sec. H.1 of the supplement are all exemplars. The oracle-ablation honesty (showing lazy greedy tracks exhaustive within 0.6% loss) is also good.\n\nSoft spots, in proportion. The target ambiguity is the main one; it directly weakens the strongest claim. The abstract's 'surpasses existing voxel selectors' should be scoped to the compared baselines, as the reader noted. Table 1 has no significance stars or confidence intervals, unlike the voxelization tables, which is an inconsistency. A few hyperparameter values (perceptual loss weights, repair gap tolerance) are only in the supplement. None of these are fatal; the central methodology holds up.\n\nWho this is for: anyone building physical assemblies from 3D data, and anyone thinking about discretization and assembly as coupled stages. It deserves a serious referee. I would accept it for review, with the target clarification as a required revision.","headline":"Careful, well-scoped brick pipeline with a real co-design idea, but the full-pipeline 'from photographs' claim hinges on an unspecified rendered target that looks like the reconstructed mesh, not the ground-truth object.","tokens_in":31072,"tokens_out":3456,"would_cite":true,"duration_ms":29594,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"ResemBrick claims that coupling budgeted occupancy completion with buildability-aware assembly lets a photo-to-brick pipeline achieve better perceptual fidelity than prior systems while uniquely producing fully stable assemblies on…","keywords":["brick reconstruction","perceptual fidelity","buildability","budgeted occupancy completion","pose-free multi-view reconstruction","voxelization","greedy placement","stability repair"],"falsifier":"Re-run the full-pipeline comparison with the ground-truth object (or the original photographs) as the render reference instead of the reconstructed mesh, and see whether ResemBrick still leads on MS-SSIM and LPIPS. Separately, run the pipeline on a held-out set deliberately rich in thin, arch-like, or bridge-like objects and count how often the exact force-balance solver certifies full stability, since the paper's own scope analysis excludes cross-layer support from the guaranteed properties.","tokens_in":29913,"feed_emoji":"🧱","tokens_out":8023,"duration_ms":67506,"temperature":0.7,"pith_summary":"The paper tries to settle a specific conflict: when a 3D object is turned into a brick model on a coarse grid, the same voxels cannot simultaneously maximize visual resemblance and structural stability. Its central claim is that this conflict is best resolved by treating discretization itself as a budget-allocation problem, given a target number of occupied voxels, choosing which surface voxels to fill for best appearance, and then handing that grid to an assembly stage that is buildability-aware by construction. ResemBrick couples a single resolution-conditioned 3D U-Net that scores candidate surface voxels in one forward pass with a greedy brick placer that rewards support, color fidelity, and look-ahead, plus a deterministic repair that grounds every floating component. The reported result is the best perceptual fidelity among prior brick-construction pipelines while uniquely reaching zero floating and zero unstable bricks on unfiltered held-out objects. If this holds, it shows that discretization and assembly should be co-designed rather than treated as independent pipeline stages.","feed_headline":"Zero floating bricks and best fidelity from casual photos","feed_subtitle":"Voxelization as budget allocation plus a stability repair makes brick models that both resemble the object and stand.","key_machinery":"The central mechanism is budgeted occupancy completion: instead of uniform voxelization, the network is asked to allocate a fixed number of occupied voxels among the surface band, the shell of voxels straddling the recovered mesh, and it does so with a fully-convolutional 3D U-Net conditioned on resolution and realized density via feature-wise linear modulation (FiLM) layers, trained in two phases: oracle distillation, then annealed straight-through refinement against rendered depth. The second mechanism is the buildability-aware greedy assembler: a layer-wise merge scores candidate library bricks by support fraction, a rescue bonus for bricks spanning overhangs, importance-rarity color salience, and a look-ahead penalty for stranded cells, followed by a two-tier floating-component repair (zero-deformation recombination, then cap/shelf bridging) that is deterministic and provably terminating.","core_discovery":"The core discovery is that the fidelity ceiling of a brick model is set by how the occupancy budget is spent, so the voxelizer and the assembler must be coupled. The paper's completion network is distilled from an offline greedy oracle that selects the surface-band voxels whose filling most reduces a geometry-aware perceptual loss, then refined with a straight-through estimator against rendered depth; the same weight set serves resolutions 24, 32, and 48 and generalizes to 13 resolutions. The assembly stage then partitions each layer into library bricks using a greedy score that rewards supported placement, rare-color preservation, and look-ahead, and a two-tier repair with cap and shelf bridges guarantees grounded connectivity by construction. On the held-out 198-object set at resolution 24, ResemBrick reports MS-SSIM 0.955 and LPIPS 0.075 versus 0.940/0.081 for BrickGPT and 0.940/0.082 for Legolization, with 100% of assemblies certified stable by an exact force-balance solver; under a matched occupancy budget, the completion network also beats all geometric voxel selectors on the perceptual metrics.","pith_inferences":["The paper never states whether the rendered reference in the full-pipeline fidelity table is the ground-truth object or the reconstructed mesh; if it is the mesh, the comparison tests brick-model-versus-its-own-source rather than photo-to-brick fidelity, so the reader should check the supplement's evaluation protocol before trusting the headline numbers.","A natural extension the paper leaves implicit: propagating a differentiable stability or brick-count signal from the assembler back into the occupancy network might push the fidelity ceiling higher, particularly for arch- or bridge-like shapes whose mid-assembly states are currently excluded from the buildability guarantee.","At the coarsest grids the network trails a sparse SDF shell on MS-SSIM, so a hybrid selector that switches by resolution could dominate both; the paper reports the gap but does not propose the hybrid."],"forward_implications":["Because one learned selector spans 13 resolutions, a user can change the brick budget without retraining the discretizer.","The repair stage is deterministic and provably terminating, so zero floating components is a structural property of the pipeline rather than a statistical outcome.","Since the completion budget upper-bounds what any downstream assembly can achieve, improvements to the allocator's fidelity directly raise the quality ceiling of brick reconstruction.","The distilled stability surrogate certifies per-brick stability in milliseconds, which makes it practical to screen candidate assemblies during construction rather than only after completion.","The four real-world hand-builds show the exported layer-by-layer instructions assemble standing models without manual edits, on the tested objects."],"supporting_citations":[{"why":"Pose-free reconstructor used as the front end, lifting casual photographs to the mesh the brick model is built from.","marker":"Xu, Gao, and Shan 2025"},{"why":"OmniObject3D dataset supplies the training and held-out 198-object evaluation set used for all pipeline numbers.","marker":"Wu et al. 2023"},{"why":"StableLego's exact Gurobi force-balance MILP certifies the reported Stab% values and provides the per-brick stability labels for the distilled surrogate.","marker":"Liu et al. 2024"},{"why":"BrickGPT is the primary full-pipeline and building-stage baseline that ResemBrick compares against on fidelity and stability.","marker":"Pun et al. 2025a"},{"why":"Legolization is a fidelity-oriented brick-construction baseline run on the same input grid and held-out set.","marker":"Luo et al. 2015"},{"why":"BrickLink Studio is the commercial converter baseline in the full-pipeline comparison.","marker":"BrickLink 2026"},{"why":"STABLE is the learned buildability-focused building-stage baseline evaluated on both StableText2Brick and OmniObject3D.","marker":"Xu et al. 2026"},{"why":"LPIPS is the perceptual distance used both in the completion oracle's training loss and as a headline evaluation metric.","marker":"Zhang et al. 2018"},{"why":"MS-SSIM is the other headline perceptual metric reported for the full-pipeline and completion comparisons.","marker":"Wang, Simoncelli, and Bovik 2003"},{"why":"StableText2Brick is the clean buildability benchmark used for the building-stage comparison.","marker":"Pun et al. 2025b"}],"fun_headline_variants":["Zero floating bricks and best perceptual fidelity","Couple voxelization and assembly for stable, faithful bricks","One network, 13 resolutions, zero unstable bricks","From casual photos to brick models that resemble and stand"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The full-pipeline fidelity claim assumes the rendered 'target' used for MS-SSIM and LPIPS is the ground-truth object appearance, not the intermediary mesh reconstructed from the photos that the brick model is built from.","fun_headline_variants_meta":{"raw":{"variants":["Zero floating bricks and best perceptual fidelity","Couple voxelization and assembly for stable, faithful bricks","One network, 13 resolutions, zero unstable bricks","From casual photos to brick models that resemble and stand"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000397,"raw_usage":{"total_tokens":2110,"prompt_tokens":1006,"completion_tokens":1104,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":622,"completion_tokens_details":{"reasoning_tokens":1042}},"tokens_in":622,"tokens_out":1104,"duration_ms":9983,"temperature":1.0,"reasoning_tokens":1042,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T14:21:58.839437+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the full-pipeline comparison with the ground-truth object (or the original photographs) as the render reference instead of the reconstructed mesh, and see whether ResemBrick still leads on MS-SSIM and LPIPS. Separately, run the pipeline on a held-out set deliberately rich in thin, arch-like, or bridge-like objects and count how often the exact force-balance solver certifies full stability, since the paper's own scope analysis excludes cross-layer support from the guaranteed properties.","supporting_citations":[{"cited_title":"Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) , pages =","cited_arxiv_id":null,"evidence_quote":"Pose-free reconstructor used as the front end, lifting casual photographs to the mesh the brick model is built from."},{"cited_title":"IEEE Robotics and Automation Letters , volume =","cited_arxiv_id":null,"evidence_quote":"StableLego's exact Gurobi force-balance MILP certifies the reported Stab% values and provides the per-brick stability labels for the distilled surrogate."},{"cited_title":"ACM Transactions on Graphics , volume =","cited_arxiv_id":null,"evidence_quote":"Legolization is a fidelity-oriented brick-construction baseline run on the same input grid and held-out set."},{"cited_title":"and Shechtman, Eli and Wang, Oliver , title =","cited_arxiv_id":null,"evidence_quote":"LPIPS is the perceptual distance used both in the completion oracle's training loss and as a headline evaluation metric."}],"review_version":1}