{"id":"f8d6d7c4-dae3-47ea-b37a-8df78c66d3af","arxiv_id":"2411.19292","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"UrbanCAD retrieves a matching CAD model from a single car image, optimizes its materials, and inserts it into reconstructed urban scenes, showing that perception models degrade when the cars are edited into out-of-distribution poses.","lead":"UrbanCAD turns a single street photo of a car into a 3D model that can be opened, relit, and dropped into new scenes, by retrieving a similar free CAD model and optimizing its materials. It aims to help autonomous driving simulators generate rare, safety-critical scenarios for testing perception systems.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'digital twin' claim rests on unvalidated per-instance geometry: CAD retrieval is semantic, not geometric, and the reported photorealistic metrics are distribution-level, so novel-view and OOD results may not correspond to the actual observed vehicle.","rationale":"The central claim that UrbanCAD creates 'digital twins' would require per-instance geometric and material agreement with the observed vehicle, including unseen sides. The paper's own Section 6 and Section 10.3 acknowledge that this condition is only approximately met. The reported photorealism metrics (Table 1) are distributional (FID/KID on 360-degree renderings of 30 models versus a real-car dataset) and therefore cannot confirm that a given asset matches a given input vehicle away from the reference view; LPIPS in Table 1 is evaluated only under matched poses. Table 6's geometry numbers are on ShapeNet, not on the KITTI-360/MVMC instances central to the paper, and do not measure rendering fidelity from novel views. Without a per-instance held-out-view check, the 'digital twin' label is unsupported, and the OOD door-opening experiment cannot separate the effect of the OOD state from the sim-to-real gap or the wrong-geometry artifact. This is exactly the weakness the reader flagged, so I agree with the conditional verdict. The paper deserves credit for disclosing the limitation and reporting failure cases, and for ablations showing the material and lighting modules matter; the concern is not that the system is useless, but that the strongest claims outrun the evidence. The proposed KITTI-360 held-out-view test would settle whether the retrieved model actually behaves as a twin on unseen sides; the closed-door control would settle whether Table 3's degradation is due to the OOD configuration. If both pass, the claims would be substantially strengthened; if not, the paper should be reframed around 'digital cousins' with predictable geometry mismatch, as concurrent work ACDC does.","tokens_in":21425,"tokens_out":6999,"duration_ms":63577,"concrete_test":"On KITTI-360, select 20 target vehicles, build each asset from one source frame, and render from a different frame in which the same physical vehicle is visible; compute masked LPIPS and silhouette IoU between the rendering and the real held-out image. Also render the same asset with doors closed versus open in identical backgrounds and measure the iIOU drop of the perception models; if the closed-door condition drops as much as the open-door condition, Table 3's OOD effect is confounded by the render gap.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The strongest claim is that UrbanCAD produces 'photorealistic 3D vehicle digital twins' whose 360-degree rendering and OOD edits (e.g., door opening) are realistic enough for safety-critical testing. This requires that the retrieved Objaverse CAD model matches the observed vehicle's geometry and part layout on all sides, including sides never visible in the single input image. Section 6 concedes the geometries are not the same, and Section 10.3 reports failures for defective or rare models. The quantitative support is weak: Table 1 measures FID/KID between 360-degree renderings of 30 models and a generic real-car dataset, a distribution-level statistic that does not verify per-instance correspondence; LPIPS is computed only under the matched input pose. Table 6 reports Chamfer distance and volume IoU on ShapeNet retrieval, not on the KITTI-360/MVMC vehicles used in the main experiments, and does not test unseen-side fidelity. Consequently, the door-opening OOD result in Table 3 may reflect the gap between the rendered look-alike and the real vehicle, or the synthetic pipeline generally, rather than a genuine out-of-distribution state of the specific vehicle. The 'digital twin' terminology and the OOD conclusion therefore overstate what is demonstrated, though the disclosed limitation and failure cases show the authors are aware of the boundary.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"UrbanCAD proposes a retrieval-optimization pipeline for creating controllable, photorealistic 3D vehicle assets from a single urban image. Given one segmented vehicle image, the method retrieves a semantically similar CAD model from Objaverse using CLIP/OpenShape features, assigns part-aware material priors from Adobe's procedural material library by recognizing CAD part semantics with ControlNet and Grounded SAM, and then optimizes the albedo of metal and rubber materials via differentiable rendering. The resulting vehicle can be rendered in 360 degrees, edited at the part level, relit, transferred to other materials, and inserted into reconstructed urban backgrounds using fisheye-based HDR lighting estimation and 3D Gaussian Splatting. The paper evaluates photorealism with FID/KID/LPIPS against reconstruction and texturing baselines, and assesses downstream perception performance on in-distribution and out-of-distribution (door-opening) scenarios.","tokens_in":21669,"tokens_out":3935,"duration_ms":35485,"significance":"If the claims hold, UrbanCAD would be a practical and scalable alternative to both handcrafted simulator assets and neural reconstruction: it preserves part-level controllability while producing renderings that are competitive with or better than state-of-the-art single-view reconstruction and texturing baselines. The paper is honest about several limitations, including geometric mismatch and failure cases, and it provides a fairly complete system description with ablations, lighting estimation comparisons, and downstream perception experiments. The main gap is that the headline 'digital twin' and photorealism claims are supported by distribution-level metrics and by matched-pose LPIPS, rather than by per-instance geometric or novel-view fidelity on the actual vehicles used in the experiments. The system is nonetheless a meaningful step toward controllable urban simulation assets, provided the evidence is strengthened.","major_comments":[{"comment":"The headline photorealism comparison rests on FID and KID computed between 360-degree renderings of 30 retrieved CAD models and 1800 real car images from [67]. These are distribution-level metrics: they do not measure whether the rendered model matches the specific input vehicle, and they are reported without error bars, confidence intervals, or significance tests. The accompanying observation that Paint3D performs better on KID shows that the ranking is not uniform across metrics, so the claim that UrbanCAD 'outperforms baselines in terms of photorealism' needs per-instance evidence or at least repeated-seed and statistical validation.","section":"§5.2, Table 1 and §8.1"},{"comment":"The paper explicitly concedes in Section 6 that 'the geometries of our created CAD models are not the same as the vehicles in the input image,' and Section 10.3 lists defective or rare models as failure cases. Yet the central 'digital twin' claim and the out-of-distribution conclusion require unseen-side geometric fidelity. Table 6 reports Chamfer distance and volume IoU on a ShapeNet-to-Objaverse retrieval task, not on the KITTI-360 or MVMC vehicles used in the main experiments, so it does not validate the actual digital twins produced in the paper. A per-instance geometry evaluation on the evaluated datasets (for example, against LiDAR scans or multi-view masks where available) is needed to justify the 'digital twin' terminology and the safety-critical OOD claims.","section":"§6, §10.3, Table 6"},{"comment":"Material optimization is deliberately restricted: only the albedo of metal and rubber materials is optimized, glass is assigned without optimization, and roughness is fixed. This is a reasonable engineering choice given the single-view setting, but the LPIPS number in Table 1 is computed only under the matched input pose, so it cannot verify that the optimized materials are correct on unseen sides. The claim of 'photorealistic 360-degree rendering' is therefore not fully supported by the optimization objective, which operates on mean/variance, Gram-matrix, and masked RGB losses rather than pixel-aligned appearance.","section":"§3.3 and supplementary §7.5"},{"comment":"The out-of-distribution conclusion is based on 150 frames of five door-opening/closing scenarios and compares against the same vehicles with closed doors. Without a real-world OOD set (for example, actual images of vehicles with open doors) or an alternative synthetic OOD generator, the performance drop in Table 3 does not establish that the drop is caused by the realism of UrbanCAD's OOD data; it may be induced by any novel object configuration. I recommend adding a control experiment or tempering the claim to 'on our generated OOD data.'","section":"§5.3, Table 3"}],"minor_comments":[{"comment":"The heading 'Background Reconstruction and Compostion' contains a typo; it should read 'Composition.'","section":"§4.2"},{"comment":"The text says 'we report the LIPIS scores'; this should be 'LPIPS scores.'","section":"supplementary §8.1"},{"comment":"The material-index IOU threshold (0.5) and the loss weights in Eq. (9) are free parameters, but no sensitivity analysis is provided for either; a brief ablation or discussion would strengthen confidence in the method's robustness.","section":"supplementary §7.3 and §7.5"},{"comment":"The geometry-quality comparison reports Chamfer distance and volume IoU without specifying the alignment procedure, units, or normalization used for the ShapeNet-to-Objaverse retrieval, which makes the numbers difficult to interpret or reproduce.","section":"supplementary §10.4, Table 6"}],"recommendation":"major_revision","confidential_remarks":"The paper is a solid systems contribution with a clear application to autonomous driving simulation, and the authors are transparent about limitations. The main risk is that the 'digital twin' and OOD realism claims exceed the current evidence, which is distribution-level and matched-pose only. The requested per-instance geometry and novel-view validation is feasible with existing KITTI-360 multi-view/LiDAR data, so I view this as a major revision rather than a rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: UrbanCAD is a capable, well-integrated system for turning a single street image of a vehicle into an editable CAD asset with plausible materials and lighting. It delivers on the controllability-photorealism trade-off better than the reconstruction baselines it compares against, and the qualitative results are genuinely nice. But the paper overstates its case when it calls the output a 'digital twin': the retrieved CAD geometry is semantically matched, not geometrically faithful to the observed vehicle, and the quantitative support for the twin claim is distribution-level, not per-instance.\n\nWhat is new: the package as a whole. Semantic CAD retrieval through OpenShape, part-aware material prior retrieval by rendering the material index map, photorealizing it with ControlNet and labeling it with Grounded SAM, then optimizing only the albedo of metal and rubber while leaving glass and roughness fixed, is a sensible and novel combination. The fisheye-based spatially varying lighting estimate is a practical engineering choice, and the 3DGS background integration makes the system usable for simulator-style novel views. The downstream perception experiments, including the in-distribution vs. door-opening OOD comparison, are a step beyond most single-image reconstruction papers. The ablation and the failure cases in the supplement give an honest view of where the method breaks.\n\nWhere it is soft: the 'digital twin' language. Section 6 concedes the geometries are not the same as the input vehicle, and Section 10.3 lists failures for defective or rare models. On the quantitative side, Table 1 reports FID/KID over 360-degree renderings of 30 models against a generic car dataset, with no error bars; that tests realism of the rendering distribution, not fidelity of a particular twin to its input. LPIPS is computed only at the matched input pose, so unseen-side fidelity is never measured. The ShapeNet geometry numbers in Table 6 do not transfer to the KITTI-360/MVMC instances used in the main experiments. The OOD door-opening result could be explained by the synthetic look of the render, not by a genuine out-of-distribution state of that specific vehicle. None of this kills the contribution, but it means the strong claims should be toned down.\n\nThe citation pattern is fine: related work is covered, including CADSim and ACDC, and the authors are appropriately transparent about the manual steps.\n\nAudience: people working on autonomous driving simulation, controllable 3D assets, and inverse rendering. It deserves a serious referee; the system is real and the limitations are disclosed. A referee should ask for per-instance geometry evaluation on the actual test sequences, error bars on FID/KID, and ideally code release, but this should not be a desk reject.","headline":"UrbanCAD is a well-built system that advances the photorealism-controllability trade-off for single-image vehicle assets, but the 'digital twin' claim outruns the evidence, and the headline metrics are distribution-level rather than per-instance.","tokens_in":22220,"tokens_out":3003,"would_cite":true,"duration_ms":25874,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"UrbanCAD builds a photorealistic, part-controllable 3D vehicle from a single urban image by retrieving a handcrafted CAD model and optimizing only its material albedo, and shows that the result can populate driving simulators with rare…","keywords":["3D vehicle reconstruction","CAD retrieval","material optimization","driving simulation","photorealistic rendering","out-of-distribution scenarios","procedural materials","urban scene"],"falsifier":"Render an UrbanCAD twin from the viewpoint opposite the input photo and compare its silhouette and visible components (wheel arches, mirrors, window shape) against a LiDAR scan or dense multi-view capture of the same real vehicle instance; if the unseen-side geometry is visibly wrong or the silhouette IoU is low, the retrieval assumption fails. A targeted test is a rare vehicle type like a heavy-duty truck, which the paper's own failure cases indicate can already produce unsatisfactory matches.","tokens_in":21212,"feed_emoji":"🚗","tokens_out":7266,"duration_ms":58962,"temperature":0.7,"pith_summary":"UrbanCAD proposes a way to build a photorealistic, part-controllable 3D digital twin of a vehicle from a single urban photograph. Instead of reconstructing geometry from pixels, it retrieves the most similar handcrafted CAD model from a large free library and then optimizes only the appearance of its metal and rubber parts to match the photo, keeping the artist-designed geometry, glass, and roughness fixed. The claimed payoff is that the result is as editable as a game-engine asset while rendering convincingly from all 360 degrees, so it can be inserted into reconstructed street scenes and used to synthesize rare safety-critical situations like opened doors. The paper argues this closes the usual trade-off between photorealism and controllability in driving simulation, and demonstrates the point by showing that perception models keep accuracy on in-distribution renderings but degrade on the out-of-distribution scenarios the method can produce.","feed_headline":"One urban photo becomes a controllable 3D car","feed_subtitle":"Retrieved CAD geometry plus optimized materials lets simulators build rare, perception-breaking scenes.","key_machinery":"The central mechanism is a retrieval-optimization loop that treats geometry, material structure, and lighting as fixed expert priors and only tunes a small set of appearance parameters. Retrieval uses a joint image–3D encoder to find the CAD model most semantically similar to the segmented vehicle image; part-aware material assignment uses ControlNet to turn material-index renderings into realistic images that Grounded SAM can segment by component name; and material optimization converts procedural node graphs into albedo, normal, and roughness textures, then minimizes a masked statistical loss, a Gram-matrix VGG loss, and an RGB loss against the reference view. For scene insertion, fisheye images are stitched into panoramas, the sky is converted to HDR, and the nearest environment map is chosen for spatially varying lighting before compositing the vehicle over a 3DGS background.","core_discovery":"The central discovery, on the paper's own terms, is that a retrieval-optimization paradigm can preserve the fine-grained priors of handcrafted CAD models while still adapting to real observations. Given a segmented input vehicle, the method first retrieves a CAD model using a cross-modal encoder that aligns images and 3D shapes in a shared latent space, then matches its pose by DINO feature similarity. It assigns artist-made procedural materials to each part by recognizing semantic components with a vision-language segmentation model on ControlNet-augmented renderings, and finally optimizes only the albedo of metal and rubber material graphs through physically based differentiable rendering, leaving glass and roughness as fixed priors. The resulting models support 360-degree rendering, component editing, material transfer, and relighting, and can be inserted into 3D Gaussian Splatting backgrounds using fisheye-derived HDR environment maps.","pith_inferences":["The method's ceiling is set by library coverage: a vehicle type absent from the CAD collection will fall back to a poor retrieval, so extending the library with generated or procedurally varied CAD bodies is the most direct path to broader applicability.","The geometry-mismatch limitation is acknowledged by the authors; a natural follow-up is to add lightweight per-part deformation or scaling to the retrieved model so that silhouette and wheelbase match the observed vehicle before appearance optimization.","The spatially varying lighting approximation (nearest fisheye environment map) will degrade for vehicles far from the capture rig; estimating a continuous lighting field from the fisheye sequence is a testable extension."],"forward_implications":["Driving simulators could populate scenes with realistic, editable vehicles derived from ordinary street photos rather than manual modeling.","Rare safety-critical configurations such as open doors, collisions, or inverted vehicles become renderable assets that can be animated in a physics engine.","Perception models can be stress-tested on out-of-distribution data generated by the same pipeline that produces in-distribution data, exposing accuracy drops.","Because retrieval preserves part-disentangled geometry, material and component edits transfer directly between vehicles."],"supporting_citations":[{"why":"Supplies the free CAD model library whose handcrafted geometry priors are retrieved.","marker":"[15]"},{"why":"Supplies the handcrafted procedural material graphs used as appearance priors.","marker":"[1]"},{"why":"Provides the image–3D aligned encoder that performs semantic CAD retrieval.","marker":"[39]"},{"why":"Segments the target vehicle from the input urban image.","marker":"[29]"},{"why":"Recognizes semantic parts in augmented renderings for part-aware material assignment.","marker":"[52]"},{"why":"Augments material-index renderings into photorealistic images that make part segmentation reliable.","marker":"[78]"},{"why":"Converts procedural node graphs into differentiable programs, enabling material optimization.","marker":"[55]"},{"why":"Provides the part-level statistical and Gram-matrix losses and the procedural-material baseline.","marker":"[71]"}],"fun_headline_variants":["One image to a controllable 3D car","Single photo yields photoreal 3D vehicle","Retrieval-optimization builds 3D cars from one image","UrbanCAD: one urban photo to a fully editable 3D car"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the free CAD library already contains a model whose geometry and part layout are close enough to the observed vehicle that recoloring only the metal and rubber parts makes it look like the real vehicle from every angle, including the occluded sides never visible in the single input photo.","fun_headline_variants_meta":{"raw":{"variants":["One image to a controllable 3D car","Single photo yields photoreal 3D vehicle","Retrieval-optimization builds 3D cars from one image","UrbanCAD: one urban photo to a fully editable 3D car"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000509,"raw_usage":{"total_tokens":2500,"prompt_tokens":989,"completion_tokens":1511,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":605,"completion_tokens_details":{"reasoning_tokens":1442}},"tokens_in":605,"tokens_out":1511,"duration_ms":11592,"temperature":1.0,"reasoning_tokens":1442,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T10:19:37.709285+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Render an UrbanCAD twin from the viewpoint opposite the input photo and compare its silhouette and visible components (wheel arches, mirrors, window shape) against a LiDAR scan or dense multi-view capture of the same real vehicle instance; if the unseen-side geometry is visibly wrong or the silhouette IoU is low, the retrieval assumption fails. A targeted test is a rare vehicle type like a heavy-duty truck, which the paper's own failure cases indicate can already produce unsatisfactory matches.","supporting_citations":[{"cited_title":"Objaverse: A universe of annotated 3d objects","cited_arxiv_id":null,"evidence_quote":"Supplies the free CAD model library whose handcrafted geometry priors are retrieved."},{"cited_title":"Real-time neural rasterization for large scenes","cited_arxiv_id":null,"evidence_quote":"Provides the image–3D aligned encoder that performs semantic CAD retrieval."},{"cited_title":"Girshick, Carsten Rother, and Piotr Doll ´ar","cited_arxiv_id":null,"evidence_quote":"Segments the target vehicle from the input urban image."},{"cited_title":"Urban radiance fields","cited_arxiv_id":null,"evidence_quote":"Recognizes semantic parts in augmented renderings for part-aware material assignment."},{"cited_title":"Airsim: High-fidelity visual and physical simulation for autonomous vehicles","cited_arxiv_id":null,"evidence_quote":"Converts procedural node graphs into differentiable programs, enabling material optimization."}],"review_version":1}