{"id":"6c659fa2-0fa6-4fa8-8b4e-35127067c253","arxiv_id":"2412.17226","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"A two-stage LiDAR diffusion model generates foreground objects from text and 3D box conditions and then completes the scene, improving object fidelity and downstream 3D detection.","lead":"OLiDM generates a LiDAR scene by first synthesizing foreground objects from text and 3D box prompts, then diffusing the background around those objects. It targets controllable simulation and 3D detection training for self-driving vehicles.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Object-level superiority on KITTI-360 is not yet established: the object training database and the object-evaluation metric are both built from the same SECOND detector, so the high CD/JSD/#Box scores may reflect training-to-the-evaluator rather than fidelity to real objects.","rationale":"Read in good faith, OLiDM's central engineering claim is a two-stage generation pipeline, and the nuScenes detection augmentation results in Table 4 provide independent evidence that the framework has practical value. However, the KITTI-360 object-level results, which anchor the 'object-aware' headline, are constructed and measured with the same SECOND detector. Since the object generator is trained to denoise SECOND-derived object patches, metrics computed by feeding generated scenes through SECOND reward reproducing that detector's inductive biases, not necessarily human-annotated objects. The paper's 'real benchmark' #Box = 9.06 is itself a SECOND count, so the target and the evaluator share the same error profile. This is a circularity/leakage risk specifically for the object-level superiority claim; scene-level FPD/JSD and downstream augmentation are less affected. Missing seeds and the abstract's 57.47% wording are additional quality issues but are not the load-bearing one. A concrete re-evaluation against ground-truth KITTI-360 annotations and a held-out detector would resolve the concern; absent that, the object-level claim should remain conditional.","tokens_in":14696,"tokens_out":7329,"duration_ms":70762,"concrete_test":"Recompute the object-level evaluation on KITTI-360 (Table 2) using official 3D ground-truth boxes instead of SECOND detections as the reference object set, and detect generated objects with a held-out detector that was not used to build the object database (e.g., CenterPoint trained on nuScenes or another KITTI-based model). If OLiDM's CD/SS/JSD/#Box advantage over LiDARGen, UltraLiDAR, and R2DM persists against ground-truth boxes and the held-out detector, the circularity concern is settled. If the advantage shrinks or reverses, the object-level claim is detector-specific.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing assumption is that SECOND detections define \"real foreground objects\" on KITTI-360. Appendix A builds the KITTI-360 object training set by running a SECOND detector (trained on KITTI) on KITTI-360, and Table 2 and Fig. 1 also use SECOND to count and evaluate generated objects. The reference value #Box = 9.06 is itself the average SECOND-detected box count in real scenes. Because OLiDM's object denoiser is trained to reproduce SECOND's detected object patches, its high object-level CD/SS/JSD/#Box numbers may partly reflect overfitting to SECOND's detection biases (typical box shapes, easy poses, false-positive patterns) rather than closeness to human-annotated objects. Baselines such as LiDARGen and R2DM were not trained on this detector-derived object distribution, so the comparison is not symmetric. This does not invalidate the scene-level FPD/JSD results or the nuScenes detection augmentation, which are separate evidence, but it weakens the headline claim of high-fidelity object generation on KITTI-360.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"OLiDM proposes an object-aware LiDAR diffusion framework that generates foreground objects first, conditioned on CLIP text embeddings and Fourier-encoded 3D boxes, and then generates a full scene from a range image using a scene controller and an Object Semantic Alignment (OSA) loss. Experiments on KITTI-360 report strong scene-level FPD/JSD, object-level Chamfer/JSD/#Box numbers, sparse-to-dense completion results, and downstream nuScenes detection improvements over GT-Aug and LiDAR-Aug. The paper's central novelty is the object-first, scene-second progressive generation, which is reasonable and potentially useful, but the object-level evaluation is weakened by a training/evaluation circularity and the abstract contains quantitative claims that do not match the reported tables.","tokens_in":14935,"tokens_out":5978,"duration_ms":55456,"significance":"If the claims hold, OLiDM would be a practical contribution because it is the first LiDAR generator in this line that separates foreground-object generation from background-scene generation and provides initial object annotations for downstream tasks. The scene-level FPD/JSD improvements and the nuScenes detection gains are externally benchmarked and deserve credit: the FPD improvement over UltraLiDAR (25.1 to 7.60 in Table 1) is large, and the CenterPoint mAP gain of 2.9 points over a strong baseline is a concrete, falsifiable result. However, the headline claim of high-fidelity object generation on KITTI-360 is not yet established, because the object training database and the object-level evaluation both rely on the same SECOND detector. The paper also omits error bars and has abstract numbers inconsistent with Tables 3 and 4. The architecture is original, and the OSA idea is clearly presented, but the experimental validation needs to be strengthened before the central object-level claim can be accepted.","major_comments":[{"comment":"The object-level evaluation is circular. Appendix A states that the KITTI-360 object training database is built by running SECOND (trained on KITTI) on KITTI-360, and Section 4.2 states that the object-level evaluation uses the same detector family to identify and count generated objects, with the reference value #Box = 9.06 being the average SECOND-detected box count in real scenes. Consequently, the low CD/JSD values and the #Box count in Table 2 may reflect the model reproducing SECOND's detection biases (typical box shapes, easy poses, false-positive patterns) rather than geometric or semantic fidelity to human-annotated objects. The comparison is also asymmetric, because LiDARGen and R2DM were not trained on this detector-derived object distribution. Please re-evaluate object-level metrics against KITTI-360 ground-truth annotations or an independent detector, and report, for example, point-to-surface distance to annotated 3D boxes and per-category recall/precision of detected generated objects.","section":"§4.2, Appendix A"},{"comment":"The abstract's quantitative claims do not match the reported tables. The claim of a '57.47% increase in semantic IoU' in sparse-to-dense completion is not a relative increase: Table 3 reports semantic IoU 22.46 for LiDARGen and 79.93 for OLiDM, which is an absolute increase of 57.47 percentage points and a relative increase of about 256%. The claim of '2.4% in mAP and 1.9% in NDS' is also not directly supported by Table 4, where PointPillars shows +1.9 mAP / +2.8 NDS and CenterPoint shows +2.9 mAP / +2.7 NDS; neither detector gives 1.9 NDS, and the mAP value is only the average of two detectors. Please state the exact aggregation rule or correct the abstract.","section":"Abstract, Tables 3 and 4"},{"comment":"No random seeds, number of runs, or confidence intervals are reported for any quantitative result. Diffusion sampling is stochastic, and several reported differences are small (e.g., JSD 0.49 vs. 0.42 in Table 1; NDS differences in Table 4), so without repeated runs it is unclear whether these differences are significant. Please run at least three seeds for the main generation and detection experiments and report mean and standard deviation, or an equivalent statistical statement.","section":"Tables 1–4"},{"comment":"The description of the object-level reference set is incomplete. The text reports CD, SS, and JSD for 'foreground objects' and states that 'we sample 1000 objects for the evaluation,' but it does not state what the reference point clouds are, how the generated objects are extracted (e.g., by cropping detected boxes), or how each generated object is matched to a real object for the Chamfer Distance computation. If the real reference objects are also the SECOND-extracted patches used in Appendix A, then the CD and SS numbers are not an independent measure of fidelity. Please specify the full evaluation protocol.","section":"§4.2, Table 2"}],"minor_comments":[{"comment":"The sentence 'R2DM is our baseline, with range image inputs (Sec.3.1 w = 1024)' is unclear: R2DM is a prior method, and Section 3.1 does not define the symbol w. If R2DM is reimplemented, state the reimplementation details and the source of the numbers in Table 1.","section":"§4.1"},{"comment":"The OSA loss in Eq. (9) is written without the scene controller or the object-conditioned features used in Eq. (8), but the text says Eq. (9) is used jointly with Eq. (8). Please clarify whether Eq. (9) is computed with the same conditioned denoiser and specify how the mask M is obtained during training (e.g., from ground-truth object boxes or from detector outputs).","section":"§3.3, Eq. (9)"},{"comment":"The text refers to 'fig. 17' when first introducing the OSA motivation, but Figure 17 is in the appendix and is not called out until later. Please renumber or move the reference so the figure is introduced at the appropriate point.","section":"§3.3"},{"comment":"The header 'Reflectence' appears to be a typo for 'Reflectance'.","section":"Table 3"},{"comment":"The notation OLiDM, OLiDM−, OLiDMT, and OLiDMB is used in Table 2, but the caption only explains the abbreviations loosely. Please define each row explicitly in the caption or in the text before the table.","section":"§4.1 and Table 2"}],"recommendation":"major_revision","confidential_remarks":"The central architectural idea is sound and the scene-level and downstream results are potentially publishable, but the object-level fidelity claim rests on a training/evaluation circularity that must be resolved with independent annotations or a different detector. The abstract also needs to be aligned with the tables, and the lack of error bars is a serious rigor issue. I would ask the authors to provide the evaluation code and seeds as part of the revision, since the claimed object-level metrics are unusually strong."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a genuinely useful contribution to LiDAR simulation, with an architectural novelty that the baselines don’t have. The object-first, scene-second pipeline, the ControlNet-style controller on range images, and the category-masked OSA loss are new relative to LiDARGen, R2DM, UltraLiDAR, and LiDM. The paper also introduces object-level evaluation metrics for LiDAR generation, which is a helpful addition. The experimental scope is broad—scene generation, sparse-to-dense completion, controllable generation, and detection augmentation on nuScenes—and the nuScenes downstream numbers are externally benchmarked and plausible. The scene-level FPD/JSD improvements on KITTI-360 stand on their own.\n\nSoft spots, in order of seriousness. First, the KITTI-360 object database is built from SECOND detections (Appendix A), and the object-level evaluation also uses SECOND to count and compare objects (Table 2, Fig. 1). The reference #Box = 9.06 is itself the average SECOND-detected count in real scenes. So the high CD/SS/#Box scores partly reflect matching SECOND’s detection biases, not necessarily fidelity to human-annotated objects. This doesn’t sink the paper—the nuScenes results and scene-level metrics are separate evidence—but it directly weakens the “high-fidelity object generation” claim as stated. Second, the abstract numbers don’t align with Tables 3 and 4: the 57.47% semantic IoU gain is an absolute percentage-point increase over LiDARGen, not a relative increase, and the 1.9% NDS figure doesn’t appear in Table 4. That’s sloppy but fixable. Third, no seeds or error bars are reported anywhere, so we can’t judge variance. Fourth, the OSA inference pathway is underspecified—it’s clear during training but unclear how the alignment is applied at generation time.\n\nOverall, the core engineering idea is sound and the paper is worth engaging with. It needs a revision that disentangles the KITTI-360 object evaluation from the training-data detector, adds variance estimates, and makes the numbers consistent. I’d send it to peer review—the architecture and the downstream augmentation results deserve referee time—and I’d cite it if I worked in this area.","headline":"Solid engineering contribution with a real two-stage object-to-scene architecture, but the KITTI-360 object-fidelity claim is undermined by using the same SECOND detector for both training data and evaluation.","tokens_in":15486,"tokens_out":2960,"would_cite":true,"duration_ms":24830,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"OLiDM generates LiDAR scenes by first generating controllable foreground objects, then conditioning the scene on them, improving fidelity and downstream detection.","keywords":["LiDAR generation","diffusion models","object-aware generation","autonomous driving","range image","3D object detection","data augmentation","point cloud"],"falsifier":"Train an independent object detector, or use human labels, to count and compare objects in OLiDM-generated KITTI-360 scenes against the real-data benchmark of 9.06 boxes per scene; if the independent count does not approach that number while SECOND does, the claimed object fidelity is an artifact of the evaluation detector. Alternatively, fine-tune a detector on OLiDM-only synthetic data and evaluate on real KITTI-360; poor real-world performance would indicate the generated objects do not capture the distribution the paper claims.","tokens_in":14469,"feed_emoji":"🚗","tokens_out":5070,"duration_ms":36263,"temperature":0.7,"pith_summary":"The paper introduces OLiDM, a diffusion framework that generates LiDAR point clouds for autonomous driving at two levels: individual foreground objects and full scenes. Its central claim is that by generating foreground objects first, with text and 3D-box conditions, and then using those objects to guide scene generation, the model produces LiDAR data that is closer to real sensor output than prior whole-scene generators. On KITTI-360 it reports lower FPD and JSD than LiDARGen, R2DM, and UltraLiDAR, and on nuScenes it reports that augmenting training data with its synthetic objects raises 3D detector mAP and NDS. The reason to care is that high-fidelity, controllable synthetic LiDAR could reduce costly real-data collection and annotation for self-driving perception.","feed_headline":"Object-first LiDAR diffusion beats prior generators by 17.5 FPD","feed_subtitle":"New model produces controllable foreground objects and scenes, improving 3D detector training on nuScenes.","key_machinery":"The central machinery is a two-stage diffusion pipeline: an object denoiser, a voxel-based diffusion transformer, that takes CLIP text embeddings and a Fourier-embedded 3D box as conditions, and a scene-level U-Net denoiser over range images, steered by a zero-convolution scene controller that injects the generated object range image. OSA computes category-wise binary masks from the object conditions and uses them to form c-channel masked range images, with an extra diffusion loss on those masked channels to align features within each semantic subspace.","core_discovery":"OLiDM's central claim is that foreground objects, not whole scenes, should be the unit of LiDAR generation. The Object-Scene Progressive Generation (OPG) module first denoises object point clouds conditioned on a text description and a 3D bounding box, then projects those objects into a range image that a scene controller feeds into the scene denoiser. The Object Semantic Alignment (OSA) module adds per-category masked channels to the range image so the diffusion loss is balanced across semantic subspaces. The paper reports that this object-first, scene-second design yields the best FPD and JSD on KITTI-360, with OLiDM exceeding UltraLiDAR by 17.5 in FPD, and that the generated objects improve downstream 3D detection on nuScenes by 2.4% mAP and 1.9% NDS.","pith_inferences":["The object-first design could be applied to other sensor modalities such as radar or event cameras, where foreground-background imbalance also exists, though the paper does not test this.","Because object crops are generated before the scene, the framework could, in principle, be extended to insert a requested number of objects per category at desired locations, enabling stress-testing of perception systems; the paper only demonstrates uniform and prompted sampling.","The reliance on SECOND detections for the KITTI-360 object database suggests that using ground-truth annotations instead could shift both training and evaluation, and it is an open question how the reported object-level metrics would change under an independent detector."],"forward_implications":["If correct, OLiDM gives a single pipeline that outputs both a LiDAR scene and the 3D object annotations for the objects it placed, since the objects are generated first.","It enables user control of object category, description, position, and size, which can be used to create rare or corner-case scenarios for perception testing.","The reported sparse-to-dense completion gains suggest the same conditional mechanism can upsample low-beam LiDAR to high-beam density with better semantic preservation.","The downstream improvement over GT-Aug indicates that object-first synthetic data can serve as a data augmentation source for 3D detectors."],"supporting_citations":[{"why":"Supplies the range-image diffusion baseline and the FPD metric used for scene-level comparison.","marker":"(Zyrianov, Zhu, and Wang 2022)"},{"why":"Provides the voxel/BEV-generation baseline that OLiDM compares against and improves upon in FPD.","marker":"(Xiong et al. 2023)"},{"why":"Provides the R2DM baseline for range-image DDPM generation and sparse-to-dense completion.","marker":"(Nakashima and Kurazume 2023)"},{"why":"Supplies the DiT-3D voxelized diffusion transformer architecture used to build the object denoiser.","marker":"(Mo et al. 2024)"},{"why":"Supplies the SECOND detector used to build the KITTI-360 object database and to count and assess foreground objects in generated scenes.","marker":"(Yan, Mao, and Li 2018)"},{"why":"Provides the zero-convolution control mechanism repurposed as the scene controller.","marker":"(Zhang, Rao, and Agrawala 2023)"}],"fun_headline_variants":["Object-first LiDAR diffusion beats prior generators by 17.5 FPD","LiDAR diffusion that generates foreground objects first boosts detection","OLiDM: object-aware LiDAR generation with semantic alignment","Generate LiDAR objects then scenes: OLiDM ups detector mAP by 2.4%","Object-scene progressive LiDAR diffusion improves semantic IoU by 57%"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The premise that SECOND detector outputs are an unbiased definition of real foreground objects is load-bearing: if SECOND systematically misses or double-counts objects, both the training database of KITTI-360 objects and the object-level evaluation are biased in the same direction.","fun_headline_variants_meta":{"raw":{"variants":["Object-first LiDAR diffusion beats prior generators by 17.5 FPD","LiDAR diffusion that generates foreground objects first boosts detection","OLiDM: object-aware LiDAR generation with semantic alignment","Generate LiDAR objects then scenes: OLiDM ups detector mAP by 2.4%","Object-scene progressive LiDAR diffusion improves semantic IoU by 57%"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000448,"raw_usage":{"total_tokens":2309,"prompt_tokens":1045,"completion_tokens":1264,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":661,"completion_tokens_details":{"reasoning_tokens":1166}},"tokens_in":661,"tokens_out":1264,"duration_ms":10785,"temperature":1.0,"reasoning_tokens":1166,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T05:41:47.597502+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train an independent object detector, or use human labels, to count and compare objects in OLiDM-generated KITTI-360 scenes against the real-data benchmark of 9.06 boxes per scene; if the independent count does not approach that number while SECOND does, the claimed object fidelity is an artifact of the evaluation detector. Alternatively, fine-tune a detector on OLiDM-only synthetic data and evaluate on real KITTI-360; poor real-world performance would indicate the generated objects do not capture the distribution the paper claims.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the range-image diffusion baseline and the FPD metric used for scene-level comparison."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the voxel/BEV-generation baseline that OLiDM compares against and improves upon in FPD."}],"review_version":1}