{"id":"ed22fb20-4dbd-49a5-8647-494f617136e3","arxiv_id":"2501.04950","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"MORDA, a synthetic dataset of South Korean digital twins with nuScenes-compatible sensors and labels, improves 2D/3D object detection on the unseen AI-Hub South Korea dataset when added to nuScenes training, while preserving nuScenes performance.","lead":"The paper introduces MORDA, a synthetic driving dataset that recreates South Korean road environments with digital-twin maps and matches nuScenes sensor and labeling setup. Adding MORDA to nuScenes training improves 2D and 3D detection on an unseen real South Korean dataset while preserving nuScenes performance.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Target-domain gains on AI-Hub are not shown to require the Korea-specific digital twins; no other synthetic dataset is evaluated on AI-Hub, so the central claim may reduce to generic synthetic-data augmentation.","rationale":"Read in good faith, the paper is a practical demonstration: adding MORDA to nuScenes training consistently improves AI-Hub mAP across 2D and 3D detectors while leaving nuScenes mAP roughly flat. The reported effects are large for several models, and the dataset construction (sensor suite, labeling rules, Korean digital twins) is described in detail. The concern is not internal inconsistency; the experiments support the empirical claim. However, the causal claim implied by the title and intro -- that the Korea-specific, target-mimicking design is what makes MORDA useful -- rests on an untested comparison. Table IV only shows MORDA outperforming other synthetic datasets on the source domain, which does not establish that the AI-Hub improvement requires Korean geography. Without a target-domain synthetic baseline, the central mechanism cannot be separated from generic synthetic-data augmentation. This is the weakest load-bearing assumption: if generic synthetic data gives the same AI-Hub gains, the proposed 'digital-twin fusion' methodology is not necessary, and the paper's contribution reduces to 'more synthetic data helps.' The proposed test settles this directly. The reader's weakest-assumption analysis points to the same gap, so agreement is 'agree'; the verdict remains CONDITIONAL pending that test.","tokens_in":12354,"tokens_out":6398,"duration_ms":61431,"concrete_test":"Use the Faster-RCNN protocol from Section V-C to train on nuScenes plus each of VKITTI2, SYNTHIA-AL, and SHIFT (matched to MORDA's 3,700-frame scale by subsampling if needed), then evaluate on AI-Hub under the exact Section V-B.2 pre-processing (front camera, 1600x900 crop, same four classes, same pseudo-2D labels). If any generic synthetic dataset attains an AI-Hub mAP gain comparable to MORDA's +6.35 (e.g., within 2 points), the Korea-specific digital-twin maps are not the load-bearing component.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that MORDA's value comes from fusing nuScenes sensor/labeling characteristics with Korean target-domain geography. To establish this, the AI-Hub improvements in Table II must be attributable to the Korean digital-twin maps rather than to any additional synthetic data. The paper never performs that comparison on the target domain: Table IV benchmarks VKITTI2, SYNTHIA-AL, and SHIFT against MORDA only on the nuScenes validation split, not on AI-Hub. Since MORDA adds 3.7K annotated frames and the other synthetic datasets add comparable or larger amounts of data, the large AI-Hub gains (e.g., +6.35 mAP for Faster-RCNN, +5.71 for SSN, +18.16 and +22.17 for CenterPoint voxel models in Table II) could in principle be produced by generic synthetic augmentation that regularizes the detector or increases object-shape diversity, without any contribution from the Korean maps. The paper explicitly defers this attribution to future work in Section VI ('identify which characteristics implemented in MORDA contributed to performance stability on DT_rg'). Thus the load-bearing assumption -- that the Korea-specific digital twins are what drives the transfer -- is untested. A secondary weakness is the absence of multiple seeds/error bars, which weakens the 'preserved' claim on nuScenes, but the missing target-domain synthetic baseline is the more fundamental gap.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces MORDA, a synthetic driving dataset built by combining digital-twin maps of South Korea with a reproduction of the nuScenes sensor suite, object categories, and 3D bounding-box annotation rules. The authors train 2D camera-based and 3D LiDAR-based detectors on nuScenes alone or nuScenes plus MORDA, then evaluate on the nuScenes validation set and on the AI-Hub dataset collected in South Korea, which is never used for training. They report that adding MORDA improves mean average precision on AI-Hub across all tested detectors (e.g., +6.35 mAP for Faster-RCNN, +5.71 for SSN, +22.17 for CenterPoint-Voxel-0.075m) while nuScenes mAP/NDS are maintained or slightly improved. A comparison with VKITTI2, SYNTHIA-AL, and SHIFT is performed only on nuScenes validation.","tokens_in":12598,"tokens_out":4484,"duration_ms":42751,"significance":"If the central claim survives, the paper offers a cost-effective alternative to collecting and labeling real data in a new deployment region, and MORDA is a large, openly documented synthetic dataset with annotations for multiple tasks. The strengths are the breadth of detectors (2D and 3D, multiple architectures), the faithful reproduction of nuScenes sensor geometry and label conventions, and the clear demonstration that adding synthetic data can rescue detectors from near-zero performance under a large domain shift (CenterPoint). The main weakness is that the target-domain gains are not shown to require the Korean digital twins, because no generic synthetic dataset is evaluated on AI-Hub.","major_comments":[{"comment":"The paper's central claim is that MORDA's value comes from fusing nuScenes characteristics with Korean target-domain geography, but the only synthetic-dataset comparison is conducted on nuScenes validation, not on AI-Hub. Table IV shows that VKITTI2, SYNTHIA-AL, and SHIFT give mAP gains of -0.3, +0.1, and +0.4 on nuScenes, yet none of these datasets is tested on AI-Hub. Therefore, the large AI-Hub improvements in Table II (e.g., +18.16 and +22.17 mAP for CenterPoint voxel models) could in principle be obtained by any additional synthetic data that regularizes the detector, without any contribution from the Korean digital-twin maps. The authors themselves defer attribution to future work in Section VI ('identify which characteristics implemented in MORDA contributed to performance stability on DT_rg'). Please add a control experiment that trains on nuScenes plus a non-Korean synthetic dataset (or synthetic data generated from non-Korean maps) and evaluates on AI-Hub, to isolate the effect of the target-specific geography.","section":"Section V-C, Table IV; Section VI"},{"comment":"All results are from a single training run per condition, and no error bars or significance tests are provided. The nuScenes 'retained or slightly enhanced' claim relies on small differences (e.g., +0.5 mAP for Faster-RCNN, +0.51 for PointPillars), which could be within run-to-run variance. Please report mean and standard deviation over at least three seeds for the main comparisons, or otherwise justify why single runs are sufficient.","section":"Section V-A.4 and Table II"}],"minor_comments":[{"comment":"OPV2V is a vehicle-to-vehicle communication dataset rather than a single-vehicle driving dataset; the comparison table would benefit from a note clarifying this distinction.","section":"Table I"},{"comment":"The sentence 'we use the implementation of the mentioned networks in MMDetection3D' should say 'we use the implementations'.","section":"Section V-A.2"},{"comment":"The red dotted ellipse mentioned in the text is not clearly visible in grayscale; please increase its contrast or add a zoomed inset.","section":"Fig. 4"},{"comment":"The term 'A Vs' is inconsistently spaced; consider using 'AVs' for uniformity.","section":"Throughout"},{"comment":"The dataset is announced via a GitLab page; please state the license and whether the generation code and assets will be released publicly.","section":"Dataset access"}],"recommendation":"major_revision","confidential_remarks":"This is an industry-led dataset contribution from MORAI Inc., and the proposed dataset is generated with the company's own simulator. The missing synthetic-control experiment on AI-Hub is the key technical gap; if the authors add a non-Korean synthetic baseline or an ablation with non-Korean maps, the central attribution claim would be substantially strengthened. I also recommend requesting error bars for the main comparisons. The paper is within the scope of the journal, but the dataset's licensing and availability conditions should be clarified."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"MORDA is a real contribution: a synthetic dataset that reproduces nuScenes' sensor suite, 10-class labeling, and annotation rules inside digital twins of South Korea. That combination is new, and the empirical core is solid. Across four detectors (Faster-RCNN, PointPillars, SSN, CenterPoint variants), adding MORDA to nuScenes training improves mAP on the unseen AI-Hub dataset, sometimes dramatically, while nuScenes validation performance is roughly preserved or slightly better. Those are consistent, useful results for anyone doing AV perception adaptation, and the dataset itself—publicly referenced, with formatted conversion to nuScenes—seems like a practical asset. The paper honestly describes its own limits, including deferring attribution to future work in Section VI.\n\nThe main soft spot is exactly what the stress test flags: the target-domain gains are not shown to require the Korean maps. Table IV compares MORDA against VKITTI2, SYNTHIA-AL, and SHIFT only on nuScenes validation, not on AI-Hub. So the +6.35 mAP for Faster-RCNN and the +18/+22 mAP for CenterPoint voxel models could plausibly come from generic synthetic-data augmentation—more object-shape diversity, regularization, or just more frames—rather than from the Korea-specific digital twins. The paper itself concedes this in Section VI, which is honest but means the central causal claim is untested. That is the load-bearing weakness, and it is real. I would not call it fatal, because the dataset's value stands even if the mechanism is partly generic, but it does need to be addressed before the claim is strong.\n\nA secondary issue is the absence of multiple seeds or error bars. The \"preserved performance\" claim on nuScenes rests on single runs, and some class-wise AP differences are small enough that noise could flip them. That is a minor concern compared to the missing target-domain control, but worth mentioning.\n\nOverall: this is a useful dataset paper with an honest but incomplete evaluation. It deserves a serious referee. I would send it to peer review, and in revision push the authors to run the comparison synthetic datasets on AI-Hub and add multiple seeds. For a reader working on domain adaptation or synthetic data for AVs, this is worth a read; I would cite it if I needed a nuScenes-compatible synthetic dataset with target-region geography.","headline":"A genuinely new synthetic dataset that improves detection on an unseen target domain, but the paper never isolates whether the Korea-specific digital twins or just extra synthetic data drive the gain.","tokens_in":768,"tokens_out":715,"would_cite":true,"duration_ms":25037,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A synthetic-fusion dataset built from South Korean digital twins and nuScenes-style sensors lets object detectors generalize to an unseen real target without losing source-domain performance.","keywords":["MORDA dataset","synthetic driving dataset","domain adaptation","digital twin","object detection","autonomous driving","nuScenes","sim-to-real transfer"],"falsifier":"Train the same detectors on nuScenes plus synthetic data generated with identical sensors and annotation rules but digital-twin maps of a different country; if the AI-Hub mAP gain is unchanged, the target-specific Korean maps are not what carries the result. A complementary check is to evaluate the same models on a second unseen target city outside Korea, where MORDA should confer no advantage if the effect is genuinely target-specific.","tokens_in":12141,"feed_emoji":"🚗","tokens_out":8037,"duration_ms":70595,"temperature":0.7,"pith_summary":"The paper tries to show that a synthetic driving dataset can prepare object detectors for a real region they have never seen, without forgetting the region they were trained on. It builds MORDA by putting nuScenes-style sensors, annotation rules, and object classes into a simulator stocked with digital-twin maps of several South Korean regions. Detectors trained on plain nuScenes plus MORDA, with no domain-adaptation trick, score substantially better on the unseen real South Korean AI-Hub dataset across 2D camera and 3D LiDAR detectors, while nuScenes scores stay flat or rise. The authors also report that MORDA beats other synthetic datasets for 2D detection with far fewer frames. If the result holds, simulating the target region plus matching the source sensor and labeling pipeline could be a cheap substitute for recollecting and relabeling data in every new deployment country.","feed_headline":"Synthetic Korean roads lift detection on unseen real data","feed_subtitle":"nuScenes-based detectors gain up to 22 mAP points on never-seen South Korean roads without losing source accuracy.","key_machinery":"The load-bearing object is the synthetic-fusion domain $D^{Src+Trg}_{Syn}$, a simulated world in which the target region's digital-twin maps are populated with the source dataset's sensor suite and labeling scheme. In MORDA this means nuScenes' six-camera and 32-beam LiDAR geometry, ten detection classes, and 3D-box conventions, including articulated bus and truck-trailer rules, are reproduced inside simulator maps of one highway and three urban South Korean locations, with static and dynamic scenes recorded at 20 Hz. The dataset is then converted to nuScenes format so that existing detectors can ingest it by simple concatenation. This pairing of target geography with source sensor and label fidelity is what the paper claims carries the adaptation, and it is what distinguishes MORDA from synthetic datasets built without a specific real target in mind.","core_discovery":"On the paper's own terms, the central discovery is that a synthetic-fusion domain, a simulator world that blends the target region's geography with the source dataset's sensing and labeling protocol, transfers to an unseen real target domain. Using nuScenes as source and South Korea as target, the paper creates MORDA, the Mixture Of Real-domain characteristics for synthetic-data-assisted Domain Adaptation: 87 scenes, 37K frames, 1.6M 3D boxes, six 1600x900 cameras and one 32-beam LiDAR following nuScenes, and digital-twin maps of one highway and three urban South Korean areas. Training simple detectors on nuScenes plus MORDA raises AI-Hub mAP from 13.35 to 19.7 for Faster R-CNN and from 7.20 to 22.93 for CenterPoint with a 0.1m voxel backbone, while nuScenes mAP improves slightly; similar gains appear for PointPillars and SSN. The paper interprets this as evidence that the simulator provides a preview of the target domain and acts as a regularizer against generalization failure.","pith_inferences":["If the target-specific digital-twin maps are the active ingredient, the same pipeline should transfer to other source-target pairs; building a non-Korean twin with identical sensors and labels and comparing AI-Hub gains would test this directly.","The AI-Hub evaluation uses converted labels and pseudo-2D boxes derived from panoptic segmentation, so part of the reported gain could reflect alignment with the converted label distribution rather than raw perception quality; a small re-annotation study would separate these.","MORDA's heavy representation of rare classes such as trucks, construction vehicles, and trailers may be doing much of the work; an ablation that matches nuScenes' class distribution in the simulator would reveal whether class balancing or geographic fidelity drives the gain.","The method suggests a general principle: synthetic data's value for domain adaptation comes from matching the source annotation protocol and the target environment simultaneously, not from photorealism alone."],"forward_implications":["Adding MORDA to nuScenes training improves 2D and 3D detection on the unseen AI-Hub target across all tested architectures, with the largest gains where baseline models collapse under domain shift.","Source-domain performance is retained or slightly improved, so the method does not trade away the original deployment domain.","MORDA outperforms VKITTI2, SYNTHIA-AL, and SHIFT for 2D detection with fewer frames, suggesting target-matched synthetic data is more efficient than larger generic synthetic data.","Because only simple concatenation is used, combining MORDA with existing unsupervised domain adaptation methods is a direct next step that could yield further gains.","The recipe, if it generalizes, gives AV developers a data-acquisition path to a new region that avoids dispatching a sensor vehicle there for labeling."],"supporting_citations":[{"why":"Defines the real-source domain, its sensor suite, ten detection classes, labeling rules, and the mAP and NDS metrics used throughout the paper.","marker":"[19]"},{"why":"Provides the simulator environment, digital-twin maps, traffic generation, and ground-truth labeling used to create MORDA.","marker":"[30]"},{"why":"Supplies one of the LiDAR 3D detectors used to evaluate MORDA's effect on nuScenes and AI-Hub.","marker":"[4]"},{"why":"Supplies a LiDAR 3D detector used in the main experiments, where MORDA yields a 5.7 mAP gain on AI-Hub.","marker":"[5]"},{"why":"Supplies the 3D detector family whose baseline collapses under domain shift and where MORDA produces the largest gains.","marker":"[6]"},{"why":"Supplies the 2D detector used in the main comparison and in the synthetic-dataset benchmark against VKITTI2, SYNTHIA-AL, and SHIFT.","marker":"[32]"},{"why":"Provides the closest prior synthetic dataset that clones a real dataset, used as a comparison point for MORDA.","marker":"[14]"},{"why":"Provides a compared synthetic dataset that reaches similar 2D mAP but with roughly eight times more frames than MORDA.","marker":"[17]"},{"why":"Provides another compared synthetic dataset in the 2D detection benchmark.","marker":"[16]"}],"fun_headline_variants":["Synthetic fusion domain boosts detection on unseen Korean roads","MORDA synthetic set lifts mAP on unseen South Korean data","Simulator-blended data helps detectors adapt to new regions","Synthetic digital twins of Korea improve cross-domain detection"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the digital-twin maps of South Korea used to build MORDA faithfully reproduce the visual and geometric properties of the real target that matter for object detection, even though the real AI-Hub data were not used to construct them.","fun_headline_variants_meta":{"raw":{"variants":["Synthetic fusion domain boosts detection on unseen Korean roads","MORDA synthetic set lifts mAP on unseen South Korean data","Simulator-blended data helps detectors adapt to new regions","Synthetic digital twins of Korea improve cross-domain detection"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000719,"raw_usage":{"total_tokens":3303,"prompt_tokens":1095,"completion_tokens":2208,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":711,"completion_tokens_details":{"reasoning_tokens":2141}},"tokens_in":711,"tokens_out":2208,"duration_ms":17378,"temperature":1.0,"reasoning_tokens":2141,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T21:21:13.906970+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train the same detectors on nuScenes plus synthetic data generated with identical sensors and annotation rules but digital-twin maps of a different country; if the AI-Hub mAP gain is unchanged, the target-specific Korean maps are not what carries the result. A complementary check is to evaluate the same models on a second unseen target city outside Korea, where MORDA should confer no advantage if the effect is genuinely target-specific.","supporting_citations":[{"cited_title":"nuScenes: A multimodal dataset for autonomous driving,","cited_arxiv_id":null,"evidence_quote":"Defines the real-source domain, its sensor suite, ten detection classes, labeling rules, and the mAP and NDS metrics used throughout the paper."},{"cited_title":"MORAI Simulator,","cited_arxiv_id":null,"evidence_quote":"Provides the simulator environment, digital-twin maps, traffic generation, and ground-truth labeling used to create MORDA."},{"cited_title":"PointPillars: Fast Encoders for Object Detection from Point Clouds,","cited_arxiv_id":null,"evidence_quote":"Supplies one of the LiDAR 3D detectors used to evaluate MORDA's effect on nuScenes and AI-Hub."},{"cited_title":"SSN: Shape Signature Networks for Multi-class Object Detection from Point Clouds,","cited_arxiv_id":null,"evidence_quote":"Supplies a LiDAR 3D detector used in the main experiments, where MORDA yields a 5.7 mAP gain on AI-Hub."},{"cited_title":"Center-based 3d object detection and tracking,","cited_arxiv_id":null,"evidence_quote":"Supplies the 3D detector family whose baseline collapses under domain shift and where MORDA produces the largest gains."},{"cited_title":"Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks,","cited_arxiv_id":null,"evidence_quote":"Supplies the 2D detector used in the main comparison and in the synthetic-dataset benchmark against VKITTI2, SYNTHIA-AL, and SHIFT."},{"cited_title":"Virtual KITTI 2,","cited_arxiv_id":null,"evidence_quote":"Provides the closest prior synthetic dataset that clones a real dataset, used as a comparison point for MORDA."},{"cited_title":"SHIFT: A Synthetic Driving Dataset for Continuous Multi-Task Domain Adaptation,","cited_arxiv_id":null,"evidence_quote":"Provides a compared synthetic dataset that reaches similar 2D mAP but with roughly eight times more frames than MORDA."},{"cited_title":"Temporal coherence for active learning in videos,","cited_arxiv_id":null,"evidence_quote":"Provides another compared synthetic dataset in the 2D detection benchmark."}],"review_version":1}