{"id":"2f6014c5-8e11-4898-9f94-42effbd0f5d7","arxiv_id":"2505.05474","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":0.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"The paper surveys 3D scene generation and organizes methods into four paradigms, with datasets, evaluation metrics, applications, and future directions.","lead":"A team from Nanyang Technological University reviews the fast-growing field of 3D scene generation and sorts the work into four paradigms: procedural rules, neural 3D construction, image-based synthesis, and video-based synthesis. The survey is a practical map for researchers and engineers working on generative AI for games, robotics, and autonomous driving.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Taxonomy's procedural/neural boundary is internally inconsistent: §2.3 defines procedural as 'without learned priors', yet §3.1.3 classifies LLM-based generation, which relies on learned large language models, as procedural; this undermines the central four-paradigm claim.","rationale":"The reader's verdict is CONDITIONAL, with the weakest assumption being that the taxonomy assumes clean assignment of methods to four paradigm buckets despite boundary-spanning hybrids such as LLM-controlled procedural generation. I agree that the boundary is the most load-bearing point, but I locate the problem more precisely: the paper is internally inconsistent about what 'procedural' means. Section 2.3 explicitly excludes learned priors from procedural generation, while Section 3.1.3 explicitly includes methods whose essential component is a learned LLM. This is not merely a matter of soft boundaries or overlapping convenient categories; it is a contradiction between the taxonomy's foundational definition and its application. The central claim—that the four paradigms are a valid and complete organizing framework—requires at least that the categories be assignable from the paper's definitions. They are not, unless the definition in Section 2.3 is amended or the LLM-based subcategory is relocated. The concrete test I propose is a simple audit of Table 2's procedural rows against the Section 2.3 definition; it will settle the question without subjective judgment about quality or completeness. I agree with the reader that Table 1 is uninformative because every cell has the same three-star rating, and that the Figure 1 search methodology is undocumented; however, those are presentational and evidential flaws, not threats to the central claim. The taxonomy self-consistency issue is the more fundamental concern. Because the paper can plausibly be fixed by clarifying the definitions or reclassifying LLM-based generation, a REJECT verdict would be too strong. CONDITIONAL remains the appropriate verdict: the survey should be accepted only if the taxonomy's category definitions are made internally consistent and the boundary assignments in Table 2 are justified against those definitions.","tokens_in":55875,"tokens_out":5694,"duration_ms":66710,"concrete_test":"Perform a definitional audit of the Procedural Generation rows in Table 2, especially the three LLM-based methods (LayoutGPT, 3D-GPT, SceneX). For each row, read the corresponding description in Section 3.1 and mark whether a trained LLM or VLM is an essential component of the generation process. If any row marked 'Procedural' has an essential learned model, the taxonomy violates Section 2.3's own 'without learned priors' definition, and the procedural/neural boundary is not sharp. The test is decisive: either no procedural row uses a learned model as an essential component, in which case the concern fails, or at least one does, in which case the definitions of Sections 2.3 and 3.1.3 must be reconciled before the four-paradigm taxonomy can be accepted as stated.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is that state-of-the-art 3D scene generation can be organized into four paradigms, with Section 3 stating that methods are classified 'based on their generation paradigms.' For that taxonomy to be load-bearing, the four categories must be well-defined and mutually exclusive. Section 2.3 defines procedural generators as 'construct[ing] structured 3D scenes through deterministic or stochastic logic without learned priors,' contrasting them with models that 'learn statistical patterns' such as AR models, VAEs, GANs, and diffusion models. Section 3.1.3 then places LLM-based generation under Procedural Generation, describing methods that use Large Language Models to generate scene layouts, scene graphs, object parameters, or Python code that drives procedural software. An LLM is a learned statistical model on any reasonable reading; it encodes priors learned from data. Under the Section 2.3 definition, these methods should not be procedural. The survey tries to paper over this by amending Section 3's opening to say procedural generation uses 'predefined rules, enforced constraints, or prior knowledge from LLMs,' but this contradicts its own foundational definition. The result is that the boundary between Procedural Generation and Neural 3D-based Generation is not sharp: both categories now permit learned components, and the distinction becomes one of output format or control mechanism rather than generation paradigm. This directly affects the central claim's completeness criterion: if a reader cannot reliably assign methods to categories from the paper's own definitions, the four-paradigm organization is not a valid organizing framework as claimed. The reader's weakest assumption identified hybrid boundary cases; this internal definitional contradiction is a sharper instance of that concern.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript surveys 3D scene generation for computer vision and graphics, proposing a hierarchical taxonomy of four generation paradigms: procedural generation, neural 3D-based generation, image-based generation, and video-based generation. It reviews core 3D scene representations (voxel grids, point clouds, meshes, neural fields, 3D Gaussians, image sequences) and generative models, then organizes representative methods under the four paradigms with subcategories. It also surveys datasets, evaluation metrics and benchmarks, downstream applications (scene editing, human-scene interaction, embodied AI, robotics, autonomous driving), and challenges and future directions. The central claim is that the four-way taxonomy is a valid and complete organizing framework for the current literature.","tokens_in":56230,"tokens_out":6394,"duration_ms":67134,"significance":"If the taxonomy is made internally consistent, this survey will be a useful and current reference: it covers a large body of work through 2025, provides two large summary tables (methods in Table 2 and datasets in Table 3), and gives a helpful overview of evaluation protocols and applications. Its strengths include breadth, the dataset-to-paradigm usage mapping, and the explicit discussion of challenges. The main weaknesses are the definitional inconsistency of the central taxonomy and the lack of comparative support in Table 1, both described below.","major_comments":[{"comment":"The definition of procedural generation in §2.3 is internally inconsistent with the classification of LLM-based methods in §3.1.3. Section 2.3 characterizes procedural generators as constructing scenes 'without learned priors' and contrasts them with models that 'learn statistical patterns,' yet §3.1.3 places LLM-based generation, which relies on large language models trained on data, under Procedural Generation, and Section 3's opening redefines procedural generation to include 'prior knowledge from LLMs.' Because an LLM is a learned statistical model on any reasonable reading, the procedural/neural boundary is not sharp: both categories now permit learned components. Since the paper's central claim is that the four paradigms organize the field 'based on their generation paradigms,' this contradiction needs to be resolved by either changing the §2.3 definition (e.g., distinguishing the control mechanism from the model class) or reassigning LLM-based methods to another paradigm.","section":"§2.3 and §3.1.3"},{"comment":"Table 1 assigns exactly three stars (⋆⋆⋆) to every category on every characteristic, so it does not support the comparative claims it is invoked for. The text repeatedly references this table (e.g., §3.1 states procedural methods 'offer high efficiency and spatial consistency'; §3.2 states neural methods 'achieve high view and semantic consistency, but their controllability and efficiency remain limited'), but an all-equal rating matrix conveys no information about trade-offs. I recommend either replacing the star ratings with differentiated, justified values (with citations or a documented rubric) or removing the table and stating the trade-offs in prose.","section":"Table 1"},{"comment":"The survey's literature statistics and its claim to be comprehensive are not backed by a documented search or selection protocol. Figure 1 reports annual paper counts by paradigm, but the text does not state which databases were queried, which query terms were used, how duplicates or multi-paradigm papers were assigned, or what inclusion/exclusion criteria were applied. As a result, the counts cannot be reproduced and the apparent trend is not verifiable. I recommend adding a methodology subsection or appendix table describing the search protocol, the screening process, and the coding of papers into the four paradigms.","section":"Figure 1 and Section 1"}],"minor_comments":[{"comment":"The condition-code list in Table 2 uses 'C' for both 'constraint' and 'camera pose'; for example, Wu et al. [27] uses C as a constraint while MagicDrive [39] uses C as a camera pose, making the table ambiguous. Please use distinct codes (e.g., CO for constraint and CP for camera pose).","section":"Table 2 caption"},{"comment":"The 'Used by' codes P, N, I, V overlap with the scene-type codes I, N, U in the same table; for instance, 'I' in the 'Used by' column denotes image-based generation while 'I' in the 'Type' column denotes indoor. Consider using different letter sets or adding a clear legend to prevent confusion.","section":"Table 3"},{"comment":"The procedural-generator update equation S_{t+1}=R(S_t,Θ) is not numbered, unlike Eq. (1) and Eq. (2); numbering it would make cross-referencing easier.","section":"Section 2.3"},{"comment":"The reference list has inconsistent formatting, e.g., [335] is cited only with a URL and [314] does not use the 'in' convention used by other entries; please align all references to the journal style.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The authors' own scene-generation papers are cited extensively in the core sections; this is normal for a survey by active researchers and does not affect my assessment. The main concern is the internal consistency of the taxonomy, which should be resolvable with a revised definition or a reassignment of LLM-based methods."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this survey is worth taking seriously. It gives the field a needed map—procedural vs. neural-3D vs. image-based vs. video-based—with subcategories that mostly fit the literature, a broad dataset table, and a sensible applications section. For a fast-moving area with scattered prior surveys, that organizational work is a real contribution, and I expect the paper to get cited as an entry point.\n\nWhat it does well: coverage. It reaches well into 2025, includes Infinigen, CityDreamer, Director3D, CAT3D, WonderJourney, recent 4D methods, and a genuinely useful set of datasets and metrics. The four paradigms are not arbitrary; most methods do fall mainly into one bucket. The self-citations to the authors' SceneDreamer/CityDreamer/GaussianCity line are fine in a survey where those papers are on-topic.\n\nSoft spots, in rough order of size. First, the procedural/LLM boundary is inconsistent. Section 2.3 defines procedural generators as operating \"without learned priors,\" but Section 3.1.3 classifies LLM-based layout/code generation as procedural, and Section 3's opening even says procedural generation can use \"prior knowledge from LLMs.\" An LLM is a learned statistical model; under the paper's own definition those methods are not procedural. This is a real internal contradiction, and it makes the taxonomy's completeness claim looser than stated. It is fixable by redefining the category as control-mechanism-based generation, or by splitting LLM methods into a hybrid subcategory. Second, Table 1 rates every paradigm with the same three stars on every axis. That conveys no comparison and should either be replaced with actual ratings plus a protocol, or deleted. Third, Figure 1's paper counts have no documented search protocol, so the growth trend is not independently reproducible. Minor: I noticed a few dataset-table nits (e.g., KITTI's scene count looks off), but nothing that changes the survey's utility.\n\nNone of this sinks the survey. The central argument—that the current literature sorts into four paradigms with identifiable sub-paradigms—holds up once the procedural definition is cleaned up. The missing comparison data and statistics methodology are presentation problems, not deep flaws.\n\nI'd send it to peer review. The field needs this reference, and a serious referee can push the authors to fix the definitions and remove or substantiate Table 1. I'd also bring it to a reading group: it's a fast way to get everyone up to speed, and the taxonomy debate is genuinely useful for framing new work.","headline":"A useful, broad survey of 3D scene generation whose four-way taxonomy is basically sound but needs a cleaner procedural/LLM boundary, a real comparison table, and search statistics before it should be cited as authoritative.","tokens_in":56717,"tokens_out":2332,"would_cite":true,"duration_ms":27878,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This survey claims that the entire field of 3D scene generation can be organized into four generation paradigms — procedural, neural 3D-based, image-based, and video-based — and that this taxonomy is complete enough to serve as a map of…","keywords":["3D scene generation","procedural generation","neural 3D-based generation","image-based generation","video-based generation","diffusion models","3D Gaussians","NeRF"],"falsifier":"Take a recent 3D scene generation paper, such as an LLM-driven procedural city builder or an iterative image-outpainting pipeline, and ask whether it can be assigned to exactly one of the four categories without arbitrary choice; if a substantial share of the literature resists unique assignment, the taxonomy's completeness claim and the comparative conclusions drawn from it give way.","tokens_in":55649,"feed_emoji":"🗺️","tokens_out":4170,"duration_ms":42399,"temperature":0.7,"pith_summary":"This survey claims that the entire body of 3D scene generation research can be organized by generation paradigm into four categories: procedural rules, neural 3D-aware generative models, 2D image generators, and video generators. The authors argue that prior surveys missed this structure because they focused on narrow subdomains or treated scenes as a side topic. A sympathetic reader would care because a valid taxonomy turns a fragmented literature into a map with clear trade-offs: each paradigm buys a different combination of realism, diversity, view consistency, controllability, and physical plausibility. The survey's comparative framework, summarized in Table 1, is intended to help researchers choose an approach and to expose where the open problems lie.","feed_headline":"Survey: all 3D scene generation falls into four paradigms","feed_subtitle":"Procedural, neural 3D, image, and video pipelines each buy different trade-offs; the map shows where gaps are.","key_machinery":"The load-bearing device is the hierarchical taxonomy built on the 'generation paradigm' — the kind of generator that turns input into a 3D scene. The four root categories are procedural, neural 3D-based, image-based, and video-based generation, each with named subcategories (e.g., rule/optimization/LLM-based; scene parameters/scene graph/semantic layout/implicit layout; holistic/iterative; two-stage/one-stage). The taxonomy carries the argument because every comparative claim in the survey, including the trade-off table and the discussion of challenges, is organized around these four buckets.","core_discovery":"The central claim is that state-of-the-art 3D scene generation methods fall into four paradigms defined by how the scene is produced: procedural generation (rules, optimization, or LLM-guided code), neural 3D-based generation (scene parameters, scene graphs, semantic layouts, or implicit layouts fed into 3D-aware generators), image-based generation (holistic panorama synthesis or iterative extrapolation), and video-based generation (two-stage or one-stage video diffusion that animates or constructs scenes over time). Each paradigm is analyzed for its technical foundation, its characteristic 3D representation, and its comparative strengths and weaknesses. The survey further claims that this taxonomy is complete for the current literature and that the field's progress can be traced as a shift from procedural control toward learned image and video priors.","pith_inferences":["The taxonomy's boundaries are already softening: LLM-based procedural generation and image-outpainting-with-3D-reconstruction are hybrids that the survey must place in one bucket by fiat, so the four paradigms may be better seen as poles of a continuum than as disjoint classes.","A testable consequence of the survey's completeness claim is that any new 3D scene generation paper should be assignable to exactly one root category; a meta-analysis of recent papers could check whether assignment is unambiguous in practice.","The survey's framing suggests the next frontier is not a fifth paradigm but the convergence of video-based realism with neural 3D representations — the paper's own future-directions section points at unified perception-generation models, which would erase the current paradigm boundaries."],"forward_implications":["If the taxonomy is correct, a researcher can locate any existing method in one of four paradigm buckets and immediately read off its expected trade-offs in realism, view consistency, controllability, and physical plausibility.","The survey's framing implies that image- and video-based paradigms currently lead in photorealism and diversity, while procedural and neural 3D-based paradigms lead in geometric and semantic consistency, so combining paradigms is a natural route to scene generation that is both realistic and 3D-coherent.","The four-way split gives the field a shared vocabulary, making it possible to compare methods across paradigms on common datasets and evaluation metrics rather than in isolated subcommunities.","The taxonomy identifies missing combinations — such as physics-aware generation and interactive generation — as open directions, which the survey lists as future work."],"supporting_citations":[{"why":"Defines the procedural modelling paradigm and its rule-based methods, anchoring the survey's first category.","marker":"[44]"},{"why":"Introduces NeRF, the neural representation that anchors much of neural 3D-based generation.","marker":"[31]"},{"why":"Introduces 3D Gaussian splatting, the other core representation for neural 3D-based generation and for image- and video-based reconstruction.","marker":"[32]"},{"why":"Anchors the image-based iterative generation paradigm with perpetual view generation from a single image.","marker":"[33]"},{"why":"Anchors the video-based generation paradigm by providing the video diffusion backbone these methods build on.","marker":"[37]"},{"why":"Shows the scalability of photorealistic procedural generation, supporting the survey's claims for the procedural paradigm.","marker":"[80]"},{"why":"A representative scene-parameter method for neural 3D-based generation, used to illustrate that subcategory.","marker":"[86]"},{"why":"Provides the latent diffusion engine used across image- and video-based generation, linking those two paradigms.","marker":"[189]"},{"why":"A prior survey on 3D representations that the authors position against, supporting their claim of a gap for scene generation.","marker":"[51]"},{"why":"A prior text-to-3D survey that they say treats scene generation peripherally, justifying the need for a scene-specific taxonomy.","marker":"[52]"}],"fun_headline_variants":["Four paradigms define all 3D scene generation","3D scenes from rules, neural nets, images, or video","Survey: four paradigms cover all 3D scene generation","Procedural, neural, image, video: the 3D scene generation map"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The taxonomy assumes every published method can be cleanly sorted into exactly one of the four paradigm buckets, even though hybrid methods — LLM-controlled procedural generation, or image outpainting followed by 3D reconstruction — span the boundaries.","fun_headline_variants_meta":{"raw":{"variants":["Four paradigms define all 3D scene generation","3D scenes from rules, neural nets, images, or video","Survey: four paradigms cover all 3D scene generation","Procedural, neural, image, video: the 3D scene generation map"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000968,"raw_usage":{"total_tokens":4130,"prompt_tokens":966,"completion_tokens":3164,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":582,"completion_tokens_details":{"reasoning_tokens":3091}},"tokens_in":582,"tokens_out":3164,"duration_ms":20800,"temperature":1.0,"reasoning_tokens":3091,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T23:01:59.680830+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a recent 3D scene generation paper, such as an LLM-driven procedural city builder or an iterative image-outpainting pipeline, and ask whether it can be assigned to exactly one of the four categories without arbitrary choice; if a substantial share of the literature resists unique assignment, the taxonomy's completeness claim and the comparative conclusions drawn from it give way.","supporting_citations":[],"review_version":1}