{"id":"5e9fd7d4-8267-42a1-8b81-087029320038","arxiv_id":"2412.17628","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A review that classifies radiance field editing into explicit, latent space, text-guided, compositional, and other categories, with application and dataset tables.","lead":"This paper surveys methods for editing neural radiance fields and 3D Gaussian splatting scenes, and organizes them into a new taxonomy by editing strategy. It is a practical map of a fast-growing subfield for researchers selecting editing methods and for newcomers learning the landscape.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The comprehensiveness claim is unverifiable: no inclusion criteria are given, and the taxonomy's three stated editing types are not used as the organizing axis, so the survey cannot currently support 'comprehensive' or 'new taxonomy'.","rationale":"The reader's weakest assumption was that coverage may be unrepresentative because no search strategy is documented. That is a legitimate process-level concern, and I agree with it partly. However, the more load-bearing and more specific problem is internal: the paper's stated taxonomy contribution is not operationalized. Section 4 announces three editing types but organizes everything by five method families, and Table 1's 'Editing Type' column uses method-family labels rather than the three types. This means the central 'new taxonomy' claim cannot be checked, applied, or compared against existing surveys even in principle. The missing selection protocol compounds the problem by making 'comprehensive' unverifiable, but the taxonomy defect alone justifies conditional acceptance: the survey is still useful as a curated review, but the claimed organizing contribution requires either a revised taxonomy that consistently assigns methods to a single category, or an explicit statement that the three editing types are orthogonal dimensions to be combined with method families. I therefore keep the reader's CONDITIONAL verdict, with the condition expanded to include a reproducible classification mapping and a documented coverage protocol.","tokens_in":21420,"tokens_out":4691,"duration_ms":44712,"concrete_test":"Construct, from the paper's own text, a complete assignment matrix for all cited editing methods: for each cited work, record (a) which of the three declared editing types (geometry/appearance/dynamic) it supports according to the text and (b) which Section 4 family the text assigns it to. Then check whether every method is assignable, whether any family mixes methods of different declared types without an explicit rule, and whether the dynamic-editing column is non-empty outside the 'Other' catch-all. If any method is unassignable or any family requires an overlapping assignment, the taxonomy is not a classifying scheme in the claimed sense.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claims are the abstract's 'comprehensive survey' and 'new taxonomy' of radiance-field editing methods. For the first to hold, the included method set must be representative or complete; for the second, the classification must be principled and consistently applied. Neither is presently checkable from the manuscript. The text gives no search venues, query terms, inclusion/exclusion criteria, or screening process in Sections 2-4, so 'comprehensive' cannot be distinguished from curated coverage. More importantly, the taxonomy is internally underspecified: Section 4 states 'we identify three different types, namely geometry editing, appearance editing and dynamic editing,' but the actual section headings (4.1 Explicit representation, 4.2 Latent space based NeRF, 4.3 Text guided editing, 4.4 Compositional, 4.5 Other) are method families, not editing types. Table 1's 'Editing Type' column likewise lists method families ('Style transfer', 'CLIP', 'Mesh', 'Text', 'Hyperspace'). There is no stated rule for mapping methods to families or for how the three declared editing types interact with those families. CLIP-NeRF appears under both latent-space editing (Section 4.2.1) and text-guided editing (Section 4.3.2) without cross-cutting analysis, and dynamic editing, one of the three declared types, appears only as a 'Video Editing' column in Table 1 and implicitly inside Compositional and Hyper-space subsections, never as a first-class category in the hierarchy. A survey whose main contribution is a taxonomy should provide a definitional test that assigns any given method to a well-defined cell; as written, the three-type claim and the five-family structure are two unconnected axes, so the 'new taxonomy' cannot be validated or used by readers.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript is a survey of radiance-field editing, focused on NeRF and 3D Gaussian Splatting. It reviews the mathematical background of volumetric rendering and the two base representations, then proposes a taxonomy of editing approaches and surveys representative models, applications, datasets, and metrics. The abstract states that the survey is comprehensive and that the taxonomy is new; the Section 4 text declares geometry, appearance, and dynamic editing as the three editing types. The paper closes with challenges and opportunities.","tokens_in":21734,"tokens_out":6945,"duration_ms":61793,"significance":"The survey addresses a timely and under-covered topic, and it collects a broad set of recent references, including 3DGS-based editing works, which is useful for researchers entering this area. The discussion of missing benchmarks, user interfaces, and editing time in Section 7 is fair and well aimed. If the taxonomy were made internally consistent and the selection of literature transparent, this would be a valuable reference. At present, however, the central claim of a 'new taxonomy' is weakened by the mismatch between the declared editing types and the actual section organization, and the 'comprehensive' claim cannot be verified from the manuscript.","major_comments":[{"comment":"The central claim of a new taxonomy is not currently supported by the manuscript's organization. Section 4 states that the taxonomy distinguishes geometry editing, appearance editing, and dynamic editing, but the subsections 4.1–4.5 are organized by method families (explicit representation, latent space, text-guided, compositional, other), and Table 1's 'Editing Type' column lists the same families (Style transfer, CLIP, Mesh, Text, Hyperspace, Composition). There is no rule stated for mapping a method to a family or to one of the three declared editing types, and dynamic editing only appears as a 'Video Editing' flag in Table 1 rather than as a first-class branch of the taxonomy. Please either reorganize Section 4 around the three declared axes or explicitly redefine the taxonomy as a two-dimensional scheme with method families on one axis and editing types on the other.","section":"Section 4, Table 1"},{"comment":"The 'comprehensive survey' claim is not verifiable. The paper gives no search venues, query terms, inclusion/exclusion criteria, or screening process in Sections 2 and 4, so the reader cannot distinguish a comprehensive collection from a curated selection. This matters because the abstract and Section 2 use 'comprehensive' as a differentiator against prior surveys. Please add a short methodology paragraph describing how the literature was collected and filtered, or replace 'comprehensive' with a more modest claim about representative coverage.","section":"Abstract; Sections 2 and 4"},{"comment":"The subsection on 3D Gaussian Splatting does not actually review 3DGS editing methods: it summarizes the representation's advantages, mentions GaussianEditor and GaussianGrouping in two sentences, and states that many methods 'will be detailed in the corresponding parts of this paper' without a concrete pointer. Since the title promises coverage of explicit radiance-field representations, this subsection should either provide a dedicated review of 3DGS editing methods (including geometry, appearance, and dynamic editing) or explicitly map each 3DGS method to the relevant later subsections.","section":"Section 4.1.3"},{"comment":"The abstract promises a comparison of approaches 'in terms of editing options and performance,' but no performance comparison is delivered. Table 1 compares editing options and control, while Section 6 only lists evaluation metrics and notes the absence of standard benchmarks. Either add a comparison table with quantitative results (rendering quality, editing time, user-study scores where available) for at least the most popular models, or revise the abstract to say that the comparison covers editing options only.","section":"Abstract; Section 6"}],"minor_comments":[{"comment":"'Peak Signal to Noise Ration' should be 'Peak Signal to Noise Ratio', and 'SSIM))' has an extra closing parenthesis.","section":"Section 6.2"},{"comment":"In the sentence 'GIRAFFE [31] directly improves GRAF by two novelties First it adds ...', there is a missing comma or period after 'novelties'; the sentence is run-on.","section":"Section 4.2.1"},{"comment":"The Panoptic Dataset row is hard to parse: the camera count '480 VGA camera, 30+ HD camera' and the FPS values '25 30' should be split into explicit columns, and the resolution column should be labeled as VGA/HD.","section":"Table 2"},{"comment":"Several references appear to lack complete bibliographic details, for example [88] 'arXiv e-prints, 2403 (2024)' does not give an article identifier or a complete year; please normalize all arXiv entries.","section":"References"},{"comment":"The notation ln nf m in Eq. (10) uses subscript and spaces inconsistently; use a single symbol such as L_NNFM.","section":"Section 6.2, Eq. (10)"}],"recommendation":"major_revision","confidential_remarks":"The stress-test concern about the taxonomy is well placed, and the comprehensiveness concern is also legitimate; both are fixable within the scope of a survey. I do not see evidence of circular reasoning or citation abuse: reference [98] is used incidentally and is not load-bearing. The manuscript is not ready for acceptance but is salvageable with a major revision that makes the taxonomy internally consistent and documents the literature-selection process."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know two things about this paper. First, it is a genuinely useful reference: it covers radiance field editing from mesh proxies through latent codes, text-guided diffusion, compositional scenes, and hyperspace methods, and it does a good job of connecting NeRF and 3DGS. Second, the taxonomy promised in the abstract – geometry, appearance, dynamic editing – is not the axis the paper actually uses. The section structure is organized by method families, Table 1's 'Editing Type' column lists families (Style transfer, CLIP, Mesh, Text, Hyperspace), and dynamic editing appears only as a video-editing checkbox plus some content inside other subsections. That is a real gap between claim and structure, and the stress-test note is correct about it.\n\nWhat the paper does well: the coverage is current and broad, with accurate descriptions of signposts like ARF's NNFM loss, DreamFusion's SDS, Instruct-NeRF2NeRF, and the NeuS-based mesh proxy line. The evaluation section honestly notes there is no consensus on editing-specific metrics and that most quality assessment is subjective. The applications and future directions are sensible and not inflated. If I needed a quick map of this subfield, I would trust this paper.\n\nSoft spots: the taxonomy's classification rule is never stated. There is no definitional test for assigning a method to a family, and the relationship between the three announced editing types and the five families is left implicit. CLIP-NeRF appearing in both 4.2.1 and 4.3.2 without cross-referencing is a symptom. There is also no literature selection protocol – no search venues, no inclusion criteria – so 'comprehensive' cannot be verified. These are fixable. The paper needs a survey methodology paragraph and either a reorganization around the three types or an explicit mapping from families to types. Minor issues: 'Peak Signal to Noise Ration' in 6.2, and some rows in Table 2 are incomplete.\n\nBottom line: this is a helpful survey for newcomers and for researchers wanting a structured overview. It is not a formal classification system as claimed. I would send it to a serious referee, not desk-reject, and would ask for revisions that tighten the taxonomy and add the selection methodology. The one self-citation [98] is used purely as an example and is not load-bearing.","headline":"A useful current survey of radiance field editing, but its taxonomy is a set of method families rather than the three editing types promised in the abstract; worth reviewing with revisions.","tokens_in":22261,"tokens_out":3885,"would_cite":true,"duration_ms":33270,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This survey claims that the fragmented literature on editing radiance fields can be organized into a unified, methodology-based taxonomy spanning geometry, appearance, and dynamic editing, and it maps NeRF and 3D Gaussian Splatting…","keywords":["Survey","Neural Radiance Fields","3D Gaussian Splatting","Radiance Field Editing","Latent Space Manipulation","Text-guided editing","Style transfer","Scene editing taxonomy"],"falsifier":"A concrete check would be to assemble a dated list of radiance-field editing papers published before the survey and try to place each into the proposed taxonomy; if a significant number fit no branch, or if two reviewers disagree on placements, the claim that the taxonomy comprehensively covers the literature fails.","tokens_in":21249,"feed_emoji":"🎨","tokens_out":7180,"duration_ms":63312,"temperature":0.7,"pith_summary":"Radiance fields store a 3D scene in a form that is hard to modify: a neural network's weights in NeRF or a set of Gaussian primitives in 3D Gaussian Splatting. This paper claims that the many scattered strategies for editing such scenes can be organized into a single tree-shaped taxonomy with three top-level editing types (geometry, appearance, and dynamic) and several methodological families, including explicit representation, latent-space manipulation, text-guided editing, compositional editing, hyper-space deformation, surface NeRF, and knowledge distillation. It also reviews pioneering models in each family, compares state-of-the-art methods, and catalogs applications from autonomous-driving simulation to face and body editing. If the taxonomy holds, it gives researchers and practitioners a shared map of what editing operations are possible, how they are achieved, and where the open problems are.","feed_headline":"Survey maps radiance-field editing into a three-branch tree","feed_subtitle":"The taxonomy spans geometry, appearance, and dynamic changes, across NeRF and 3D Gaussian Splatting methods.","key_machinery":"The main instrument of the survey is its taxonomy: a tree that classifies editing methods by the operation they perform and the mechanism they use, independent of whether the scene is a NeRF or a 3D Gaussian Splatting model. It formalizes editing as replacing an original radiance field $L(r,d)$ by a modified field $L'(r,d)$ over all points and directions, then distinguishes implicit editing, which retrains or modifies the network weights $F_{\\theta}$, from explicit editing, which directly manipulates Gaussian parameters $G_i = \\langle \\mu_i, S_i, R_i, \\alpha_i, c_i \\rangle$. The taxonomy does the work of turning a large set of individual papers into comparable categories, so that methods as different as bending rays through a mesh proxy and optimizing a CLIP loss can be discussed as sibling strategies.","core_discovery":"The paper's central claim is that radiance-field editing, though scattered across NeRF and 3DGS papers, is a coherent research area that can be classified by editing methodology rather than by underlying representation. It proposes a tree taxonomy in which every method falls under geometry, appearance, or dynamic editing, and within those, into families: explicit representations (mesh proxies, editable spatial encodings, and Gaussian primitives), latent-space techniques (conditional generative fields and style transfer), text-guided editing (score distillation, CLIP guidance, and iterative dataset update), compositional scene decomposition, hyper-space deformation, surface-based NeRF editing, and knowledge distillation. The paper also argues that many strategies introduced for implicit NeRF editing carry over to explicit 3DGS editing, and that text-to-image generative models have made editing more accessible. It further claims that the field currently lacks agreed-upon editing-specific benchmarks, with most evaluation relying on rendering metrics such as PSNR, SSIM, and LPIPS plus subjective comparisons.","pith_inferences":["The survey's representation-agnostic framing implies the taxonomy should also accommodate future representations; any new scene encoding that can be edited can likely be slotted into an existing family rather than requiring a new one.","An implication the survey leaves implicit is that editing speed, not edit quality, is the current bottleneck: most families rely on lengthy optimization, so the methods that cut optimization time could dominate over time.","The iterative-dataset-update family suggests a general recipe -- use a 2D editor to modify rendered views, then retrain -- that could extend beyond radiance fields to any differentiable renderer.","A testable extension would be to score each family on edit locality and view consistency, since the survey notes that current metrics do not yet capture those 3D-specific requirements."],"forward_implications":["Because the taxonomy is representation-agnostic, a method validated on NeRF can be assessed for transfer to 3DGS within the same category, and vice versa.","The dynamic editing branch links editing to dynamic radiance fields, so object motion and topology changes can share machinery with video editing.","Text-guided editing is the most accessible route for non-expert users, with SDS, CLIP, and instruction-based dataset updates as the main levers.","The lack of editing-specific metrics means that reported progress mostly rests on rendering quality plus subjective judgment; a shared benchmark would make comparisons sharper.","The classified methods directly support applications such as synthetic scenario generation for autonomous driving, human face and body editing, and stylized asset creation."],"supporting_citations":[{"why":"Supplies NeRF, the implicit radiance field representation that most editing methods in the survey modify.","marker":"[3]"},{"why":"Supplies 3D Gaussian Splatting, the explicit representation whose editability motivates the survey's dual coverage.","marker":"[10]"},{"why":"Introduces conditional generative radiance fields with disentangled shape and appearance codes, anchoring the latent-space family.","marker":"[29]"},{"why":"Introduces the Nearest Neighbor Feature Matching loss that the style-transfer family and later text-guided models build on.","marker":"[64]"},{"why":"Introduces Score Distillation Sampling, the diffusion-model guidance used by the text-guided editing family.","marker":"[76]"},{"why":"Provides the CLIP image-text latent space used by CLIP-based editing methods to steer edits from prompts.","marker":"[77]"},{"why":"Introduces Instruct-NeRF2NeRF and the iterative dataset update strategy for instruction-based 3D editing.","marker":"[78]"},{"why":"Introduces hyper-space coordinates for topologically varying scenes, the basis of the hyper-space editing family.","marker":"[28]"},{"why":"Introduces NeuMesh's mesh plus local implicit field, used for geometry and texture editing and by later editors such as DreamEditor.","marker":"[44]"},{"why":"Demonstrates hierarchical splatting for editing 3D Gaussians, a reference point for the 3DGS editing branch.","marker":"[53]"}],"fun_headline_variants":["Editing radiance fields: a tree taxonomy","Radiance-field editing: a three-branch taxonomy","Survey sorts radiance-field editing by methodology","Geometry, appearance, dynamics: the editing taxonomy","Editing NeRF and 3DGS: a survey's taxonomy"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The survey's claim of being comprehensive rests on the untested premise that the papers it discusses are a complete and representative sample of the editing literature, since it does not state a search strategy or inclusion criteria.","fun_headline_variants_meta":{"raw":{"variants":["Editing radiance fields: a tree taxonomy","Radiance-field editing: a three-branch taxonomy","Survey sorts radiance-field editing by methodology","Geometry, appearance, dynamics: the editing taxonomy","Editing NeRF and 3DGS: a survey's taxonomy"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000389,"raw_usage":{"total_tokens":2025,"prompt_tokens":892,"completion_tokens":1133,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":508,"completion_tokens_details":{"reasoning_tokens":1059}},"tokens_in":508,"tokens_out":1133,"duration_ms":10344,"temperature":1.0,"reasoning_tokens":1059,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T05:19:17.397079+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A concrete check would be to assemble a dated list of radiance-field editing papers published before the survey and try to place each into the proposed taxonomy; if a significant number fit no branch, or if two reviewers disagree on placements, the claim that the taxonomy comprehensively covers the literature fails.","supporting_citations":[{"cited_title":"In: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp","cited_arxiv_id":null,"evidence_quote":"Introduces Instruct-NeRF2NeRF and the iterative dataset update strategy for instruction-based 3D editing."}],"review_version":1}