{"id":"b9723608-8e67-40eb-8ed2-12776245ee4b","arxiv_id":"2508.14879","paper_version":2,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":5.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":1,"one_line_summary":"A multimodal LLM trained on a large paired dataset turns point clouds into executable, semantically decomposed Blender Python scripts for shape reconstruction and editing.","lead":"MeshCoder maps 3D point clouds to editable Blender Python scripts, so a shape can be rebuilt and then modified by editing code. If it works, editing 3D models could become as simple as editing a program, which matters for reverse engineering, design, and LLM-based 3D reasoning.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Self-generated evaluation risk: the abstract builds the paired dataset from the same Blender API set that defines the target code; without an independent test split, the claimed superiority over baselines may reflect distribution memorization rather than shape-to-code reconstruction.","rationale":"The reader's weakest assumption included both the information-theoretic underdetermination of point clouds and the risk that a self-built dataset using the same APIs for evaluation makes metrics optimistic. I focus on the second concern because it is directly testable and is not contradicted by the abstract. The first concern is also real but much harder to settle experimentally; it would require showing that multiple valid code decompositions produce near-identical point clouds, which the unreadable full text does not allow us to assess. Because the supplied body is an undecodable corruption with a different arXiv watermark, I cannot confirm whether the paper already includes held-out categories, independent baselines, or leakage checks. That keeps the appropriate verdict at UNVERDICTED, matching the reader. I am not alleging deliberate misrepresentation; the concern is about evidence availability and evaluation design, and the proposed test is intended to resolve it rather than prejudge it.","tokens_in":13952,"tokens_out":6167,"duration_ms":72818,"concrete_test":"Release or inspect the exact train/test splits. Independently sample point clouds from a held-out corpus of Blender scripts authored by third parties (or from reverse-engineered real objects) that do not use MeshCoder's API set. Run the trained model, execute the generated scripts in headless Blender, and compare (i) fraction of successfully executing scripts and (ii) geometric error (e.g., Chamfer distance) against results on the original self-generated test split. If the self-generated split is substantially easier, the central superiority claim is distribution-specific.","verdict_should_be":"UNCHANGED","load_bearing_attack":"For the central claim to hold, MeshCoder must map point clouds to executable code and outperform existing shape-to-code methods on a meaningful benchmark. The abstract says the paired dataset is constructed using the same 'comprehensive set of expressive Blender Python APIs' that the model is trained to emit. If the evaluation objects are also generated by this same API/procedural pipeline, then train and test code share a narrow syntactic and semantic distribution; a model can score well by memorizing API idioms and parameter patterns without recovering geometry from the point cloud. The supplied full text is mojibake and carries a watermark for arXiv:2508.14880v3, not this paper, so the evaluation section cannot be inspected. The abstract contains no baseline names, metrics, or leakage controls, leaving this distribution overlap unchecked.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes MeshCoder, an LLM-based framework that reconstructs 3D objects from point clouds into executable Blender Python scripts. The authors introduce a set of Blender Python APIs, use these APIs to build a large-scale paired object-code dataset in which each object's code is decomposed into semantic parts, and train a multimodal LLM to translate point clouds into such scripts. The abstract claims superior performance in shape-to-code reconstruction and improved LLM reasoning about 3D shapes via the code-based representation. The submitted full text, however, is heavily corrupted, and the abstract contains no quantitative results, baselines, dataset statistics, or error analysis; consequently the central claims cannot currently be verified.","tokens_in":13952,"tokens_out":2705,"duration_ms":34190,"significance":"If the claims are substantiated, the framework would be a genuinely useful contribution to programmatic 3D reconstruction and editing: the idea of a large paired dataset of semantically decomposed Blender code, together with an expressive API set, is promising and well aligned with recent interest in code as an intermediate representation for shape generation. The paper also has a plausible downstream motivation, namely improving LLM reasoning about 3D shapes through code. However, as submitted, the contribution is conditional: no evidence is visible that the trained model outperforms existing shape-to-code methods, and the unreadable full text prevents inspection of the method, dataset, ablations, and metrics. The project homepage is a useful pointer, but it is not a substitute for a self-contained, verifiable manuscript.","major_comments":[{"comment":"The central claim of 'superior performance in shape-to-code reconstruction tasks' is unsupported by any number in the abstract: no metric, baseline, dataset size, or error measure is reported. Since the full text is unreadable, there is no way to check whether this claim is backed by experiments. Please provide, at minimum, quantitative comparisons on a defined benchmark with named baselines and standard reconstruction metrics (e.g., IoU, chamfer distance, code-execution accuracy).","section":"Abstract"},{"comment":"The paired dataset is built using the same 'comprehensive set of expressive Blender Python APIs' that the model is trained to emit. If the evaluation objects are also generated from the same procedural API pipeline, train and test code will share a narrow syntactic/semantic distribution, and the model could score well by memorizing API idioms rather than recovering geometry from the point cloud. The manuscript must describe the evaluation split, the source of test objects, and any leakage-control measures. Without this, the claimed superiority is not identifiable.","section":"Dataset construction (described in Abstract)"},{"comment":"The submitted full text is corrupted mojibake; equations, tables, and figures are largely unreadable, and the watermark reads 'arXiv:2508.14880v3 [cs.CL] 1 Sep 2025', which does not match the paper's arXiv ID (2508.14879). This prevents inspection of the proposed architecture, training procedure, dataset statistics, and experimental results. A correctly encoded and version-consistent manuscript is required before the technical content can be assessed.","section":"Full text"},{"comment":"The mapping from an unorganized point cloud to an editable, semantically decomposed program assumes that the input contains enough information to determine a useful programmatic decomposition. The paper does not address the ambiguity of this mapping (e.g., multiple programs can produce the same shape, and part semantics are not uniquely defined by geometry). The authors should discuss this information-theoretic limitation and provide evidence that the model resolves it consistently, e.g., through human evaluation or editing tests on ambiguous shapes.","section":"Abstract / Task definition"}],"minor_comments":[{"comment":"The rendering contains many garbled characters and misaligned headings throughout. A clean PDF is essential.","section":"Full text"},{"comment":"The project homepage URL is mentioned but not printed in a machine-readable form; include the full URL in the abstract or footnotes.","section":"Abstract"},{"comment":"The watermark/arXiv identifier is inconsistent with the paper number. Please correct the metadata to avoid version confusion.","section":"Front matter"},{"comment":"Section headings appear duplicated or out of order in the rendered text; verify the final layout and numbering.","section":"Full text"}],"recommendation":"uncertain","confidential_remarks":"The manuscript as received appears to be a corrupted/unreadable file, and the watermark references a different arXiv paper. This may be a production error rather than a substantive flaw, but it makes the current submission impossible to review. I recommend asking the authors to resubmit a correctly encoded, version-consistent PDF with full experimental details. The topic is within the journal's scope and the proposed idea is plausible, but no assessment of correctness or significance can be made from this file."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Honestly, this submission isn't reviewable as it stands. The full text is corrupted encoding (mojibake) and the watermark is for a different arXiv paper, so the only substantiative content I can assess is the abstract.\n\nWhat's worth taking seriously: the proposed pipeline — a purpose-built Blender Python API library, a large paired object-code dataset with semantic part decomposition, and a multimodal LLM translating point clouds to executable scripts — is a credible and potentially useful combination. If the system works as described, it would give users editable, semantically decomposed 3D programs from raw point clouds, which has clear value for reverse engineering and shape editing. The claim that code-based representations improve LLM 3D reasoning is also a reasonable hypothesis to test.\n\nThe soft spots are significant. The abstract contains no metrics, no baselines, and no dataset statistics, so 'superior performance' is unsubstantiated. The stress-test worry about train/test leakage is legitimate: if the evaluation objects come from the same procedural API pipeline that generated the training data, the model might excel by memorizing API idioms rather than recovering geometry from the point cloud. Nothing in the abstract rules this out. There's also the information-theoretic question of whether a point cloud alone determines the semantic parts and topology needed to regenerate the object; that deserves a direct discussion.\n\nBut the real blocker is the corrupted full text. I can't check the experiments, the ablations, or the dataset. As submitted, the paper provides no verifiable evidence beyond the abstract. That doesn't mean the work is wrong; it means I can't tell.\n\nWho's this for? Researchers in programmatic 3D reconstruction, LLM-based CAD, and shape understanding. They should watch for a clean version. As it stands, I'd send it back for a readable PDF and concrete evaluation details, not to peer review. The research idea has merit, but this submission is not reviewer-ready.","headline":"The idea is plausible and potentially useful, but the supplied full text is unreadable and the abstract has no numbers, so this version isn't reviewable.","tokens_in":14659,"tokens_out":4392,"would_cite":false,"duration_ms":45233,"reading_group":"maybe","serious_thinker":"unclear","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"MeshCoder maps point clouds to editable Blender Python scripts.","keywords":["point cloud reconstruction","Blender Python APIs","shape-to-code","programmatic 3D modeling","multimodal LLM","semantic part decomposition","3D shape understanding","editable geometry code"],"falsifier":"Take one mesh, generate two semantically different Blender programs that both reproduce it (for example, a hole made by boolean subtraction versus by a swept profile), sample point clouds from each, and see whether MeshCoder outputs the same code or the corresponding ground-truth code. If the output is not stable across valid programs, the part-level semantic code is not recoverable from geometry alone.","tokens_in":13665,"feed_emoji":"🧊","tokens_out":5684,"duration_ms":63292,"temperature":0.7,"pith_summary":"MeshCoder aims to establish that a raw 3D point cloud can be translated directly into an executable Blender Python script, rather than into a mesh or a niche CAD language. It builds an expressive set of Blender Python functions, constructs a large dataset pairing point clouds with code split into semantic parts, and trains a multimodal language model to generate that code from the input cloud. If this works, reverse engineering and shape editing become a matter of editing code, and LLMs can reason about shapes through code instead of raw points. The paper reports that this code-based approach outperforms prior shape-to-code methods and improves LLM performance on 3D shape understanding.","feed_headline":"Point clouds become editable Blender code","feed_subtitle":"A trained multimodal LLM writes executable, part-structured Python from raw 3D scans for editing and reasoning.","key_machinery":"The load-bearing mechanism is a code-as-geometry representation with three parts. First, an expressive Blender Python API set: a vocabulary of function calls that can synthesize complex solids, booleans, and modifiers. Second, a large paired object-code dataset in which each object's script is decomposed into semantic parts—blocks of code corresponding to distinct components such as legs, handles, or bodies. Third, a multimodal LLM trained to produce this part-structured code from a point-cloud input. The semantic-part separation is what converts a generated program from a flat shape description into an editable, interpretable model.","core_discovery":"MeshCoder's central claim is that shape reconstruction should be treated as structured code generation: a point cloud maps to an executable Blender Python program whose statements are grouped by semantic part. The paper argues that this representation lifts the two constraints of earlier work—limited domain-specific languages and small datasets—by giving the model a broad API vocabulary and a large paired corpus to learn from. On the paper's own terms, the result is a model that reconstructs complex objects with editable, geometrically and topologically meaningful code, and that same code format improves LLM reasoning about 3D shape.","pith_inferences":["This suggests code may be a better token-level interface for geometric reasoning than point clouds, since programs name operations and relationships explicitly; part-level queries such as 'remove the handle' would be a natural next task.","A testable extension is to measure how fidelity degrades as objects move outside the training API vocabulary, for example on open-source models not generated by the authors' pipeline; that would separate the method from the dataset.","If semantic-part decomposition is reliable, the model could support procedural modeling workflows—regenerating a whole shape after numerically editing a part—without retraining.","Because many geometries admit more than one valid program, a practical extension would treat code not as a unique ground truth but as one editable decomposition, letting users choose among valid alternatives."],"forward_implications":["Point-cloud-to-code reconstruction can move beyond small DSLs: richer API coverage plus large paired data makes complex, non-primitive objects expressible as programs.","Because outputs are executable Blender Python scripts, editing a shape becomes code modification—change a parameter, a boolean, or a part block to alter geometry and topology.","Using code as the representation gives LLMs a compact, structural handle on 3D shape, improving downstream 3D understanding compared with operating on raw points.","The pipeline can serve reverse engineering and shape-editing workflows where an editable, re-runnable model is more useful than a static mesh."],"supporting_citations":[],"fun_headline_variants":["Point clouds to editable Blender scripts via LLM","MeshCoder turns 3D scans into part-structured code","LLM writes editable Python from 3D point cloud","From points to programs: MeshCoder's edit-friendly code"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The point cloud alone must contain enough information to pin down a semantically decomposed program that recreates the object; if the same geometry can be validly encoded by many different programs, the model can learn only one mapping and the claimed fidelity has no unique target.","fun_headline_variants_meta":{"raw":{"variants":["Point clouds to editable Blender scripts via LLM","MeshCoder turns 3D scans into part-structured code","LLM writes editable Python from 3D point cloud","From points to programs: MeshCoder's edit-friendly code"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00019,"raw_usage":{"total_tokens":1169,"prompt_tokens":727,"completion_tokens":442,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":471,"completion_tokens_details":{"reasoning_tokens":374}},"tokens_in":471,"tokens_out":442,"duration_ms":5007,"temperature":1.0,"reasoning_tokens":374,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T18:13:20.423556+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take one mesh, generate two semantically different Blender programs that both reproduce it (for example, a hole made by boolean subtraction versus by a swept profile), sample point clouds from each, and see whether MeshCoder outputs the same code or the corresponding ground-truth code. If the output is not stable across valid programs, the part-level semantic code is not recoverable from geometry alone.","supporting_citations":[],"review_version":1}