{"id":"eeb15777-a374-4245-8a7e-f14bfc68e659","arxiv_id":"2607.17519","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"high","formal_verification":"none","parameter_count":4,"one_line_summary":"GeAR recovers more geometrically plausible and painterly-faithful 3D Gaussian reconstructions from classical paintings by separating geometry grounding from appearance restitution.","lead":"A training-free two-stage pipeline that first converts classical paintings into a geometry-oriented intermediate image, then refines a 3D Gaussian scene to restore painterly appearance, is shown to beat standard single-image 3D reconstruction baselines on a new 10,160-image benchmark. A smart reader should care because it addresses a concrete domain gap—artistic imagery violates the natural-image assumptions most 3D reconstruction networks rely on—and it introduces a reusable","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Geometry claims rest on unvalidated VEND/DDI; appearance-only refinement nearly doubles DDI, suggesting metric artifact rather than geometric improvement.","rationale":"The reader's weakest assumption focuses on the 'Sculpture' translation prompt and its alignment/physical coherence, with the unvalidated nature of VEND/DDI cited as supporting evidence. I agree that the translation alignment is an important risk, but I see the more load-bearing issue as the validity of the geometry metrics themselves. If VEND/DDI do not correspond to actual geometric plausibility, then even a perfectly aligned translation cannot support the headline geometry claim. The internal evidence from Table 3—where an appearance-only stage increases DDI by more than 50%—makes this concern concrete and suggests the metric may reward high-frequency appearance artifacts rather than depth integrity. This is a correctness risk, not a stylistic disagreement: the paper's only quantitative geometry evaluation is built on these unvalidated metrics. A synthetic ground-truth validation study would settle whether Table 1 reflects real geometric gains or metric inflation. Because the paper is otherwise structurally sound, the proposed framework is plausible, and the missing evidence is fixable, I do not move the verdict beyond the reader's CONDITIONAL assessment; I keep the verdict unchanged.","tokens_in":12844,"tokens_out":4584,"duration_ms":41929,"concrete_test":"Create a synthetic CP3D validation set: render 100+ 3D assets from Objaverse/ShapeNet with known ground-truth meshes under painterly stylization (flattened shading, weak perspective, symbolic layout). Run TRELLIS, GeAR w/o AR, and GeAR w/ AR on these renders; compute VEND/DDI as defined by the authors and standard geometric errors (Chamfer distance, normal consistency, F-score at several tau thresholds, or rendered depth L1 against true depth). Then compare rankings: if GeAR w/ AR improves DDI while its Chamfer/normal-consistency error relative to GeAR w/o AR worsens, Table 1's geometry claim is a metric artifact. Also publish VEND/DDI definitions so the metric can be independently scrutinized.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—GeAR recovers more geometrically plausible 3D scenes—rests on Table 1's VEND/DDI gains, but these metrics are not defined in the available manuscript ('Detailed definitions ... provided in the appendix') and are never validated against known geometry. The internal evidence makes the artifact risk concrete: Table 3 compares GeAR w/o AR and GeAR w/ AR on 1000 samples. Appearance Restitution (Eq. 16) is an appearance-only diffusion-editing stage, anchored to the grounded scene to suppress geometric drift. Yet it raises DDI from 6.68 to 10.24, a 53% increase. If DDI measures 'depth detail integrity,' an appearance-only edit should not nearly double it unless the metric is rewarding added high-frequency texture/opacity variation rather than real depth structure. VEND's own definition ('volumetric extent and directional diversity of surface normals') can likewise be inflated by fragmented or noisy Gaussians. Thus the quantitative evidence for the geometry half of the central claim is not trustworthy; even if the 'Sculpture' translation is perfectly alignment-preserving, the paper has not shown that the metric improvements correspond to more plausible geometry.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces Classical Painting-to-3D (CP3D), a task and benchmark for reconstructing 3D Gaussian scenes from a single classical painting while jointly maintaining geometric plausibility and fidelity to the source artwork. The proposed method, GeAR, is a training-free two-stage pipeline: Geometry Grounding first translates the painting into a more geometry-oriented image via edge-guided style transfer and illumination-ratio manipulation (Eqs. 2–13), then reconstructs a Gaussian scene with TRELLIS; Appearance Restitution then edits multi-view renders with a diffusion model and refines the Gaussians with anchor-constrained optimization (Eqs. 14–16). The authors build HeriArch, a 10,160-image benchmark, and report quantitative gains in VEND/DDI geometry metrics, MLLM appearance scores, and a 300-participant user study over SplatterImage, LGM, and TRELLIS.","tokens_in":13125,"tokens_out":6808,"duration_ms":58719,"significance":"If the claims were fully supported, the main contribution would be an effective inference-time recipe for a genuinely hard domain: classical paintings violate natural-image assumptions, and decoupling geometric recoverability from painterly appearance is a plausible design solution. HeriArch, with 10,160 images across six artistic traditions, is a potentially valuable resource, and the user study (300+ participants, including experts) is a useful qualitative complement. The comparison against generic enhancement baselines (Table 6) is also a good sanity check. However, the central quantitative evidence for geometric improvement is currently unsupported: the novelty metrics are unvalidated, the stage ablation is consistent with metric contamination by texture, and the key prompt is selected on the evaluation benchmark. The significance therefore remains conditional until these concerns are addressed.","major_comments":[{"comment":"The central claim of geometric plausibility rests entirely on VEND and DDI, whose definitions are deferred to an appendix and which are never validated against ground-truth geometry. Both definitions are of the form that can be inflated by non-geometric signal: VEND rewards volumetric extent and normal directional diversity, and DDI rewards high-frequency geometric variation. Table 3 makes the risk concrete: Appearance Restitution, an appearance-only refinement stage whose optimization is anchored to the grounded geometry by L_anchor (Eq. 16), increases DDI_avg from 6.68 to 10.24 (+53%) and VEND from 85.84 to 87.61. If geometry is held approximately constant by the anchor term, a depth-detail metric should not nearly double from texture/color editing. The authors should validate VEND/DDI on scenes with known 3D geometry (e.g., synthetic renders of 3D models with and without added texture","section":"§5.2, Table 1; §5.1, Eq. (16), Table 3"},{"comment":"The Sculpture prompt T_g is a design choice that is selected by comparing prompt styles on the HeriArch benchmark and taking the one that maximizes VEND/DDI (Table 4: Sculpture 11.05 DDI vs 2.46 for Pencil sketch). The same benchmark is then used in Table 1 to claim that GeAR outperforms the baselines. This is a selection-on-the-test-set problem and can materially inflate the reported advantage. The prompt should be fixed a priori on a validation split, or the final comparison should be on a held-out set that was not used for any prompt/ablation choice. Without this, the 'consistent outperformance' claim is not properly supported.","section":"§5.6, Table 4; Eq. (3); §5.2, Table 1"},{"comment":"No variance or confidence interval is reported for any geometry metric, even though Table 3 reports N=1000. Several differences in Table 5 (e.g., VEND 85.29 vs 85.27 vs 86.01) are much smaller than plausible sample-to-sample noise at that scale. The paper should report means with standard deviations or bootstrap CIs over at least several independent runs (or resampling of the benchmark), and state the number of runs. Without this, the reader cannot tell whether the reported gains are significant.","section":"Tables 1, 3, 5"},{"comment":"The availability of the appendix is essential: the full text states that 'Detailed definitions of VEND, DDI, and the full evaluation protocols are provided in the appendix,' but the appendix is not included in the reviewed version. Because the main quantitative results depend on these definitions, the submission is not currently reproducible. Please include the appendix (or define the metrics in the main text) in the revision.","section":"§5.1"}],"minor_comments":[{"comment":"The model is called GeAR throughout the text but GEAR in the title; unify.","section":"Title/Abstract"},{"comment":"The caption contains a formatting artifact: 'Splatter [26Image ] LGM'.","section":"Fig. 4 caption"},{"comment":"R denotes both the log-ratio illumination field and the rendered views; use different symbols to avoid confusion.","section":"Eq. (4) vs Eq. (14)"},{"comment":"The user study should report the raw counts behind Fig. 3(b) and confidence intervals for the preference percentages; the expertise distribution in Fig. 3(a) is informative but not sufficient.","section":"§5.4"},{"comment":"L_edit and L_anchor are not defined in the main text, and lambda_anc is deferred to the appendix; the losses should be specified.","section":"Eq. (16)"},{"comment":"The use of NIQE for stylized painterly content, although described as a supplementary diagnostic, deserves one sentence of justification since NIQE was not designed for non-photographic images.","section":"Table 3"},{"comment":"The paper promises that code and dataset will be released; please include a working link or anonymized repository in the revision to support reproducibility.","section":"Abstract/Conclusion"}],"recommendation":"major_revision","confidential_remarks":"To the editor: The paper has a potentially interesting idea and a useful dataset, but the main claim of geometric improvement currently rests on unvalidated metrics and a test-set-derived prompt. If the authors can provide metric validation (e.g., synthetic scenes with known geometry) and a clean validation split, I would be supportive; I would not accept the paper in its present form. The issues are substantial but fixable within the manuscript's scope."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague, look at this paper for the problem framing, not for the headline numbers. It proposes CP3D — reconstructing a 3D Gaussian scene from a single classical painting — and GeAR, a training-free two-stage pipeline: first translate the painting into a more geometry-friendly image (sculpture-like shading) and run TRELLIS; then restore the original painterly appearance via diffusion editing anchored to the grounded geometry. It also introduces HeriArch, a 10,160-image benchmark of murals, ukiyo-e, thangkas, Persian miniatures, and more. That benchmark and task definition are genuine contributions to cultural-heritage 3D.\n\nThe method is well designed. The core idea — separate geometric recoverability from appearance fidelity — is sensible. Paintings lack the shading cues monocular reconstruction expects, so converting them to a sculpture-like image before reconstruction is reasonable. Their prompt-style ablation (Sculpture vs. Pencil sketch vs. Woodcut) supports that. The appearance-restitution stage is likewise motivated.\n\nBut the central claim — that GeAR recovers more plausible geometry — rests on two metrics, VEND and DDI, whose definitions are in an appendix we can't see. They are never validated against ground truth. The smoking gun is Table 3: adding Appearance Restitution, an appearance-only stage, raises DDI from 6.68 to 10.24 — a 53% jump. If DDI measures depth-detail integrity, an appearance edit should not nearly double it. The metric is likely rewarding added high-frequency texture or opacity variation, not real depth. So the geometry gains are not demonstrated.\n\nAlso missing: error bars (Table 3 has 1000 samples, no variance), code and dataset (promised but absent), the MLLM judge is unnamed, and the user study only gives a preference histogram. There's also benchmark circularity: the Sculpture prompt and the alpha range are selected by maximizing VEND/DDI on HeriArch and then reported on the same HeriArch, with no held-out split.\n\nNone of this kills the contribution. The task and benchmark are worth building on, and the method makes sense as a baseline. But the paper overstates: the geometry half of the claim is currently unsupported. A serious referee should demand the appendix, code, a held-out split, and a validation of VEND/DDI against known geometry on a synthetic proxy (e.g., render 3D scenes with painting-like stylization and check whether the metrics track the true depths).\n\nWho this is for: researchers in heritage 3D and single-image reconstruction will find the benchmark and task formulation valuable. It deserves a serious referee but with the clear expectation of a major revision.\n\nMy recommendation: send to peer review, not desk reject, but treat the geometry claim as unproven until the metrics are validated and the artifacts are released.","headline":"A useful new task and a sensible two-stage recipe, but the quantitative evidence for the geometry claim does not hold up yet.","tokens_in":13639,"tokens_out":4278,"would_cite":true,"duration_ms":36115,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A classical painting can become a plausible 3D scene by splitting the task: first convert the flat artwork into a sculpture-like image with coherent shading, reconstruct 3D Gaussians from it, then restore the original brushwork under multi-","keywords":["3D reconstruction from single image","classical paintings","3D Gaussian splatting","cultural heritage","appearance fidelity","geometry grounding","novel view synthesis","benchmark dataset"],"falsifier":"Take a set of classical paintings whose depicted scenes have known 3D structure (for instance, paintings made from a real scene or a 3D model, or a 3D scene rendered in painterly style). Compare the depth maps recovered by GeAR from the painting with depth recovered from the original photograph. If grounding does not reduce depth error relative to direct reconstruction from the painting, the central claim is falsified.","tokens_in":12699,"feed_emoji":"🎨","tokens_out":4916,"duration_ms":42188,"temperature":0.7,"pith_summary":"The paper's central claim is that a single classical painting — despite stylized perspective, flat shading, and ambiguous depth — can be turned into a plausible, explorable 3D scene if the task is split into two stages rather than forced through one representation. In the first stage, the painting is translated into a sculpture-like version with more coherent lighting and shading cues, from which a pretrained single-view 3D reconstruction model produces stable geometry. In the second stage, that grounded scene is edited through diffusion-based multi-view appearance targets, restoring painterly texture and detail while an anchor constraint keeps the geometry from drifting. If the claim holds, the practical consequence is a training-free recipe for cultural-heritage digitization: artworks that currently defeat natural-image reconstruction pipelines can be reconstructed without retraining, and the same two-stage separation may transfer to other stylized image domains. The authors support the claim with a new benchmark of 10,160 artworks and with geometry metrics, image-quality-judge evaluations, and a user study.","feed_headline":"Classical paintings become 3D scenes in two training-free steps","feed_subtitle":"A sculpture-style shading pass stabilizes depth; a multi-view texture pass restores the original brushwork.","key_machinery":"The load-bearing mechanism is the two-stage separation itself, with the adaptive illumination-grounding operation as the concrete engine of the first stage. The method computes a log-ratio illumination field R = log(L+ε) − log(gmean(L+ε)) from the translated image, smooths it locally and at multiple scales with boundary-aware weights, then fuses the smooth and multi-scale fields with a contrast-adaptive coefficient α_ms = clip(α_base · α_target/(σ_R + ε), α_min, α_max). Recombining this grounded illumination with the original painting's reflectance (and re-adding high-frequency brushwork detail) yields an input whose shading cues are more consistent for a pretrained single-view reconstructio","core_discovery":"The discovery on the paper's own terms is that geometric recoverability and painterly fidelity are not conflicting demands on a single representation but separable stages. Geometry Grounding maps the painting to a geometry-oriented image using an edge-conditioned style translation (with a 'sculpture' prompt) plus an illumination-grounding step: it computes a log-ratio illumination field, smooths it with boundary-aware multi-scale filtering, adaptively weights the smoothing by global contrast, and recombines it with the original reflectance and high-frequency details. Feeding this grounded image to a pretrained monocular Gaussian reconstruction model stabilizes the recovered structure. Appear","pith_inferences":["A sharp reader should treat the sculpture-prompt step as an empirical bet: the claim that the translation carries 'physically consistent lighting' is asserted, not verified against geometric ground truth. A direct test — reconstructing from a photograph-like rendering of the same painted scene and comparing depth error — would settle whether grounding improves real geometry or just shape-diversity","The two-stage separation could transfer beyond paintings: any stylized single image (caricature, anime cel, stained-glass window) with weak photometric cues might benefit from the same grounding-then-restitution split, provided a 'grounding' prompt can be found that preserves layout while adding shading coherence.","The benchmark's lack of 3D ground truth means current metrics measure plausibility, not accuracy. If future work adds a small set of scenes with known geometry (e.g., a 3D scene rendered in painterly style), the whole evaluation could be anchored to actual correspondence rather than relative preference.","Releasing the dataset and code could let the community stress-test whether the method's gains hold under different diffusion-editing backbones and different single-view reconstruction models, since the pipeline is training-free and modular."],"forward_implications":["If GeAR is right, classical paintings that currently produce flat, unstable reconstructions from single-view models can be reconstructed without any training or fine-tuning, using only prompt engineering, illumination grounding, and a generic editing model.","The two-stage recipe should preserve geometry gains through the appearance stage: the paper's ablation shows geometric metrics rise after Appearance Restitution, not fall.","The method applies across a wide span of artistic traditions — murals, ukiyo-e prints, Tibetan thangkas, Persian miniatures, court-lady paintings — suggesting the approach is not tied to one painting style.","Grounding should outperform generic preprocessing (histogram equalization, retinex) because it specifically constructs geometry-compatible shading rather than merely enhancing contrast.","Appearance fidelity and geometric plausibility can be measured separately and both improved, which gives the new task a concrete evaluation protocol for future work."],"fun_headline_variants":["Two training-free steps turn classic art into 3D scenes","Separating geometry and texture revives paintings in 3D","GeAR: grounding geometry, then restoring brushwork for 3D","From canvas to 3D: a training-free two-stage pipeline","Classical paintings go 3D without model training"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The claim rests on the premise that a text-to-image model prompted with 'sculpture' turns a painting into an image whose shading and illumination are physically coherent enough that the geometry inferred from it is genuinely better, not just higher on shape-variation scores.","fun_headline_variants_meta":{"raw":{"variants":["Two training-free steps turn classic art into 3D scenes","Separating geometry and texture revives paintings in 3D","GeAR: grounding geometry, then restoring brushwork for 3D","From canvas to 3D: a training-free two-stage pipeline","Classical paintings go 3D without model training"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000551,"raw_usage":{"total_tokens":2488,"prompt_tokens":786,"completion_tokens":1702,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":530,"completion_tokens_details":{"reasoning_tokens":1615}},"tokens_in":530,"tokens_out":1702,"duration_ms":11155,"temperature":1.0,"reasoning_tokens":1615,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T17:42:32.058396+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a set of classical paintings whose depicted scenes have known 3D structure (for instance, paintings made from a real scene or a 3D model, or a 3D scene rendered in painterly style). Compare the depth maps recovered by GeAR from the painting with depth recovered from the original photograph. If grounding does not reduce depth error relative to direct reconstruction from the painting, the central claim is falsified.","supporting_citations":[],"review_version":1}