Pith. sign in

REVIEW 2 major objections 4 references

Shadows visible in input images guide a 3D completion model to recover geometry of unseen urban regions and enable consistent relighting.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.3

2026-06-30 13:17 UTC pith:WNLHYM4K

load-bearing objection SRUG pairs shadow-guided 3D completion with iterative LMM material decomposition for urban relighting, but the shadow-to-geometry step remains the weakest link under sparse inputs. the 2 major comments →

arxiv 2605.24700 v3 pith:WNLHYM4K submitted 2026-05-23 cs.CV cs.GR

SRUG: Shadow-Guided Relightable Urban Scene with Generation Model

classification cs.CV cs.GR
keywords relightable urban scenesshadow-guided completionmaterial decomposition3D scene generationphysically-based relightingnovel view synthesisunbounded scenes
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Urban scenes captured in images or videos often leave large areas unobserved, yet those hidden parts cast shadows that affect visible surfaces. The paper introduces a framework that uses observed shadows to steer a generative 3D model toward filling in missing geometry so that new shadows remain physically plausible. It adds an iterative loop that consults a large material model to decompose surface properties despite sparse views and complex lighting. A physically-based illumination model then combines the completed geometry and materials to support relighting under new conditions. The result is improved novel-view synthesis and relighting quality compared with prior methods that ignore invisible geometry.

Core claim

SRUG recovers unobserved geometry in unbounded urban scenes by feeding visible shadows into a 3D completion model, then decomposes materials through repeated supervision from a large material model, and finally applies a physically-based lighting model that accounts for complex urban illumination to produce reliable relighting.

What carries the argument

Shadow-guided 3D completion model that uses observed shadows to constrain generation of invisible geometry, paired with an iterative material decomposition loop driven by large-material-model supervision and a physically-based lighting model.

Load-bearing premise

Visible shadows in the input images contain enough information to accurately determine the shape of regions the camera never sees.

What would settle it

A controlled capture of an urban scene with known hidden geometry where the shadows predicted after completion and relighting deviate measurably from ground-truth shadows under changed lighting.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • Invisible regions receive geometry that produces shadows consistent with the input images.
  • Material properties can be recovered reliably even when input views are sparse and illumination is complex.
  • The combined geometry and materials support relighting under arbitrary new light conditions.
  • Both novel-view synthesis and relighting quality exceed those of prior methods that do not model hidden geometry.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The same shadow-to-geometry link could be tested on non-urban unbounded scenes such as forests or construction sites.
  • If the 3D completion model is swapped for a different generative backbone, the shadow guidance and lighting model may transfer directly.
  • Video input could supply additional shadow motion cues that further constrain the completion step.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 0 minor

Summary. The paper proposes SRUG, a framework for creating relightable urban scenes from sparse images or videos. It uses visible shadows to guide a 3D completion model for recovering geometry in unobserved regions, applies an iterative material decomposition scheme supervised by a large material model (LMM), and introduces a physically-based lighting model for complex urban illumination. The abstract claims that extensive quantitative evaluations and visual comparisons show outperformance over existing methods in novel view synthesis and relighting tasks.

Significance. If the central claims hold with rigorous validation, the work could advance practical relightable reconstruction for unbounded urban environments by explicitly handling invisible geometry via shadow cues and mitigating material ambiguities under sparse views. However, the absence of any reported metrics, datasets, ablations, or implementation details in the provided description prevents assessment of whether these advances are realized.

major comments (2)
  1. [Abstract / method description] The central claim that shadows from invisible regions can guide a 3D completion model to recover accurate geometry (rather than any of the many 3D configurations consistent with the same 2D shadow boundaries) is load-bearing for the entire pipeline, yet the abstract supplies no disambiguation mechanism, validation experiments, or quantitative evidence that the completion prior suffices under sparse views in unbounded scenes. This directly matches the weakest assumption identified in the stress-test note.
  2. [Abstract] The abstract asserts quantitative outperformance in novel view synthesis and relighting but provides no metrics, datasets, baselines, or ablation details. Without these, the claim of superiority cannot be evaluated and is therefore unsupported in the manuscript as presented.

Simulated Author's Rebuttal

2 responses · 0 unresolved

We thank the referee for the constructive feedback on the abstract and central claims. We address each major comment below and will revise the manuscript accordingly.

read point-by-point responses
  1. Referee: [Abstract / method description] The central claim that shadows from invisible regions can guide a 3D completion model to recover accurate geometry (rather than any of the many 3D configurations consistent with the same 2D shadow boundaries) is load-bearing for the entire pipeline, yet the abstract supplies no disambiguation mechanism, validation experiments, or quantitative evidence that the completion prior suffices under sparse views in unbounded scenes. This directly matches the weakest assumption identified in the stress-test note.

    Authors: The abstract is high-level by design. The full manuscript details the disambiguation via the generation model, which enforces physical consistency between observed shadows and completed geometry under the shadow-guided prior; this is further supported by the iterative LMM-based material decomposition. Quantitative validation appears in the experiments section with metrics on novel view synthesis and relighting under sparse views. We will revise the abstract to briefly note the generation model for disambiguation and reference the validation experiments. revision: yes

  2. Referee: [Abstract] The abstract asserts quantitative outperformance in novel view synthesis and relighting but provides no metrics, datasets, baselines, or ablation details. Without these, the claim of superiority cannot be evaluated and is therefore unsupported in the manuscript as presented.

    Authors: The abstract summarizes results whose details (specific metrics, datasets, baselines, and ablations) are reported in the full manuscript body. However, the referee correctly notes that these are absent from the provided abstract text. We will revise the abstract to include key quantitative highlights or qualify the outperformance claim to ensure it is supported within the abstract itself. revision: yes

Circularity Check

0 steps flagged

No circularity: framework components described without equations or self-referential reductions

full rationale

The provided abstract and description outline a proposed framework (shadow-guided 3D completion, iterative LMM-based material decomposition, physically-based lighting model) but contain no equations, parameter-fitting steps, predictions derived from fitted inputs, or self-citations that reduce the central claims to their own inputs by construction. No load-bearing derivation chain is exhibited that matches any of the enumerated circularity patterns. The method is presented as a novel combination of existing ideas applied to urban relighting, with no evidence of self-definitional loops or renamed known results in the given text.

Axiom & Free-Parameter Ledger

0 free parameters · 0 axioms · 0 invented entities

No free parameters, axioms, or invented entities can be identified from the abstract alone; the method description remains at the level of high-level components without mathematical or modeling details.

pith-pipeline@v0.9.1-grok · 5768 in / 1136 out tokens · 47885 ms · 2026-06-30T13:17:12.934621+00:00 · methodology

0 comments
read the original abstract

Creating relightable urban scenes from images or videos is widely useful but highly ill-posed. Urban environments are typically unbounded and extend beyond the visible regions. As a result, many portions of the scene remain unobserved, yet these invisible regions can cast shadows onto visible areas. Reasonably modeling shadows cast by such invisible regions is challenging and poses a significant obstacle to creating relightable urban scenes. At the same time, sparse input views and complex illumination conditions further complicate relighting, as they introduce severe ambiguities in material decomposition. In this paper, we propose Shadow-guided Relightable Urban Scene with Generation model (SRUG), a novel framework designed to address relighting challenges in urban scenes. SRUG leverages shadows to guide a 3D completion model for recovering the geometry of invisible regions, promoting the synthesis of physically reasonable shadows. In addition, SRUG employs an iterative material decomposition scheme that applies the large material model (LMM) to provide material supervision and iteratively decompose the scene's material properties, enabling robust material decomposition. Building upon these components, we introduce a physically-based lighting model that captures the complex illumination of urban scenes and supports reliable relighting. Extensive quantitative evaluations and visual comparisons demonstrate that our method outperforms existing approaches in both novel view synthesis and relighting tasks.

Figures

Figures reproduced from arXiv: 2605.24700 by Beibei Wang, Jian Yang, Jin Xie, Yonghao Zhao, Zexin Yin.

Figure 1
Figure 1. Figure 1: We propose SRUG, a novel framework for constructing relightable urban scenes from multi-view images or videos. SRUG reconstructs 3D scene [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Overview of the SRUG framework. We first initialize a Gaussian-based scene representation, including the geometry and appearance of visible regions within the scene. Based on the initialized Gaussians, we construct relightable urban scenes through two key components: (1) a shadow-guided invisible geometry completion module. It employs differentiable shadow mapping to use shadows as supervisory signals for … view at source ↗
Figure 3
Figure 3. Figure 3: In contrast to standard shadow mapping, DGSM replaces the binary [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Visual comparison of NVS, relighting, and material–lighting decomposition on the KITTI-360 dataset. Our method effectively mitigates the lighting [PITH_FULL_IMAGE:figures/full_fig_p006_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Ablation studies of LMM-based material decomposition on the KITTI [PITH_FULL_IMAGE:figures/full_fig_p007_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Comparison of NVS, relighting, and material decomposition on the synthetic dataset. Under novel lighting conditions, our method achieves the most [PITH_FULL_IMAGE:figures/full_fig_p010_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: We conduct an ablation study of the shadow-guided invisible geometry completion module on both synthetic datasets and KITTI-360. Incorporating [PITH_FULL_IMAGE:figures/full_fig_p010_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: Comparison of NVS and material decomposition on the TandT100 dataset (left part) and the TandT50 dataset (right part). [PITH_FULL_IMAGE:figures/full_fig_p011_8.png] view at source ↗
Figure 9
Figure 9. Figure 9: Comparison of NVS and material decomposition on the Waymo dataset. Existing methods, including GS-IR, GShader, and R3DG, struggle to handle [PITH_FULL_IMAGE:figures/full_fig_p011_9.png] view at source ↗
Figure 10
Figure 10. Figure 10: Visual ablation of the iterative material update scheme on the syn [PITH_FULL_IMAGE:figures/full_fig_p016_10.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

4 extracted references · 4 canonical work pages

  1. [1]

    Mark Boss, Raphael Braun, Varun Jampani, Jonathan T Barron, Ce Liu, and Hendrik Lensch

    Mip-NeRF 360: Unbounded Anti-Aliased Neural Radiance Fields.CVPR(2022). Mark Boss, Raphael Braun, Varun Jampani, Jonathan T Barron, Ce Liu, and Hendrik Lensch. 2021a. Nerd: Neural reflectance decomposition from image collections. In Proceedings of the IEEE/CVF International Conference on Computer Vision. 12684– 12694. Mark Boss, Varun Jampani, Raphael Bra...

  2. [2]

    Unirelight: Learning joint decomposition and synthesis for video relighting.arXiv preprint arXiv:2506.15673, 2025

    Shadow map filtering with gaussian shadow maps. InProceedings of the 10th International Conference on Virtual Reality Continuum and Its Applications in Industry. 75–82. Ayaan Haque, Matthew Tancik, Alexei A Efros, Aleksander Holynski, and Angjoo Kanazawa. 2023. Instruct-nerf2nerf: Editing 3d scenes with instructions. InProceed- ings of the IEEE/CVF intern...

  3. [3]

    Multi-view relighting using a geometry-aware network.ACM Trans. Graph. 38, 4 (2019), 78–1. Ben Poole, Ajay Jain, Jonathan T Barron, and Ben Mildenhall. 2023. DreamFusion: Text-to-3D using 2D Diffusion. InThe Eleventh International Conference on Learning Representations. Albert Pumarola, Enric Corona, Gerard Pons-Moll, and Francesc Moreno-Noguer. 2021. D-n...

  4. [4]

    2024 , publisher =

    Pixelwise View Selection for Unstructured Multi-View Stereo. InEuropean Conference on Computer Vision (ECCV). Pratul P Srinivasan, Boyang Deng, Xiuming Zhang, Matthew Tancik, Ben Mildenhall, and Jonathan T Barron. 2021. Nerv: Neural reflectance and visibility fields for relighting and view synthesis. InProceedings of the IEEE/CVF Conference on Computer Vi...