REVIEW 3 major objections 4 minor 47 references
Visual Geometry Foundation-Aware Gaussians for Single-Frame Surround-View Driving Reconstruction
T0 review · 3 major / 4 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read VGGD shifts geometric reasoning to a frozen VGGT frontend and reports the best rendering quality on the nuScenes single-frame surround-view benchmark.
desk verdict A solid incremental surround-view reconstruction paper whose rendering gain is believable but whose geometric-consistency claim rests on pseudo-depth circularity. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing objects are: the frozen VGGT geometry foundation, whose frame and global tokens encode cross-view structure and whose DINO-aligned tokens retain local appearance; the Dual-Path Neck, which routes geometry-token-only features to depth and geometry prediction while routing full tokens to appearance features, then fuses the two streams by addition; and Scale Warmup, a gated L1 loss on predicted depth active only for the first S_warm training steps, preventing scale drift under ego-pose changes without permanently constraining the depth network. These feed a hybrid pixel–volume 3DGS decoder: the pixel head consumes full points and features for image-aligned detail and far-range content, while the volume head consumes a masked, driving-centric bounded subset for 3D completion. The mechanism's role is to make the 3D support of the Gaussians reliable before any rendering loss is applied.
What would settle it
Evaluate rendered depth consistency against LiDAR point clouds or multi-view triangulated ground truth on nuScenes instead of pseudo relative depth. If VGGD's PCC advantage over Omni-Scene shrinks or reverses under true metric evaluation, or if removing Scale Warmup no longer hurts when trained with LiDAR labels, the paper's geometric-consistency claim is not supported.
Extended reading notes
Core claim
The central claim is that sparse surround-view reconstruction should shift geometric modeling from the decoding head to a pretrained geometry foundation frontend. VGGT, frozen, supplies per-view DINO-aligned, frame, and global tokens that carry transferable multi-view structure; a Dual-Path Neck then decouples a geometry-consistent path, which predicts depth and 3D support, from an appearance-aware path, which preserves texture and semantic cues for weakly observed regions. A transient Scale Warmup supervises early depth with a coarse reference for the first 2,000 steps, anchoring metric scale before end-to-end rendering refinement takes over. The hybrid pixel–volume Gaussian decoder, taken from the Omni-Scene line but fed directly by both neck paths, converts the adapted features and back-projected points into a renderable 3D Gaussian scene. The paper reports that this combination beats per-scene optimization methods, general-purpose feed-forward models, and prior driving-centric pipelines on the nuScenes single-frame protocol, with the largest gains coming when warmup and both neck paths are present.
Load-bearing premise
The result stands on the assumption that the pseudo-depth maps used for warmup supervision and for evaluating geometric consistency (from Metric3D-v2 and Depth Anything V2) are unbiased proxies for true metric 3D structure in low-overlap driving scenes.
Editorial extensions
If this is right
- Freezing a geometry foundation and adapting only the neck and decoder is enough to reach state-of-the-art single-frame surround-view rendering, so future systems can avoid training large reconstruction backbones from scratch.
- Transient warmup on coarse depth outperforms both no warmup and persistent depth supervision, indicating that scale anchoring matters mainly during early learning.
- Decoupling geometry and appearance features improves both rendering quality and depth consistency, whereas fusing them into a single direct Gaussian head trades one for the other.
- The same decoder receives better inputs from the neck; the reported gains are attributable to the frontend and the adaptation, not to a more expressive Gaussian head.
Reading between the lines
- Because the method freezes VGGT, the same adapter could in principle be attached to newer geometry foundations; if their tokens are better calibrated to driving cameras, the warmup schedule may shorten or disappear.
- The PCC evaluation is correlational and shift-invariant, so the reported geometric-consistency gains do not establish metric-scale accuracy; a LiDAR-based evaluation would test whether the warmup actually produces metric depth.
- The dual-path design suggests a more general recipe for sparse-view reconstruction: keep a geometry-dedicated feature stream and a separate appearance stream, and fuse late, rather than forcing one representation to serve both roles.
- Applying the same frontend-shift idea to dynamic surround-view reconstruction is a natural next step, since the paper explicitly leaves scene dynamics out of scope.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper proposes VGGD, a feed-forward 3D Gaussian Splatting method for single-frame surround-view driving reconstruction on nuScenes. VGGD takes six surround images, extracts frozen VGGT multi-view tokens, passes them through a dual-path neck that separates geometry and appearance features, predicts per-view depth with a transient Scale Warmup, and decodes a hybrid pixel-volume 3DGS scene. The paper reports state-of-the-art rendering quality (PSNR 24.85, SSIM 0.758, LPIPS 0.216) and improved relative geometric consistency (PCC 0.808) over Omni-Scene and other baselines, with ablations supporting Scale Warmup and the dual-path design.
Significance. The main contribution is architectural: shifting geometric reasoning from the decoder to a frozen VGGT frontend with driving-oriented adaptation. If the reported results are robust, this is a useful direction for sparse surround-view reconstruction, and the paper is transparent about its training objective, the pseudo-depth references, and the exact loss weights. The rendering comparisons use standard external benchmarks, and the internal ablations are coherent and consistent with the paper's narrative. However, the geometric-consistency headline is supported only by a pseudo-depth correlation, and the ablation evidence is obtained at a much shorter training budget than the final model. The significance is therefore conditional on stronger geometric validation and on demonstrating that the small gains over Omni-Scene are not within run-to-run variability.
major comments (3)
- [Sec. V-A.2, Sec. V-A.4, Eqs. (7) and (9)] The geometric-consistency claim is evaluated with PCC against Depth Anything V2 pseudo relative depth, while the same geometry pathway is trained against Metric3D-v2 pseudo-depth. This creates a circularity risk: the evaluation metric rewards rendered depths that resemble an off-the-shelf monocular depth prior, not metric or multi-view 3D consistency. PCC is invariant to global shift and positive scaling, so the claimed suppression of scale drift by Scale Warmup cannot be validated by PCC. I request validation against LiDAR point clouds or multi-view triangulated geometry, using scale-sensitive metrics such as absolute/relative depth error after alignment or reprojection error, or at minimum an explicit demonstration that the gains persist under scale-sensitive metrics.
- [Sec. V-D.1 and Tables I-II] All ablations are trained for 10k iterations with a different final learning rate (1e-5) than the final model (100k iterations, 1e-6). Consequently Table II validates components only in the short-training regime; it does not show that the gains of Scale Warmup or Dual-Path Neck persist at the reported deployment budget. For example, SW1-None at 10k yields 21.92 PSNR, but nothing in the paper rules out that longer training closes the gap. The ablations should be rerun at the full 100k schedule, or the main table should also be reported at the ablation budget.
- [Table I] No error bars or repeated-seed statistics are reported, and the headline gains over Omni-Scene are small (PSNR +0.58, SSIM +0.022, LPIPS -0.021, PCC +0.004). Given that the PCC difference is four-thousandths on a pseudo-depth correlation, the claim of improved geometric consistency is not supported as statistically meaningful. Please report standard deviations over multiple training seeds and, for the PCC difference, a significance test or confidence interval.
minor comments (4)
- [Sec. IV-B] The dimensions and downsampling of the three VGGT token types are not given, and the adapter architecture for Ψgeo and Ψapp is described only in words; please specify these details for reproducibility.
- [Fig. 5 and Sec. V-D] The qualitative panels labeled Geo1–Geo3 and App1–App3 are not mapped to any variant in Table II, so the figure cannot be interpreted unambiguously.
- [Algorithm 1 and Eq. (7)] The coarse depth reference \tilde D^v_t is introduced before its generation is described; the text only clarifies in Sec. V-A.4 that Metric3D-v2 pseudo-depth is used. Please move that definition earlier or add an explicit pointer.
- [Table I] The footnote says STORM is included as a spatio-temporal reference; the authors should state whether STORM was retrained or evaluated under the same single-frame protocol as the other baselines, since STORM's native input is spatio-temporal.
Circularity Check
No significant circularity: the rendering benchmark is external, and the pseudo-depth-based geometry evaluation is a validity caveat, not a derivation that reduces to the paper's own inputs.
full rationale
The paper's central novel-view rendering claim is evaluated with PSNR/SSIM/LPIPS against ground-truth target images on nuScenes (Table I), an external benchmark not produced by the model or its training losses; the architecture (VGGT frontend, Dual-Path Neck, Scale Warmup, pixel-volume decoder) is a feed-forward construction whose components are compared by ablations under the same external rendering metrics. The only potential loop is geometric consistency: training depth losses use Metric3D-v2 pseudo-depth (Sec. V-A.4, Eqs. 7 and 9) while the PCC metric compares rendered depth to Depth Anything V2 pseudo relative depth (Sec. V-A.2). These are different monocular estimators and different target views, so the evaluation is not the same quantity as the training target by construction; the concern is that pseudo-depth is not metric ground truth and PCC is scale-invariant, which weakens the geometric-consistency claim as external validation but does not make the derivation circular. The self-citation to VGD [6] appears only as related work and is not load-bearing. Overall, the main results are self-contained against independent rendering benchmarks, so no circular step is established.
Assumptions & free parameters
free parameters (3)
- Loss weights lambda_perc, lambda_dep, lambda_vol_dep, lambda_warm =
0.05, 0.01, 0.01, 0.01
- Scale warmup duration S_warm =
2000 iterations
- Driving-centric cuboid range B =
not reported
assumptions (5)
- domain assumption VGGT tokens transfer to surround-view driving geometry.
- domain assumption Metric3D-v2 pseudo depths are accurate enough to anchor scale and supervise rendered depth.
- domain assumption Depth Anything V2 pseudo relative depth is a valid proxy for 3D geometric consistency.
- standard math Differentiable 3DGS rendering and standard back-projection are correct.
- domain assumption The single-frame surround-view task is sufficiently well-posed for metric structure inference.
Cite this review
Pith. "Pith review of Visual Geometry Foundation-Aware Gaussians for Single-Frame Surround-View Driving Reconstruction." pith.science (2026). https://pith.science/paper/47ER6MMO
@misc{pith2026260810682,
author = {Pith},
title = {Pith review of: Visual Geometry Foundation-Aware Gaussians for Single-Frame Surround-View Driving Reconstruction},
year = {2026},
howpublished = {\url{https://pith.science/paper/47ER6MMO}},
note = {Machine review of arXiv:2608.10682}
}
read the original abstract
Single-frame surround-view reconstruction faces severe geometric instability and rendering artifacts due to minimal inter-camera overlap. While existing methods rely on complex decoders or auxiliary cues, they remain bottlenecked by the weak geometric capacity of upstream features. We argue that leveraging pretrained visual geometry priors strengthens upstream representations and alleviates the geometric ambiguity in sparse surround views. To this end, we propose VGGD, a visual geometry foundation-aware 3D Gaussian Splatting framework for feed-forward surround-view driving reconstruction, which shifts geometric modeling to the frontend and adapts foundation priors to the driving camera setting. First, VGGD leverages VGGT to provide transferable multi-view geometric prior tokens. Next, we introduce a Dual-Path Neck to decouple geometry-consistent and appearance-aware representations, improving appearance completion in weakly observed regions. We further apply Scale Warmup to stabilize early geometry learning and suppress scale drift under ego-pose changes. Finally, we use a hybrid pixel--volume Gaussian decoder to produce a renderable 3D Gaussian scene for novel-view synthesis. Experiments on the nuScenes single-frame benchmark show that VGGD achieves the best overall rendering quality among the compared methods and improves relative geometric consistency.
Figures
Figures from the paper (2 more)
Reference graph
Works this paper leans on
-
[1]
Omni-scene: Omni-gaussian representation for ego-centric sparse-view scene reconstruction,
D. Wei, Z. Li, and P. Liu, “Omni-scene: Omni-gaussian representation for ego-centric sparse-view scene reconstruction,” inIEEE Conf. Comput. Vis. Pattern Recog., 2025, pp. 22 317–22 327
work page 2025
-
[2]
Q. Tian, X. Tan, Y . Xie, and L. Ma, “Drivingforward: Feed-forward 3d gaussian splatting for driving scene reconstruction from flexible surround- view input,” inProc. AAAI Conf. Artif. Intell., vol. 39, no. 7, 2025, pp. 7374–7382
work page 2025
-
[3]
Storm: Spatio-temporal recon- struction model for large-scale outdoor scenes,
J. Yang, J. Huang, B. Ivanovic, Y . Chen, Y . Wang, B. Li, Y . You, A. Sharma, M. Igl, P. Karkuset al., “Storm: Spatio-temporal recon- struction model for large-scale outdoor scenes,” inInt. Conf. Learn. Represent., vol. 2025, 2025, pp. 50 446–50 465
work page 2025
-
[4]
Y . Ye, Z. Zhang, J. Lin, S. Sun, C. Peng, and W. Gao, “Autodrive-p3: Uni- fied chain of perception–prediction–planning thought via reinforcement fine-tuning,”arXiv preprint arXiv:2603.28116, 2026
arXiv 2026
-
[5]
nuscenes: A multimodal dataset for autonomous driving,
H. Caesar, V . Bankiti, A. H. Lang, S. V ora, V . E. Liong, Q. Xu, A. Krishnan, Y . Pan, G. Baldan, and O. Beijbom, “nuscenes: A multimodal dataset for autonomous driving,” inIEEE Conf. Comput. Vis. Pattern Recog., 2020, pp. 11 621–11 631
work page 2020
-
[6]
Vgd: Visual geometry gaussian splatting for feed-forward surround-view driving reconstruction,
J. Lin, K. Wang, S. Wang, S. Fan, G. Li, and W. Gao, “Vgd: Visual geometry gaussian splatting for feed-forward surround-view driving reconstruction,”arXiv preprint arXiv:2510.19578, 2025
arXiv 2025
-
[7]
pixelsplat: 3d gaussian splats from image pairs for scalable generalizable 3d reconstruction,
D. Charatan, S. L. Li, A. Tagliasacchi, and V . Sitzmann, “pixelsplat: 3d gaussian splats from image pairs for scalable generalizable 3d reconstruction,” inIEEE Conf. Comput. Vis. Pattern Recog., 2024, pp. 19 457–19 467
work page 2024
-
[8]
Mvsplat: Efficient 3d gaussian splatting from sparse multi-view images,
Y . Chen, H. Xu, C. Zheng, B. Zhuang, M. Pollefeys, A. Geiger, T.-J. Cham, and J. Cai, “Mvsplat: Efficient 3d gaussian splatting from sparse multi-view images,” inEur . Conf. Comput. Vis., 2024, pp. 370–386
work page 2024
Show all 47 references
-
[9]
3d gaussian splatting for real-time radiance field rendering
B. Kerbl, G. Kopanas, T. Leimk ¨uhler, G. Drettakiset al., “3d gaussian splatting for real-time radiance field rendering.”ACM Trans. Graph., vol. 42, no. 4, pp. 139–1, 2023
2023
-
[10]
Depth anything v2,
L. Yang, B. Kang, Z. Huang, Z. Zhao, X. Xu, J. Feng, and H. Zhao, “Depth anything v2,”Adv. Neural Inform. Process. Syst., vol. 37, pp. 21 875–21 911, 2024
2024
-
[11]
Metric3d v2: A versatile monocular geometric foundation model for zero-shot metric depth and surface normal estimation,
M. Hu, W. Yin, C. Zhang, Z. Cai, X. Long, H. Chen, K. Wang, G. Yu, C. Shen, and S. Shen, “Metric3d v2: A versatile monocular geometric foundation model for zero-shot metric depth and surface normal estimation,”IEEE Trans. Pattern Anal. Mach. Intell., vol. 46, no. 12, pp. 10 57...
2024
-
[12]
Vggt: Visual geometry grounded transformer,
J. Wang, M. Chen, N. Karaev, A. Vedaldi, C. Rupprecht, and D. Novotny, “Vggt: Visual geometry grounded transformer,” inIEEE Conf. Comput. Vis. Pattern Recog., 2025, pp. 5294–5306
2025
-
[13]
Dust3r: Geometric 3d vision made easy,
S. Wang, V . Leroy, Y . Cabon, B. Chidlovskii, and J. Revaud, “Dust3r: Geometric 3d vision made easy,” inIEEE Conf. Comput. Vis. Pattern Recog., 2024, pp. 20 697–20 709
2024
-
[14]
Grounding image matching in 3d with mast3r,
V . Leroy, Y . Cabon, and J. Revaud, “Grounding image matching in 3d with mast3r,” inEur . Conf. Comput. Vis., 2024, pp. 71–91
2024
-
[15]
Emerging properties in self-supervised vision transformers,
M. Caron, H. Touvron, I. Misra, H. J ´egou, J. Mairal, P. Bojanowski, and A. Joulin, “Emerging properties in self-supervised vision transformers,” inInt. Conf. Comput. Vis., 2021, pp. 9650–9660
2021
-
[16]
Dinov2: Learning robust visual features without supervision,
M. Oquab, T. Darcet, T. Moutakanni, H. V o, M. Szafraniec, V . Khalidov, P. Fernandez, D. Haziza, F. Massa, A. El-Noubyet al., “Dinov2: Learning robust visual features without supervision,”arXiv preprint arXiv:2304.07193, 2023
2023 arXiv
-
[17]
Masked autoencoders are scalable vision learners,
K. He, X. Chen, S. Xie, Y . Li, P. Doll ´ar, and R. Girshick, “Masked autoencoders are scalable vision learners,” inIEEE Conf. Comput. Vis. Pattern Recog., 2022, pp. 16 000–16 009
2022
-
[18]
Learning transferable visual models from natural language supervision,
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clarket al., “Learning transferable visual models from natural language supervision,” inInt. Conf. Mach. Learn., 2021, pp. 8748–8763
2021
-
[19]
Vision transformers for dense prediction,
R. Ranftl, A. Bochkovskiy, and V . Koltun, “Vision transformers for dense prediction,” inInt. Conf. Comput. Vis., 2021, pp. 12 179–12 188
2021
-
[20]
Depth anything: Unleashing the power of large-scale unlabeled data,
L. Yang, B. Kang, Z. Huang, X. Xu, J. Feng, and H. Zhao, “Depth anything: Unleashing the power of large-scale unlabeled data,” inIEEE Conf. Comput. Vis. Pattern Recog., 2024, pp. 10 371–10 381
2024
-
[21]
Sparsenerf: Distilling depth ranking for few-shot novel view synthesis,
G. Wang, Z. Chen, C. C. Loy, and Z. Liu, “Sparsenerf: Distilling depth ranking for few-shot novel view synthesis,” inInt. Conf. Comput. Vis., 2023, pp. 9065–9076
2023
-
[22]
Dngaussian: Optimizing sparse-view 3d gaussian radiance fields with global-local depth normalization,
J. Li, J. Zhang, X. Bai, J. Zheng, X. Ning, J. Zhou, and L. Gu, “Dngaussian: Optimizing sparse-view 3d gaussian radiance fields with global-local depth normalization,” inIEEE Conf. Comput. Vis. Pattern Recog., 2024, pp. 20 775–20 785
2024
-
[23]
Nerf: Representing scenes as neural radiance fields for view synthesis,
B. Mildenhall, P. P. Srinivasan, M. Tancik, J. T. Barron, R. Ramamoorthi, and R. Ng, “Nerf: Representing scenes as neural radiance fields for view synthesis,”Commun. ACM, vol. 65, no. 1, pp. 99–106, 2021
2021
-
[24]
Fsgs: Real-time few-shot view synthesis using gaussian splatting,
Z. Zhu, Z. Fan, Y . Jiang, and Z. Wang, “Fsgs: Real-time few-shot view synthesis using gaussian splatting,” inEur . Conf. Comput. Vis., 2024, pp. 145–163
2024
-
[25]
Sugar: Surface-aligned gaussian splatting for efficient 3d mesh reconstruction and high-quality mesh rendering,
A. Gu ´edon and V . Lepetit, “Sugar: Surface-aligned gaussian splatting for efficient 3d mesh reconstruction and high-quality mesh rendering,” inIEEE Conf. Comput. Vis. Pattern Recog., 2024, pp. 5354–5363
2024
-
[26]
FieldGS: Unsigned distance fields in gaussian splatting for enhanced surface reconstruction,
D. Wu, L. Zhou, L. Liu, F. Qu, W. Zhang, L. Song, and M. Wang, “FieldGS: Unsigned distance fields in gaussian splatting for enhanced surface reconstruction,”IEEE Trans. Multimedia, pp. 1–11, 2026
2026
-
[27]
Gaussian splatting slam,
H. Matsuki, R. Murai, P. H. Kelly, and A. J. Davison, “Gaussian splatting slam,” inIEEE Conf. Comput. Vis. Pattern Recog., 2024, pp. 18 039– 18 048
2024
-
[28]
MSA-Splatting: Multi- scale adaptive gaussian splatting for high-fidelity view synthesis,
Y . Zhao, G. Chen, B. Wu, Q. Jin, and T. Zeng, “MSA-Splatting: Multi- scale adaptive gaussian splatting for high-fidelity view synthesis,”IEEE Trans. Multimedia, vol. 28, pp. 5357–5367, 2026
2026
-
[29]
4d gaussian splatting for real-time dynamic scene rendering,
G. Wu, T. Yi, J. Fang, L. Xie, X. Zhang, W. Wei, W. Liu, Q. Tian, and X. Wang, “4d gaussian splatting for real-time dynamic scene rendering,” inIEEE Conf. Comput. Vis. Pattern Recog., 2024, pp. 20 310–20 320
2024
-
[30]
Spacetime gaussian feature splatting for real-time dynamic view synthesis,
Z. Li, Z. Chen, Z. Li, and Y . Xu, “Spacetime gaussian feature splatting for real-time dynamic view synthesis,” inIEEE Conf. Comput. Vis. Pattern Recog., 2024, pp. 8508–8520
2024
-
[31]
4DGStream: Variable bitrate dynamic gaussian splatting streaming,
Z. Liang, D. Zhang, L. Shen, M. Zhang, J. Zhang, B. Ju, M. Dasari, F. Wang, and J. Liu, “4DGStream: Variable bitrate dynamic gaussian splatting streaming,”IEEE Trans. Multimedia, pp. 1–15, 2026
2026
-
[32]
Unisplat: Unified spatio-temporal fusion via 3d latent scaffolds for dynamic driving scene reconstruction,
C. Shi, S. Shi, X. Lyu, C. Liu, K. Sheng, B. Zhang, and L. Jiang, “Unisplat: Unified spatio-temporal fusion via 3d latent scaffolds for dynamic driving scene reconstruction,”arXiv preprint arXiv:2511.04595, 2025
2025
-
[33]
Vgocc: Learning visual-geometric gaussians for vision-centric 3d driving occupancy prediction,
J. Lin, X. Guo, K. Wang, Y . Ye, X. Liang, Y . Peng, and W. Gao, “Vgocc: Learning visual-geometric gaussians for vision-centric 3d driving occupancy prediction,”arXiv preprint arXiv:2607.18078, 2026
2026 arXiv
-
[34]
Omninwm: Omniscient driving navigation world models,
B. Li, Z. Ma, D. Du, B. Peng, Z. Liang, Z. Liu, C. Ma, Y . Jin, H. Zhao, W. Zenget al., “Omninwm: Omniscient driving navigation world models,” arXiv preprint arXiv:2510.18313, 2025
2025 arXiv
-
[35]
Gem-occ: From visual geometry evidence to embodied semantic occupancy memory,
H. Zhu, B. Li, X. Guo, H. Liu, B. Peng, M. Yuan, X. Jin, W. Zeng, and C. W. Chen, “Gem-occ: From visual geometry evidence to embodied semantic occupancy memory,”arXiv preprint arXiv:2607.05543, 2026
2026 arXiv
-
[36]
Emernerf: Emergent spatial-temporal scene decomposition via self-supervision,
J. Yang, B. Ivanovic, O. Litany, X. Weng, S. W. Kim, B. Li, T. Che, D. Xu, S. Fidler, M. Pavoneet al., “Emernerf: Emergent spatial-temporal scene decomposition via self-supervision,” inInt. Conf. Learn. Represent., vol. 2024, 2024, pp. 16 739–16 766
2024
-
[37]
Periodic vibration gaussian: Dynamic urban scene reconstruction and real-time rendering,
Y . Chen, C. Gu, J. Jiang, X. Zhu, and L. Zhang, “Periodic vibration gaussian: Dynamic urban scene reconstruction and real-time rendering,” Int. J. Comput. Vis., vol. 134, no. 3, p. 83, 2026
2026
-
[38]
Deformable 3d gaussians for high-fidelity monocular dynamic scene reconstruction,
Z. Yang, X. Gao, W. Zhou, S. Jiao, Y . Zhang, and X. Jin, “Deformable 3d gaussians for high-fidelity monocular dynamic scene reconstruction,” inIEEE Conf. Comput. Vis. Pattern Recog., 2024, pp. 20 331–20 341
2024
-
[39]
Learning to render novel views from wide-baseline stereo pairs,
Y . Du, C. Smith, A. Tewari, and V . Sitzmann, “Learning to render novel views from wide-baseline stereo pairs,” inIEEE Conf. Comput. Vis. Pattern Recog., 2023, pp. 4970–4980
2023
-
[40]
Murf: Multi-baseline radiance fields,
H. Xu, A. Chen, Y . Chen, C. Sakaridis, Y . Zhang, M. Pollefeys, A. Geiger, and F. Yu, “Murf: Multi-baseline radiance fields,” inIEEE Conf. Comput. Vis. Pattern Recog., 2024, pp. 20 041–20 050
2024
-
[41]
Lgm: Large multi-view gaussian model for high-resolution 3d content creation,
J. Tang, Z. Chen, X. Chen, T. Wang, G. Zeng, and Z. Liu, “Lgm: Large multi-view gaussian model for high-resolution 3d content creation,” in Eur . Conf. Comput. Vis., 2024, pp. 1–18
2024
-
[42]
Gs-lrm: Large reconstruction model for 3d gaussian splatting,
K. Zhang, S. Bi, H. Tan, Y . Xiangli, N. Zhao, K. Sunkavalli, and Z. Xu, “Gs-lrm: Large reconstruction model for 3d gaussian splatting,” inEur . Conf. Comput. Vis., 2024, pp. 1–19
2024
-
[43]
Scube: Instant large-scale scene reconstruction using voxsplats,
X. Ren, Y . Lu, H. Liang, Z. Wu, H. Ling, M. Chen, S. Fidler, F. Williams, and J. Huang, “Scube: Instant large-scale scene reconstruction using voxsplats,”Adv. Neural Inform. Process. Syst., vol. 37, pp. 97 670–97 698, 2024
2024
-
[44]
Image quality assessment: from error visibility to structural similarity,
Z. Wang, A. C. Bovik, H. R. Sheikh, and E. P. Simoncelli, “Image quality assessment: from error visibility to structural similarity,”IEEE Trans. Image Process., vol. 13, no. 4, pp. 600–612, 2004
2004
-
[45]
The unreasonable effectiveness of deep features as a perceptual metric,
R. Zhang, P. Isola, A. A. Efros, E. Shechtman, and O. Wang, “The unreasonable effectiveness of deep features as a perceptual metric,” in IEEE Conf. Comput. Vis. Pattern Recog., 2018, pp. 586–595
2018
-
[46]
Pearson correlation coefficient,
J. Benesty, J. Chen, Y . Huang, and I. Cohen, “Pearson correlation coefficient,” inNoise reduction in speech processing, 2009, pp. 1–4
2009
-
[47]
4d driving scene generation with stereo forcing,
H. Lu, Z. Ma, G. Jiang, W. Ge, B. Li, Y . Cai, W. Zheng, Y . Zhang, and Y . Chen, “4d driving scene generation with stereo forcing,”arXiv preprint arXiv:2509.20251, 2025
2025
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.