Pith. sign in

REVIEW 4 major objections 5 minor 2 cited by

SEED4D: A Synthetic Ego--Exo Dynamic 4D Data Generator, Driving Dataset and Benchmark

T0 review · 4 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read SEED4D is a large-scale synthetic ego-exo dynamic 4D dataset and generator for autonomous driving, built so models can learn 3D and 4D scene reconstruction from both driver and outside cameras.

desk verdict A genuinely useful synthetic ego-exo driving dataset and generator, saddled with a benchmark section whose rankings should not be trusted until the protocol is fixed. read the letter →

arxiv 2412.00730 v2 pith:6TG6OYMN submitted 2024-12-01 cs.CV

classification cs.CV
keywords syntheticdataego-exoviews4Dreconstructionautonomousdrivingnovelviewsynthesisfew-image-to-3DLiDARbenchmark
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tries to establish that the missing ingredient for joint 3D and 4D perception in driving—a large collection of dynamic urban scenes captured at the same time from egocentric (in-vehicle) and exocentric (outside-vehicle) cameras—can be produced synthetically at scale. It presents an open, customizable generator built on an urban driving simulator, plus two datasets generated with it: 2,002 static scenes with 212k images and 10,458 dynamic trajectories with 16.8M images, every frame carrying ground-truth depth, flow, segmentation, and LiDAR. The paper then defines benchmarks for few-image-to-3D reconstruction (building a 3D scene from a handful of images), novel view synthesis, and monocular depth estimation, and leaves 4D prediction as an open challenge that the dynamic dataset is sized to support.

What carries the argument

The load-bearing object is the data generator: a customizable pipeline built on an open urban driving simulator that can place any number of cameras on any vehicle or anywhere in the scene and record them over time. Exocentric cameras are arranged on a half-sphere around each vehicle using a spherical Fibonacci lattice, a point distribution scheme that spaces the viewpoints evenly; every exo camera keeps a fixed relative pose to its vehicle. Camera poses are exported in a radiance-field-friendly format with intrinsics, extrinsics, and distortion, so the generated scenes drop directly into existing radiance-field tooling. Two datasets produced by this generator carry the argument: a static set built for few-image-to-3D and a dynamic set of 10-second trajectories built for temporal reconstruction and prediction.

What would settle it

Train a few-image-to-3D model on SEED4D and evaluate it zero-shot on real street images with known geometry; if its reconstruction and depth error is no better than a model trained on unrelated real images, the claim that this synthetic ego-exo data closes a real supervision gap is falsified.

Watch

Extended reading notes

Core claim

The paper's central claim is that no existing autonomous-driving dataset offers the combination of complex, dynamic, multi-view ego-exo data, and that SEED4D is the first large-scale answer. The paper demonstrates the claim by defining benchmark protocols that only such data make possible: outward-facing ego images serve as input, while 100 inward-facing exo images supervise novel view synthesis on the static scenes. Scores are reported for per-scene radiance-field methods, monocular metric depth estimators, and few-image-to-3D models; the benchmarks are demonstrations of the new evaluation regime rather than assertions about which method is best. The dynamic dataset extends the same idea through time, providing 10-second multi-view trajectories intended to push reconstruction methods toward actual 4D prediction.

Load-bearing premise

The value of the whole resource rests on the simulator's synthetic scenes being a workable stand-in for real urban driving—both in how they look and in how vehicles move—so that what models learn on SEED4D transfers to real roads.

Editorial extensions

If this is right

  • Few-image-to-3D models can be trained with ego views as input and exo views as supervision, a protocol no current real-world driving dataset supports.
  • The dynamic set provides a common 10-second multi-view testbed for 4D reconstruction and forecasting, a task the paper notes has no vision-based method ready to run on it yet.
  • Because the generator can reproduce the camera geometry of several established driving sensor suites, experiments on SEED4D can be configured to match familiar real-world hardware.
  • Every image comes with pixel-aligned ground truth, so depth, flow, and segmentation models can be evaluated on exactly the same scenes and viewpoints as reconstruction models.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The generator's camera-placement machinery is not tied to road scenes, so the same tool could produce ego-exo data for infrastructure cameras, pedestrian views, or off-road environments, an extension the paper leaves to future work.
  • A direct test of the resource's value would be pre-training a few-image-to-3D model on SEED4D and measuring its zero-shot transfer to real road images, with and without style transfer.
  • If 4D prediction matures, the dynamic dataset's fixed 10-second trajectories and dense ground truth make it a natural controlled setting for comparing appearance-based video predictors with explicit 3D reconstruction methods.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper introduces SEED4D, a CARLA-based data generator for synthetic ego-exo dynamic 4D driving data, together with two released datasets: a static 3D dataset of 2,002 scenes with about 212k images, and a dynamic 4D dataset of roughly 10.5k trajectories with about 16.8M images. Each image is accompanied by camera poses, depth, semantic and instance segmentation, optical flow, and LiDAR, and the data generator supports camera setups mimicking NuScenes, KITTI360, and Waymo. The authors also benchmark existing methods for multi-view novel view synthesis, monocular metric depth estimation, and single-shot few-image-to-3D reconstruction on the static dataset. The central claim is that SEED4D fills a gap: no existing autonomous-driving dataset provides large-scale multi-view ego-exo data suitable for supervising few-image-to-3D and 4D prediction models.

Significance. If the generator and datasets are released as described, they constitute a useful community resource: they supply the missing non-ego supervision views for driving scenes at a scale unmatched by existing ego-exo autonomous-driving data, with dense ground-truth annotations and NeRFStudio-compatible poses. Concrete strengths include the open-source generator, the reproducible generation pipeline, the explicit train/test town split, and an honest Limitations section acknowledging that CARLA output is not photorealistic and that vehicle dynamics are limited. The benchmark half of the paper, however, is not yet a reliable source of evidence: as detailed below, Table 4 compares methods under incompatible protocols, the masking treatment of depth baselines is internally inconsistent, and no variance information is reported anywhere. Because benchmarks are a stated contribution, the manuscript needs revision before the few-image-to-3D benchmark claim is supported.

major comments (4)
  1. [§4, Table 4] The few-image-to-3D benchmark compares methods under incompatible protocols. The per-scene optimization methods (K-Planes, NeRFacto, SplatFacto) are evaluated on an unspecified subset of five scenes, while the feed-forward methods (PixelNeRF, SplatterImage, 6Img-to-3D) are trained on the training towns and evaluated on the Town02 test split. The resulting ranking cannot be attributed to method quality because the two groups are tested on different scenes under different data regimes. Please specify the evaluation scenes (which towns, how selected, which input views per scene) and either run all methods on an identical test set with identical inputs or restructure the table so that the two evaluations are clearly separated.
  2. [§4, Table 4 caption] The depth baseline masking is internally inconsistent. The masked variants ZoeDepth‡ and Metric3D‡ improve PSNR from 5.466 to 14.202 and from 6.314 to 13.699, respectively, yet the caption states that only the unmasked values are used for ranking. This places the depth baselines at a systematic disadvantage because the unmasked values penalize holes from point-cloud rasterization, while no masking protocol is applied to the other methods. Either apply the same masking protocol to all methods and rank on masked values, or remove the ‡ rows from the main comparison table and discuss them separately.
  3. [§4, Table 4, DRMSE column] The DRMSE column conflates two different quantities: for ZoeDepth and Metric3D it is monocular depth estimation error computed against ground-truth depth maps, while for 6Img-to-3D (and potentially pixelNeRF and SplatterImage) it is depth rendered from the reconstructed geometry. These measure different things, so the cross-method DRMSE comparison is not meaningful. The column should be recomputed under a unified protocol that defines the same ground-truth depth target and the same way of extracting depth from each method's output, or it should be removed.
  4. [§4, Tables 2–4] No variance information is reported for any benchmark result, and the number of evaluation scenes is undocumented for Table 2 while the five-scene subset used for K-Planes, NeRFacto, and SplatFacto in Table 4 is not described (which towns, how selected, whether the same scenes are used as in Table 2). For a benchmark paper, per-scene results or at least standard deviations and an explicit scene list are necessary for reproducibility and for judging whether the reported differences (e.g., K-Planes at 25.744 PSNR versus SplatFacto at 24.458 in Table 2) are significant.
minor comments (5)
  1. [§3.2, §7.1, Abstract] The counting convention for the ego cameras should be stated explicitly: the dataset totals (212k static images, 16.8M dynamic images) correspond to 6 ego views plus the exo views per scene or timestep, i.e., excluding the additional 110-degree rear camera that is described as part of the 'six plus one' setup.
  2. [§4, Table 4] The text lists 'SplatFacto-big' among the methods evaluated on five scenes, but no such row appears in Table 4; in addition, the SplatFacto row in Table 4 cites reference [112] while Table 2 cites [47] for the same method. Please align the text, the table, and the citations.
  3. [§11.2, Algorithm 2] The yaw formula in Algorithm 2 (line 7) is written as yaw = sign(x) * arccos(y / (x^2 + y^2)^0.5), which is not the standard arccos(x / sqrt(x^2 + y^2)); please rewrite the formula unambiguously and check that the implementation matches the intended spherical Fibonacci orientation.
  4. [§1] There is a duplicated phrase: 'due to the CARLA Simulator [25], our data generator and the data generator provide reliable ground truth annotations'; the second occurrence of 'the data generator' should be removed.
  5. [§4] The sentence 'results are shown Table in 3' should read 'results are shown in Table 3'.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity found: the dataset generator and benchmarks are self-contained; the authors' prior 6Img-to-3D appears only as one of several independent baselines.

full rationale

The paper's central deliverables are a CARLA-based data generator, two procedurally generated datasets, and benchmarks of existing methods. The generator is described operationally: camera setups are obtained by 'directly checking the camera poses or taken from the provided descriptions' of NuScenes/KITTI360/Waymo, while exo poses come from a spherical Fibonacci lattice, and all scene/sensor parameters are hand-specified. Nothing is fitted to benchmark outcomes, and no prediction is derived from its own inputs. The static and dynamic datasets are generated independently of the benchmark results, and the benchmarks evaluate held-out Town02 with externally developed methods (PixelNeRF, SplatterImage, K-Planes, NeRFacto, SplatFacto, ZoeDepth, Metric3D). The authors' own 6Img-to-3D is included as one of several baselines, but this is an evaluation use, not load-bearing evidence; the resource and benchmark claims would stand without it. The stated limitations in Section 5 (non-photorealistic CARLA imagery, limited vehicle dynamics) are honest scope caveats rather than circular reasoning. Potential protocol concerns around Table 4, such as per-scene optimizers being run on a five-scene subset while feed-forward methods are evaluated on Town02, and the decision to rank unmasked depth baselines despite also reporting masked variants, are methodological validity issues rather than instances of circularity, and they do not affect the generation pipeline.

Assumptions & free parameters 2 free parameters · 3 assumptions · 0 invented entities

The central claim is a dataset and benchmark contribution, not a derivation, so there are no fitted parameters in the traditional sense. The free parameters listed are hand-chosen dataset configuration choices that influence the difficulty and properties of the benchmarks. The axioms describe the trust placed in CARLA's sensor simulation and in the standard geometric construction for camera placement.

free parameters (2)
  • exocentric sphere radius = unstated
    The distance of exo cameras from the ego vehicle is described as a constant ('The cameras maintain the same absolute distance to the center vehicle') but no value is given; this geometry affects view overlap and reconstruction difficulty.
  • number of pedestrians and vehicles per scene = 20 pedestrians, 20 non-ego vehicles (static); 21 vehicles (dynamic)
    Hand-chosen scene complexity parameters; no ablation shows their effect on benchmark difficulty.
assumptions (3)
  • domain assumption CARLA simulator provides accurate ground-truth depth, optical flow, segmentation, and LiDAR consistent with the rendered RGB images.
    The dataset's supervision signals (Di, Ssemi, Sinsi, Pi) are generated by CARLA and treated as ground truth; any inconsistency between rendered RGB and these sensor outputs would propagate to models trained on the dataset. Invoked throughout Section 3.2.
  • standard math Spherical Fibonacci lattice evenly distributes the specified number of exocentric viewpoints on a half-sphere.
    The exo camera placement relies on the spherical Fibonacci lattice (Algorithm 1); this is a standard construction but the paper does not prove coverage or view diversity for the chosen N.
  • ad hoc to paper The chosen camera poses (e.g., FoV 90 degrees, half-sphere radius) are sufficient for reconstructing or predicting the scene from exocentric views.
    The number and placement of exo views (100 static, 10 dynamic) are hand-chosen design decisions with no ablations or theoretical justification; the benchmark results are the only evidence.

how reviews work

0 comments
Cite this review

Pith. "Pith review of SEED4D: A Synthetic Ego--Exo Dynamic 4D Data Generator, Driving Dataset and Benchmark." pith.science (2026). https://pith.science/paper/6TG6OYMN

@misc{pith2026241200730,
  author       = {Pith},
  title        = {Pith review of: SEED4D: A Synthetic Ego--Exo Dynamic 4D Data Generator, Driving Dataset and Benchmark},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/6TG6OYMN}},
  note         = {Machine review of arXiv:2412.00730}
}
read the original abstract

Models for egocentric 3D and 4D reconstruction, including few-shot interpolation and extrapolation settings, can benefit from having images from exocentric viewpoints as supervision signals. No existing dataset provides the necessary mixture of complex, dynamic, and multi-view data. To facilitate the development of 3D and 4D reconstruction methods in the autonomous driving context, we propose a Synthetic Ego--Exo Dynamic 4D (SEED4D) data generator and dataset. We present a customizable, easy-to-use data generator for spatio-temporal multi-view data creation. Our open-source data generator allows the creation of synthetic data for camera setups commonly used in the NuScenes, KITTI360, and Waymo datasets. Additionally, SEED4D encompasses two large-scale multi-view synthetic urban scene datasets. Our static (3D) dataset encompasses 212k inward- and outward-facing vehicle images from 2k scenes, while our dynamic (4D) dataset contains 16.8M images from 10k trajectories, each sampled at 100 points in time with egocentric images, exocentric images, and LiDAR data. The datasets and the data generator can be found at https://seed4d.github.io/.

Figures

Figures reproduced from arXiv: 2412.00730 by the authors.

Figure 1
Figure 1. The SEED4D dataset contains synthetic egocentric–exocentric dynamic 4D data and pose information (top). We benchmark [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Data Generator Capabilities. A. Showing example nuScenes images [13] and images generated by our data generator with similar intrinsic and extrinsic pose information. B. Removing dynamic vehicles from the scene. The parked vehicles remain. C. Generating egocentric views for all vehicles in a scene. multiple RGB cameras, LiDAR information, and 3D bound￾ing boxes [18]. An exception is the Cityscapes [19] dataset that … view at source ↗
Figure 3
Figure 3. Overview of sensor data contained within the SEED4D datasets. [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figures from the paper (11 more)
Figure 4
Figure 4. Figure 4: Egocentric and exocentric sensor configuration. The six [PITH_FULL_IMAGE:figures/full_fig_p005_4.png]
Figure 5
Figure 5. Figure 5: The left images show two scenes from the static dataset, the right images show two time points from the dynamic dataset. The [PITH_FULL_IMAGE:figures/full_fig_p006_5.png]
Figure 6
Figure 6. Figure 6: Qualitative Results for the single-shot few image scene reconstruction methods. [PITH_FULL_IMAGE:figures/full_fig_p016_6.png]
Figure 7
Figure 7. Figure 7: Qualitative Results for the multi-view per scene optimization methods. [PITH_FULL_IMAGE:figures/full_fig_p016_7.png]
Figure 8
Figure 8. Figure 8: Example style transfer results. ’C’ denotes Carla images, ’C-to-N’ indicates a transfer from the Carla domain to the NuScenes [PITH_FULL_IMAGE:figures/full_fig_p019_8.png]
Figure 9
Figure 9. Figure 9: Colored 3D point cloud generated from RGB images, depth maps, camera intrinsics, and extrinsics. [PITH_FULL_IMAGE:figures/full_fig_p020_9.png]
Figure 10
Figure 10. Figure 10: An overview of six egocentric cameras and their associated sensor measurements. [PITH_FULL_IMAGE:figures/full_fig_p020_10.png]
Figure 11
Figure 11. Figure 11: Samples from the Static Ego–Exo Dataset showing towns 1 to 4. The egocentric images show front left, front center, and front [PITH_FULL_IMAGE:figures/full_fig_p021_11.png]
Figure 12
Figure 12. Figure 12: Samples from the Static Ego–Exo Dataset showing towns 5 to 7 and 10HD. The egocentric images show front left, front center, [PITH_FULL_IMAGE:figures/full_fig_p022_12.png]
Figure 13
Figure 13. Figure 13: Samples from the Dynamic Ego–Exo Dataset showing towns 1 to 4 for timepoints 5, 20, and 65. The egocentric images show [PITH_FULL_IMAGE:figures/full_fig_p023_13.png]
Figure 14
Figure 14. Figure 14: Samples from the Dynamic Ego–Exo Dataset showing towns 1 to 4 for timepoints 5, 20, and 65. The egocentric images show [PITH_FULL_IMAGE:figures/full_fig_p024_14.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. EgoMe: A New Dataset and Challenge for Following Me via Egocentric View in Real World

    cs.CV 2025-01 conditional novelty 7.0 of 10

    EgoMe is a large-scale egocentric dataset of 7,902 paired observe-then-imitate videos with gaze, IMU, and multi-level annotations, plus six benchmarks for imitation learning.

  2. sshELF: Single-Shot Hierarchical Extrapolation of Latent Features for 3D Reconstruction from Sparse-Views

    cs.CV 2025-02 conditional novelty 6.0 of 10

    sshELF reconstructs full 360-degree outdoor scenes from six sparse views in 0.18 seconds by generating intermediate virtual views before decoding 3D Gaussian primitives.

Reference graph

Works this paper leans on

136 extracted references · 60 canonical work pages · cited by 2 Pith papers

  1. [1]

    Large-scale data for multiple-view stereopsis

    Henrik Aanæs, Rasmus Ramsbøl Jensen, George V ogiatzis, Engin Tola, and Anders Bjorholm Dahl. Large-scale data for multiple-view stereopsis. International Journal of Com- puter Vision, pages 1–16, 2016. 3

  2. [2]

    Youtube- 8m: A large-scale video classification benchmark

    Sami Abu-El-Haija, Nisarg Kothari, Joonseok Lee, Apos- tol (Paul) Natsev, George Toderici, Balakrishnan Varadara- jan, and Sudheendra Vijayanarasimhan. Youtube- 8m: A large-scale video classification benchmark. In arXiv:1609.08675, 2016. 3

  3. [3]

    Big bird: A large, fine-grained, bigram relatedness dataset for examining semantic composition

    Shima Asaadi, Saif Mohammad, and Svetlana Kiritchenko. Big bird: A large, fine-grained, bigram relatedness dataset for examining semantic composition. In Jill Burstein, Christy Doran, and Thamar Solorio, editors, NAACL-HLT (1), pages 505–516. Association for Computational Lin- guistics, 2019. 3

  4. [4]

    Frozen in time: A joint video and image encoder for end-to-end retrieval

    Max Bain, Arsha Nagrani, G ¨ul Varol, and Andrew Zisser- man. Frozen in time: A joint video and image encoder for end-to-end retrieval. In IEEE International Conference on Computer Vision, 2021. 3

  5. [5]

    Behley, M

    J. Behley, M. Garbade, A. Milioto, J. Quenzel, S. Behnke, J. Gall, and C. Stachniss. Towards 3D LiDAR-based se- mantic scene understanding of 3D point cloud sequences: The SemanticKITTI Dataset. The International Journal on Robotics Research, 40(8-9):959–967, 2021. 4

  6. [6]

    Behley, M

    J. Behley, M. Garbade, A. Milioto, J. Quenzel, S. Behnke, C. Stachniss, and J. Gall. SemanticKITTI: A Dataset for Semantic Scene Understanding of LiDAR Sequences. In Proc. of the IEEE/CVF International Conf. on Computer Vision (ICCV), 2019. 4

  7. [7]

    Zoedepth: Zero-shot trans- fer by combining relative and metric depth

    Shariq Farooq Bhat, Reiner Birkl, Diana Wofk, Peter Wonka, and Matthias M ¨uller. Zoedepth: Zero-shot trans- fer by combining relative and metric depth. CoRR, abs/2302.12288, 2023. 7, 8

  8. [8]

    Behave: Dataset and method for tracking human object in- teractions

    Bharat Lal Bhatnagar, Xianghui Xie, Ilya Petrov, Cristian Sminchisescu, Christian Theobalt, and Gerard Pons-Moll. Behave: Dataset and method for tracking human object in- teractions. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, jun 2022. 3

Show all 136 references
  1. [9]

    Federica Bogo, Javier Romero, Gerard Pons-Moll, and Michael J. Black. Dynamic FAUST: Registering human bodies in motion. In IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), July 2017. 3

  2. [10]

    Deepdeform: Learning non-rigid rgb-d reconstruction with semi-supervised data

    Alja ˇz Bo ˇziˇc, Michael Zollh ¨ofer, Christian Theobalt, and Matthias Nießner. Deepdeform: Learning non-rigid rgb-d reconstruction with semi-supervised data. 2020. 3

  3. [11]

    Immersive light field video with a layered mesh representation

    Michael Broxton, John Flynn, Ryan Overbeck, Daniel Er- ickson, Peter Hedman, Matthew DuVall, Jason Dourgarian, Jay Busch, Matt Whalen, and Paul Debevec. Immersive light field video with a layered mesh representation. ACM Transactions on Graphics (Proc. SIGGRAPH), 39(4):86:1– 8...

  4. [12]

    Virtual kitti 2

    Yohann Cabon, Naila Murray, and Martin Humenberger. Virtual kitti 2. arXiv preprint, arXiv:2001.10773, 2020. 4

  5. [13]

    Lang, Sourabh V ora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yu Pan, Giancarlo Baldan, and Oscar Beijbom

    Holger Caesar, Varun Bankiti, Alex H. Lang, Sourabh V ora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yu Pan, Giancarlo Baldan, and Oscar Beijbom. nuscenes: A multi- modal dataset for autonomous driving. In CVPR, 2020. 3, 4, 5, 6

  6. [14]

    Berk Calli, Aaron Walsman, Arjun Singh, Siddhartha Srini- vasa, Pieter Abbeel, and Aaron M. Dollar. Benchmarking in manipulation research: Using the yale-cmu-berkeley object and model set. IEEE Robotics amp; Automation Magazine, 22(3):36–52, Sept. 2015. 3

  7. [15]

    Hexplane: A fast representa- tion for dynamic scenes

    Ang Cao and Justin Johnson. Hexplane: A fast representa- tion for dynamic scenes. CVPR, 2023. 2

  8. [16]

    A short note about kinetics- 600, 2018

    Joao Carreira, Eric Noland, Andras Banki-Horvath, Chloe Hillier, and Andrew Zisserman. A short note about kinetics- 600, 2018. 3

  9. [17]

    Chang, Thomas A

    Angel X. Chang, Thomas A. Funkhouser, Leonidas J. Guibas, Pat Hanrahan, Qi-Xing Huang, Zimo Li, Silvio Savarese, Manolis Savva, Shuran Song, Hao Su, Jianxiong Xiao, Li Yi, and Fisher Yu. Shapenet: An information-rich 3d model repository. CoRR, abs/1512.03012, 2015. 2, 3

  10. [18]

    Argoverse: 3d tracking and forecasting with rich maps

    Ming-Fang Chang, John W Lambert, Patsorn Sangkloy, Jagjeet Singh, Slawomir Bak, Andrew Hartnett, De Wang, Peter Carr, Simon Lucey, Deva Ramanan, and James Hays. Argoverse: 3d tracking and forecasting with rich maps. In Conference on Computer Vision and Pattern Recognition (CVP...

  11. [19]

    The cityscapes dataset for semantic urban scene understanding

    Marius Cordts, Mohamed Omran, Sebastian Ramos, Timo Rehfeld, Markus Enzweiler, Rodrigo Benenson, Uwe Franke, Stefan Roth, and Bernt Schiele. The cityscapes dataset for semantic urban scene understanding. In Proc. of the IEEE Conference on Computer Vision and Pattern Recognitio...

  12. [20]

    Chang, Manolis Savva, Maciej Hal- ber, Thomas Funkhouser, and Matthias Nießner

    Angela Dai, Angel X. Chang, Manolis Savva, Maciej Hal- ber, Thomas Funkhouser, and Matthias Nießner. Scannet: Richly-annotated 3d reconstructions of indoor scenes. In Proc. Computer Vision and Pattern Recognition (CVPR), IEEE, 2017. 3

  13. [21]

    Scaling egocentric vision: The epic- kitchens dataset

    Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Sanja Fidler, Antonino Furnari, Evangelos Kazakos, Da- vide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, and Michael Wray. Scaling egocentric vision: The epic- kitchens dataset. In European Conference on Computer Vi...

  14. [22]

    Objaverse-xl: A universe of 10m+ 3d objects

    Matt Deitke, Ruoshi Liu, Matthew Wallingford, Huong Ngo, Oscar Michel, Aditya Kusupati, Alan Fan, Chris- tian Laforte, Vikram V oleti, Samir Yitzhak Gadre, Eli VanderBilt, Aniruddha Kembhavi, Carl V ondrick, Georgia Gkioxari, Kiana Ehsani, Ludwig Schmidt, and Ali Farhadi. Obja...

  15. [23]

    Objaverse: A universe of annotated 3d objects

    Matt Deitke, Dustin Schwenk, Jordi Salvador, Luca Weihs, Oscar Michel, Eli VanderBilt, Ludwig Schmidt, Kiana Ehsani, Aniruddha Kembhavi, and Ali Farhadi. Objaverse: A universe of annotated 3d objects. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Rec...

  16. [24]

    KITTI-CARLA: a KITTI-like dataset generated by CARLA Simulator

    Jean-Emmanuel Deschaud. KITTI-CARLA: a KITTI-like dataset generated by CARLA Simulator. arXiv e-prints ,

  17. [25]

    CARLA: An open urban driving simulator

    Alexey Dosovitskiy, German Ros, Felipe Codevilla, Anto- nio Lopez, and Vladlen Koltun. CARLA: An open urban driving simulator. In Proceedings of the 1st Annual Confer- ence on Robot Learning, pages 1–16, 2017. 2, 4

  18. [26]

    McHugh, and Vincent Vanhoucke

    Laura Downs, Anthony Francis, Nate Koenig, Brandon Kinman, Ryan Hickman, Krista Reymann, Thomas B. McHugh, and Vincent Vanhoucke. Google Scanned Ob- jects: A High-Quality Dataset of 3D Scanned Household Items. In 2022 International Conference on Robotics and Automation (ICRA) ...

  19. [27]

    Tenen- baum, and Jiajun Wu

    Yilun Du, Yinan Zhang, Hong-Xing Yu, Joshua B. Tenen- baum, and Jiajun Wu. Neural radiance flow for 4d view synthesis and video processing. In Proceedings of the IEEE/CVF International Conference on Computer Vision ,

  20. [28]

    Multi-level neural scene graphs for dynamic urban environments

    Tobias Fischer, Lorenzo Porzi, Samuel Rota Bul `o, Marc Pollefeys, and Peter Kontschieder. Multi-level neural scene graphs for dynamic urban environments. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024. 4

  21. [29]

    K- planes: Explicit radiance fields in space, time, and appear- ance

    Sara Fridovich-Keil, Giacomo Meanti, Frederik Rahbæk Warburg, Benjamin Recht, and Angjoo Kanazawa. K- planes: Explicit radiance fields in space, time, and appear- ance. In CVPR, 2023. 2, 7, 8, 3

  22. [30]

    Virtual worlds as proxy for multi-object tracking anal- ysis

    Adrien Gaidon, Qiao Wang, Yohann Cabon, and Eleonora Vig. Virtual worlds as proxy for multi-object tracking anal- ysis. In Proceedings of the IEEE conference on Computer Vision and Pattern Recognition, pages 4340–4349, 2016. 4

  23. [31]

    Vision meets robotics: The kitti dataset

    Andreas Geiger, Philip Lenz, Christoph Stiller, and Raquel Urtasun. Vision meets robotics: The kitti dataset. Interna- tional Journal of Robotics Research (IJRR), 2013. 3, 4

  24. [32]

    Are we ready for autonomous driving? the kitti vision benchmark suite

    Andreas Geiger, Philip Lenz, and Raquel Urtasun. Are we ready for autonomous driving? the kitti vision benchmark suite. In CVPR, pages 3354–3361. IEEE Computer Society,

  25. [33]

    Geiger, P

    A. Geiger, P. Lenz, and R. Urtasun. Are we ready for Au- tonomous Driving? The KITTI Vision Benchmark Suite. In Proc. of the IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), pages 3354–3361, 2012. 4

  26. [34]

    6img-to-3d: Few-image large- scale outdoor driving scene reconstruction

    Th ´eo Gieruc, Marius K ¨astingsch¨afer, Sebastian Bernhard, and Mathieu Salzmann. 6img-to-3d: Few-image large- scale outdoor driving scene reconstruction. arXiv preprint, arXiv:2404.12378, 2024. 1, 2, 7, 8, 3

  27. [35]

    Measurement of areas on a sphere us- ing fibonacci and latitude–longitude lattices

    Alvaro Gonzalez. Measurement of areas on a sphere us- ing fibonacci and latitude–longitude lattices. Mathematical geosciences, 42:49–64, 01 2010. 4

  28. [36]

    Kristen Grauman, Andrew Westbury, Lorenzo Torresani, Kris Kitani, Jitendra Malik, Triantafyllos Afouras, Ku- mar Ashutosh, Vijay Baiyya, Siddhant Bansal, Bikram Boote, Eugene Byrne, Zachary Chavis, Joya Chen, Feng Cheng, Fu-Jen Chu, Sean Crane, Avijit Dasgupta, Jing Dong, Mar ...

  29. [37]

    Photorealistic video generation with diffusion models

    Agrim Gupta, Lijun Yu, Kihyuk Sohn, Xiuye Gu, Meera Hahn, Fei-Fei Li, Irfan Essa, Lu Jiang, and Jos ´e Lezama. Photorealistic video generation with diffusion models. In Computer Vision – ECCV 2024: 18th European Confer- ence, Milan, Italy, September 29–October 4, 2024, Pro- ce...

  30. [38]

    Real- time deep dynamic characters

    Marc Habermann, Lingjie Liu, Weipeng Xu, Michael Zoll- hoefer, Gerard Pons-Moll, and Christian Theobalt. Real- time deep dynamic characters. ACM Transactions on Graphics, 40(4):1–16, July 2021. 3

  31. [39]

    Flexible diffusion model- ing of long videos

    William Harvey, Saeid Naderiparizi, Vaden Masrani, Chris- tian Weilbach, and Frank Wood. Flexible diffusion model- ing of long videos. In S. Koyejo, S. Mohamed, A. Agar- wal, D. Belgrave, K. Cho, and A. Oh, editors, Advances in Neural Information Processing Systems, volume 35,...

  32. [40]

    Synthehicle: Multi-vehicle multi-camera tracking in virtual cities

    Fabian Herzog, Junpeng Chen, Torben Teepe, Johannes Gilg, Stefan H ¨ormann, and Gerhard Rigoll. Synthehicle: Multi-vehicle multi-camera tracking in virtual cities. In Proceedings of the IEEE/CVF Winter Conference on Ap- plications of Computer Vision (WACV) Workshops , pages 1–...

  33. [41]

    Y .-T. Hu, J. Wang, R. A. Yeh, and A. G. Schwing. SAIL- VOS 3D: A Synthetic Dataset and Baselines for Object De- tection and 3D Mesh Reconstruction from Video Data. In CVPR, 2021. 3

  34. [42]

    Multimodal unsupervised image-to-image translation

    Xun Huang, Ming-Yu Liu, Serge Belongie, and Jan Kautz. Multimodal unsupervised image-to-image translation. In Computer Vision – ECCV 2018: 15th European Confer- ence, Munich, Germany, September 8–14, 2018, Proceed- ings, Part III , page 179–196, Berlin, Heidelberg, 2018. Sprin...

  35. [43]

    VBench: Comprehensive benchmark suite for video generative mod- els

    Ziqi Huang, Yinan He, Jiashuo Yu, Fan Zhang, Chenyang Si, Yuming Jiang, Yuanhan Zhang, Tianxing Wu, Qingyang 10 Jin, Nattapol Chanpaisit, Yaohui Wang, Xinyuan Chen, Limin Wang, Dahua Lin, Yu Qiao, and Ziwei Liu. VBench: Comprehensive benchmark suite for video generative mod- e...

  36. [44]

    Neo 360: Neural fields for sparse view synthesis of outdoor scenes

    Muhammad Zubair Irshad, Sergey Zakharov, Katherine Liu, Vitor Guizilini, Thomas Kollar, Adrien Gaidon, Zsolt Kira, and Rares Ambrus. Neo 360: Neural fields for sparse view synthesis of outdoor scenes. In Interntaional Confer- ence on Computer Vision (ICCV), 2023. 3, 4, 7

  37. [45]

    Large scale multi-view stereopsis eval- uation

    Rasmus Jensen, Anders Dahl, George V ogiatzis, Engil Tola, and Henrik Aanæs. Large scale multi-view stereopsis eval- uation. In 2014 IEEE Conference on Computer Vision and Pattern Recognition, pages 406–413, 2014. 3

  38. [46]

    Spherical fibonacci mapping

    Benjamin Keinert, Matthias Innmann, Michael S ¨anger, and Marc Stamminger. Spherical fibonacci mapping. ACM Trans. Graph., 34(6), nov 2015. 4

  39. [47]

    3d gaussian splatting for real-time radiance field rendering

    Bernhard Kerbl, Georgios Kopanas, Thomas Leimk ¨uhler, and George Drettakis. 3d gaussian splatting for real-time radiance field rendering. ACM Transactions on Graphics, 42(4), July 2023. 7, 3

  40. [48]

    Content disentanglement for semantically consistent synthetic-to- real domain adaptation

    Mert Keser, Artem Savkin, and Federico Tombari. Content disentanglement for semantically consistent synthetic-to- real domain adaptation. 2021 IEEE/RSJ International Con- ference on Intelligent Robots and Systems (IROS) , pages 3844–3849, 2021. 2, 8, 4

  41. [49]

    Point cloud forecasting as a proxy for 4d occupancy forecasting

    Tarasha Khurana, Peiyun Hu, David Held, and Deva Ra- manan. Point cloud forecasting as a proxy for 4d occupancy forecasting. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023. 1, 2

  42. [50]

    Kingma and Jimmy Ba

    Diederik P. Kingma and Jimmy Ba. Adam: A method for stochastic optimization. CoRR, abs/1412.6980, 2014. 3

  43. [51]

    Tanks and temples: benchmarking large-scale scene reconstruction

    Arno Knapitsch, Jaesik Park, Qian-Yi Zhou, and Vladlen Koltun. Tanks and temples: benchmarking large-scale scene reconstruction. ACM Trans. Graph., 36(4), jul 2017. 2, 3

  44. [52]

    EgoGen: An Egocentric Synthetic Data Generator

    Gen Li, Kaifeng Zhao, Siwei Zhang, Xiaozhong Lyu, Mi- hai Dusmanu, Yan Zhang, Marc Pollefeys, and Siyu Tang. EgoGen: An Egocentric Synthetic Data Generator. InIEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2024. 3

  45. [53]

    Qiao, Dahua Lin, Siqian Liu, Junchi Yan, Jianping Shi, and Ping Luo

    Hongyang Li, Chonghao Sima, Jifeng Dai, Wenhai Wang, Lewei Lu, Huijie Wang, Enze Xie, Zhiqi Li, Hanming Deng, Haonan Tian, Xizhou Zhu, Li Chen, Tianyu Li, Yulu Gao, Xiangwei Geng, Jianqiang Zeng, Yang Li, Jiazhi Yang, Xiaosong Jia, Bo Yu, Y . Qiao, Dahua Lin, Siqian Liu, Junch...

  46. [54]

    Xld: A cross-lane dataset for benchmarking novel driving view synthesis.arXiv preprint, arXiv:2406.18360, 2024

    Hao Li, Ming Yuan, Yan Zhang, Chenming Wu, Chen Zhao, Chunyu Song, Haocheng Feng, Errui Ding, Dingwen Zhang, and Jingdong Wang. Xld: A cross-lane dataset for benchmarking novel driving view synthesis.arXiv preprint, arXiv:2406.18360, 2024. 4

  47. [55]

    New- combe, and Zhaoyang Lv

    Tianye Li, Mira Slavcheva, Michael Zollh ¨ofer, Simon Green, Christoph Lassner, Changil Kim, Tanner Schmidt, Steven Lovegrove, Michael Goesele, Richard A. New- combe, and Zhaoyang Lv. Neural 3d video synthesis from multi-view video. In CVPR, pages 5511–5521. IEEE, 2022. 2, 3

  48. [56]

    Sora generates videos with stunning geometrical consistency, 2024

    Xuanyi Li, Daquan Zhou, Chenxu Zhang, Shaodong Wei, Qibin Hou, and Ming-Ming Cheng. Sora generates videos with stunning geometrical consistency, 2024. 2

  49. [57]

    4dcomplete: Non-rigid motion estimation beyond the observable surface

    Yang Li, Hikari Takehara, Takafumi Taketomi, and Bo Zheng nd Matthias Nießner. 4dcomplete: Non-rigid motion estimation beyond the observable surface. IEEE In- ternational Conference on Computer Vision (ICCV), 2021. 3

  50. [58]

    Neural scene flow fields for space-time view syn- thesis of dynamic scenes

    Zhengqi Li, Simon Niklaus, Noah Snavely, and Oliver Wang. Neural scene flow fields for space-time view syn- thesis of dynamic scenes. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2021. 2

  51. [59]

    KITTI-360: A novel dataset and benchmarks for urban scene understand- ing in 2d and 3d.Pattern Analysis and Machine Intelligence (PAMI), 2022

    Yiyi Liao, Jun Xie, and Andreas Geiger. KITTI-360: A novel dataset and benchmarks for urban scene understand- ing in 2d and 3d.Pattern Analysis and Machine Intelligence (PAMI), 2022. 3, 5

  52. [60]

    KITTI-360: A novel dataset and benchmarks for urban scene understand- ing in 2d and 3d.Pattern Analysis and Machine Intelligence (PAMI), 2022

    Yiyi Liao, Jun Xie, and Andreas Geiger. KITTI-360: A novel dataset and benchmarks for urban scene understand- ing in 2d and 3d.Pattern Analysis and Machine Intelligence (PAMI), 2022. 4

  53. [61]

    Dl3dv-10k: A large-scale scene dataset for deep learning-based 3d vision

    Lu Ling, Yichen Sheng, Zhi Tu, Wentian Zhao, Cheng Xin, Kun Wan, Lantao Yu, Qianyu Guo, Zixun Yu, Yawen Lu, et al. Dl3dv-10k: A large-scale scene dataset for deep learning-based 3d vision. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , p...

  54. [62]

    Jianbang Liu, Xinyu Mao, Yuqi Fang, Delong Zhu, and Max Q.-H. Meng. A survey on deep-learning approaches for vehicle trajectory prediction in autonomous driving. 2021 IEEE International Conference on Robotics and Biomimetics (ROBIO), pages 978–985, 2021. 1

  55. [63]

    Sora: A review on background, technology, limitations, and opportunities of large vision models

    Yixin Liu, Kai Zhang, Yuan Li, Zhiling Yan, Chujie Gao, Ruoxi Chen, Zhengqing Yuan, Yue Huang, Hanchi Sun, Jianfeng Gao, Lifang He, and Lichao Sun. Sora: A review on background, technology, limitations, and opportunities of large vision models. arXiv preprint, arXiv:2402.17177,

  56. [64]

    SGDR: stochastic gradi- ent descent with warm restarts

    Ilya Loshchilov and Frank Hutter. SGDR: stochastic gradi- ent descent with warm restarts. In 5th International Con- ference on Learning Representations, ICLR 2017, Toulon, France, April 24-26, 2017, Conference Track Proceedings. OpenReview.net, 2017. 3

  57. [65]

    Comport, Kefan Chen, and Srinath Sridhar

    Cheng-You Lu, Peisen Zhou, Angela Xing, Chandradeep Pokhariya, Arnab Dey, Ishaan Nikhil Shah, Rugved Ma- vidipalli, Dylan Hu, Andrew I. Comport, Kefan Chen, and Srinath Sridhar. Diva-360: The dynamic visual dataset for immersive neural fields. In Proceedings of the IEEE/CVF Co...

  58. [66]

    Aria everyday ac- tivities dataset, 2024

    Zhaoyang Lv, Nicholas Charron, Pierre Moulon, Alexan- der Gamino, Cheng Peng, Chris Sweeney, Edward Miller, Huixuan Tang, Jeff Meissner, Jing Dong, Kiran Somasun- daram, Luis Pesqueira, Mark Schwesinger, Omkar Parkhi, 11 Qiao Gu, Renzo De Nardi, Shangyi Cheng, Steve Saarinen, ...

  59. [67]

    Srinivasan, Rodrigo Ortiz-Cayon, Nima Khademi Kalantari, Ravi Ramamoorthi, Ren Ng, and Abhishek Kar

    Ben Mildenhall, Pratul P. Srinivasan, Rodrigo Ortiz-Cayon, Nima Khademi Kalantari, Ravi Ramamoorthi, Ren Ng, and Abhishek Kar. Local light field fusion: Practical view syn- thesis with prescriptive sampling guidelines. ACM Trans- actions on Graphics (TOG), 2019. 2, 3

  60. [68]

    Srinivasan, Rodrigo Ortiz-Cayon, Nima Khademi Kalantari, Ravi Ramamoorthi, Ren Ng, and Abhishek Kar

    Ben Mildenhall, Pratul P. Srinivasan, Rodrigo Ortiz-Cayon, Nima Khademi Kalantari, Ravi Ramamoorthi, Ren Ng, and Abhishek Kar. Local light field fusion: practical view syn- thesis with prescriptive sampling guidelines. 38(4), July

  61. [69]

    Srinivasan, Matthew Tancik, Jonathan T

    Ben Mildenhall, Pratul P. Srinivasan, Matthew Tancik, Jonathan T. Barron, Ravi Ramamoorthi, and Ren Ng. Syn- thetic nerf dataset. 2, 3

  62. [70]

    Al-Jarrah, Mehrdad Dianati, Paul Jennings, and Alexandros Mouzakitis

    Sajjad Mozaffari, Omar Y . Al-Jarrah, Mehrdad Dianati, Paul Jennings, and Alexandros Mouzakitis. Deep learning- based vehicle behavior prediction for autonomous driving applications: A review. IEEE Transactions on Intelligent Transportation Systems, 23(1):33–47, Jan. 2022. 1

  63. [71]

    Learning object-centric representations of multi-object scenes from multiple views

    Li Nanbo, Cian Eastwood, and Robert B Fisher. Learning object-centric representations of multi-object scenes from multiple views. In Advances in Neural Information Pro- cessing Systems, 2020. 3

  64. [72]

    CARLA-GeAR: a Dataset Generator for a Systematic Evaluation of Adver- sarial Robustness of Vision Models

    Federico Nesti, Giulio Rossolini, Gianluca D’Amico, Alessandro Biondi, and Giorgio Buttazzo. CARLA-GeAR: a Dataset Generator for a Systematic Evaluation of Adver- sarial Robustness of Vision Models. arXiv e-prints, page arXiv:2206.04365, June 2022. 4

  65. [73]

    A review on deep learning techniques for video prediction

    Sergiu Oprea, Pablo Martinez-Gonzalez, Alberto Garcia- Garcia, John Alejandro Castro-Vargas, Sergio Orts- Escolano, Jose Garcia-Rodriguez, and Antonis Argyros. A review on deep learning techniques for video prediction. IEEE Transactions on Pattern Analysis and Machine Intel- l...

  66. [74]

    Barron, Sofien Bouaziz, Dan B Goldman, Steven M

    Keunhong Park, Utkarsh Sinha, Jonathan T. Barron, Sofien Bouaziz, Dan B Goldman, Steven M. Seitz, and Ricardo Martin-Brualla. Nerfies: Deformable neural radiance fields. ICCV, 2021. 2, 3

  67. [75]

    Barron, Sofien Bouaziz, Dan B Goldman, Ricardo Martin- Brualla, and Steven M

    Keunhong Park, Utkarsh Sinha, Peter Hedman, Jonathan T. Barron, Sofien Bouaziz, Dan B Goldman, Ricardo Martin- Brualla, and Steven M. Seitz. Hypernerf: A higher- dimensional representation for topologically varying neural radiance fields. ACM Trans. Graph., 40(6), dec 2021. 2, 3

  68. [76]

    Carla2real: a tool for reducing the sim2real gap in carla simulator.arXiv preprint, arXiv:2410.18238, 2024

    Stefanos Pasios and Nikos Nikolaidis. Carla2real: a tool for reducing the sim2real gap in carla simulator.arXiv preprint, arXiv:2410.18238, 2024. 8

  69. [77]

    Peebles and Saining Xie

    William S. Peebles and Saining Xie. Scalable diffusion models with transformers. 2023 IEEE/CVF International Conference on Computer Vision (ICCV), pages 4172–4182,

  70. [78]

    Common objects in 3d: Large-scale learning and eval- uation of real-life 3d category reconstruction

    Jeremy Reizenstein, Roman Shapovalov, Philipp Hen- zler, Luca Sbordone, Patrick Labatut, and David Novotn ´y. Common objects in 3d: Large-scale learning and eval- uation of real-life 3d category reconstruction. In 2021 IEEE/CVF International Conference on Computer Vision, ICCV...

  71. [79]

    Richter, Hassan Abu Alhaija, and Vladlen Koltun

    Stephan R. Richter, Hassan Abu Alhaija, and Vladlen Koltun. Enhancing photorealism enhancement. IEEE Transactions on Pattern Analysis and Machine Intelligence, 45:1700–1715, 2021. 8

  72. [80]

    Richter, Zeeshan Hayder, and Vladlen Koltun

    Stephan R. Richter, Zeeshan Hayder, and Vladlen Koltun. Playing for benchmarks. In IEEE International Conference on Computer Vision, ICCV 2017, Venice, Italy, October 22- 29, 2017, pages 2232–2241, 2017. 4

  73. [81]

    Richter, Vibhav Vineet, Stefan Roth, and Vladlen Koltun

    Stephan R. Richter, Vibhav Vineet, Stefan Roth, and Vladlen Koltun. Playing for data: Ground truth from com- puter games. In Bastian Leibe, Jiri Matas, Nicu Sebe, and Max Welling, editors, European Conference on Computer Vision (ECCV) , volume 9906 of LNCS, pages 102–118. Spri...

  74. [82]

    German Ros, Laura Sellart, Joanna Materzynska, David Vazquez, and Antonio M. Lopez. The synthia dataset: A large collection of synthetic images for semantic segmenta- tion of urban scenes. In Proceedings of the IEEE Confer- ence on Computer Vision and Pattern Recognition (CVPR...

  75. [83]

    Forsyth, and Anand Bhattad

    Ayush Sarkar, Hanlin Mai, Amitabh Mahapatra, Svetlana Lazebnik, D.A. Forsyth, and Anand Bhattad. Shadows don’t lie and lines can’t bend! generative models don’t know projective geometry...for now. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogni...

  76. [84]

    Schuldt, I

    C. Schuldt, I. Laptev, and B. Caputo. Recognizing hu- man actions: a local svm approach. In Proceedings of the 17th International Conference on Pattern Recognition,

  77. [85]

    Harp: Autoregressive latent video pre- diction with high-fidelity image generator

    Younggyo Seo, Kimin Lee, Fangchen Liu, Stephen James, and Pieter Abbeel. Harp: Autoregressive latent video pre- diction with high-fidelity image generator. In ICIP, pages 3943–3947. IEEE, 2022. 1, 2

  78. [86]

    Airsim: High-fidelity visual and physical sim- ulation for autonomous vehicles

    Shital Shah, Debadeepta Dey, Chris Lovett, and Ashish Kapoor. Airsim: High-fidelity visual and physical sim- ulation for autonomous vehicles. In Field and Service Robotics, 2017. 4

  79. [87]

    Safety-enhanced autonomous driving using interpretable sensor fusion transformer

    Hao Shao, Letian Wang, RuoBing Chen, Hongsheng Li, and Yu Liu. Safety-enhanced autonomous driving using interpretable sensor fusion transformer. arXiv preprint , arXiv:2207.14024, 2022. 5

  80. [88]

    Common pets in 3d: Dynamic new-view synthesis of real-life deformable categories

    Samarth Sinha, Roman Shapovalov, Jeremy Reizenstein, Ignacio Rocco, Natalia Neverova, Andrea Vedaldi, and David Novotny. Common pets in 3d: Dynamic new-view synthesis of real-life deformable categories. CVPR, 2023. 3

  81. [89]

    Unsupervised learning of video representations using lstms

    Nitish Srivastava, Elman Mansimov, and Ruslan Salakhudi- nov. Unsupervised learning of video representations using lstms. In International conference on machine learning , pages 843–852. PMLR, 2015. 3

  82. [90]

    Julian Straub, Thomas Whelan, Lingni Ma, Yufan Chen, Erik Wijmans, Simon Green, Jakob J. Engel, Raul Mur- Artal, Carl Ren, Shobhit Verma, Anton Clarkson, Mingfei 12 Yan, Brian Budge, Yajie Yan, Xiaqing Pan, June Yon, Yuyang Zou, Kimberly Leon, Nigel Carter, Jesus Briales, Tyle...

  83. [91]

    Scala- bility in perception for autonomous driving: Waymo open dataset

    Pei Sun, Henrik Kretzschmar, Xerxes Dotiwalla, Aure- lien Chouard, Vijaysai Patnaik, Paul Tsui, James Guo, Yin Zhou, Yuning Chai, Benjamin Caine, Vijay Vasude- van, Wei Han, Jiquan Ngiam, Hang Zhao, Aleksei Timo- feev, Scott Ettinger, Maxim Krivokon, Amy Gao, Aditya Joshi, She...

  84. [92]

    Pix3d: Dataset and meth- ods for single-image 3d shape modeling

    Xingyuan Sun, Jiajun Wu, Xiuming Zhang, Zhoutong Zhang, Chengkai Zhang, Tianfan Xue, Joshua B Tenen- baum, and William T Freeman. Pix3d: Dataset and meth- ods for single-image 3d shape modeling. In IEEE Confer- ence on Computer Vision and Pattern Recognition (CVPR),

  85. [93]

    Richard Swinbank and R. Purser. Fibonacci grids: A novel approach to global modelling. Quarterly Journal of the Royal Meteorological Society, 132:1769 – 1793, 02 2006. 4

  86. [94]

    Splatter image: Ultra-fast single-view 3d recon- struction

    Stanislaw Szymanowicz, Christian Rupprecht, and Andrea Vedaldi. Splatter image: Ultra-fast single-view 3d recon- struction. In The IEEE/CVF Conference on Computer Vi- sion and Pattern Recognition (CVPR), 2024. 1, 7, 8

  87. [95]

    Mildenhall, Pratul Srinivasan, Jonathan T

    Matthew Tancik, Vincent Casser, Xinchen Yan, Sabeek Pradhan, Ben P. Mildenhall, Pratul Srinivasan, Jonathan T. Barron, and Henrik Kretzschmar. Block-nerf: Scalable large scene neural view synthesis. In 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR),...

  88. [96]

    Nerfstudio: A mod- ular framework for neural radiance field development

    Matthew Tancik, Ethan Weber, Evonne Ng, Ruilong Li, Brent Yi, Terrance Wang, Alexander Kristoffersen, Jake Austin, Kamyar Salahi, Abhik Ahuja, David Mcallister, Justin Kerr, and Angjoo Kanazawa. Nerfstudio: A mod- ular framework for neural radiance field development. In Specia...

  89. [97]

    Occ3d: a large-scale 3d occupancy prediction benchmark for autonomous driving

    Xiaoyu Tian, Tao Jiang, Longfei Yun, Yucheng Mao, Huitong Yang, Yue Wang, Yilun Wang, and Hang Zhao. Occ3d: a large-scale 3d occupancy prediction benchmark for autonomous driving. In Proceedings of the 37th In- ternational Conference on Neural Information Processing Systems, N...

  90. [98]

    Rtmv: A ray-traced multi-view synthetic dataset for novel view synthesis

    Jonathan Tremblay, Moustafa Meshry, Alex Evans, Jan Kautz, Alexander Keller, Sameh Khamis, Charles Loop, Nathan Morrical, Koki Nagano, Towaki Takikawa, and Stan Birchfield. Rtmv: A ray-traced multi-view synthetic dataset for novel view synthesis. IEEE/CVF European Conference o...

  91. [99]

    EPIC Fields: Marrying 3D Geometry and Video Understanding

    Vadim Tschernezki, Ahmad Darkhalil, Zhifan Zhu, David Fouhey, Iro Larina, Diane Larlus, Dima Damen, and An- drea Vedaldi. EPIC Fields: Marrying 3D Geometry and Video Understanding. In Proceedings of the Neural Infor- mation Processing Systems (NeurIPS), 2023. 3

  92. [100]

    Recovering ac- curate 3d human pose in the wild using imus and a mov- ing camera

    Timo von Marcard, Roberto Henschel, Michael Black, Bodo Rosenhahn, and Gerard Pons-Moll. Recovering ac- curate 3d human pose in the wild using imus and a mov- ing camera. In European Conference on Computer Vision (ECCV), sep 2018. 3

  93. [101]

    Embodiedscan: A holistic multi-modal 3d perception suite towards embodied ai

    Tai Wang, Xiaohan Mao, Chenming Zhu, Runsen Xu, Ruiyuan Lyu, Peisen Li, Xiao Chen, Wenwei Zhang, Kai Chen, Tianfan Xue, Xihui Liu, Cewu Lu, Dahua Lin, and Jiangmiao Pang. Embodiedscan: A holistic multi-modal 3d perception suite towards embodied ai. In IEEE Confer- ence on Comp...

  94. [102]

    Stylediffusion: Controllable disentangled style transfer via diffusion mod- els, 2023

    Zhizhong Wang, Lei Zhao, and Wei Xing. Stylediffusion: Controllable disentangled style transfer via diffusion mod- els, 2023. 2, 8

  95. [103]

    Ki- tani

    Xinshuo Weng, Junyu Nan, Kuan-Hui Lee, Rowan McAl- lister, Adrien Gaidon, Nicholas Rhinehart, and Kris M. Ki- tani. S2net: Stochastic sequential pointcloud forecasting. In Computer Vision – ECCV 2022: 17th European Con- ference, Tel Aviv, Israel, October 23–27, 2022, Proceed- ...

  96. [104]

    Inverting the pose forecasting pipeline with SPF2: sequential pointcloud forecasting for sequential pose forecasting

    Xinshuo Weng, Jianren Wang, Sergey Levine, Kris Kitani, and Nicholas Rhinehart. Inverting the pose forecasting pipeline with SPF2: sequential pointcloud forecasting for sequential pose forecasting. In Jens Kober, Fabio Ramos, and Claire J. Tomlin, editors, 4th Conference on Ro...

  97. [105]

    Synscapes: A pho- torealistic synthetic dataset for street scene parsing

    Magnus Wrenninge and Jonas Unger. Synscapes: A pho- torealistic synthetic dataset for street scene parsing. arXiv preprint, arXiv:1810.08705, 2018. 4

  98. [106]

    A survey on occupancy perception for au- tonomous driving: The information fusion perspective

    Huaiyuan Xu, Junliang Chen, Shiyu Meng, Yi Wang, and Lap-Pui Chau. A survey on occupancy perception for au- tonomous driving: The information fusion perspective. In- formation Fusion, 114:102671, 2025. 1

  99. [107]

    Temporally consistent transformers for video gen- eration

    Wilson Yan, Danijar Hafner, Stephen James, and Pieter Abbeel. Temporally consistent transformers for video gen- eration. In Proceedings of the 40th International Confer- ence on Machine Learning, ICML’23. JMLR.org, 2023. 1, 2, 3

  100. [108]

    Videogpt: Video generation using vq-vae and transformers, 2021

    Wilson Yan, Yunzhi Zhang, Pieter Abbeel, and Aravind Srinivas. Videogpt: Video generation using vq-vae and transformers, 2021. 1, 2

  101. [109]

    Street gaussians: Modeling dynamic urban scenes with gaussian splatting

    Yunzhi Yan, Haotong Lin, Chenxu Zhou, Weijie Wang, Haiyang Sun, Kun Zhan, Xianpeng Lang, Xiaowei Zhou, and Sida Peng. Street gaussians: Modeling dynamic urban scenes with gaussian splatting. In ECCV, 2024. 2

  102. [110]

    Generalized Predictive Model for Autonomous Driving

    Jiazhi Yang, Shenyuan Gao, Yihang Qiu, Li Chen, Tianyu Li, Bo Dai, Kashyap Chitta, Penghao Wu, Jia Zeng, Ping 13 Luo, Jun Zhang, Andreas Geiger, Yu Qiao, and Hongyang Li. Generalized Predictive Model for Autonomous Driving. In Proceedings of the IEEE/CVF Conference on Computer...

  103. [111]

    Blendedmvs: A large-scale dataset for generalized multi-view stereo net- works

    Yao Yao, Zixin Luo, Shiwei Li, Jingyang Zhang, Yufan Ren, Lei Zhou, Tian Fang, and Long Quan. Blendedmvs: A large-scale dataset for generalized multi-view stereo net- works. Computer Vision and Pattern Recognition (CVPR),

  104. [112]

    Mathematical sup- plement for the gsplat library

    Vickie Ye and Angjoo Kanazawa. Mathematical sup- plement for the gsplat library. arXiv preprint , arXiv:2312.02121, 2023. 8

  105. [113]

    Met- ric3d: Towards zero-shot metric 3d prediction from a sin- gle image

    Wei Yin, Chi Zhang, Hao Chen, Zhipeng Cai, Gang Yu, Kaixuan Wang, Xiaozhi Chen, and Chunhua Shen. Met- ric3d: Towards zero-shot metric 3d prediction from a sin- gle image. In Proceedings of the IEEE/CVF International Conference on Computer Vision , pages 9043–9053, 2023. 7, 8

  106. [114]

    pixelNeRF: Neural radiance fields from one or few images

    Alex Yu, Vickie Ye, Matthew Tancik, and Angjoo Kanazawa. pixelNeRF: Neural radiance fields from one or few images. In CVPR, 2021. 1, 7, 8

  107. [115]

    Bdd100k: A diverse driving dataset for hetero- geneous multitask learning

    Fisher Yu, Haofeng Chen, Xin Wang, Wenqi Xian, Yingy- ing Chen, Fangchen Liu, Vashisht Madhavan, and Trevor Darrell. Bdd100k: A diverse driving dataset for hetero- geneous multitask learning. In CVPR, pages 2633–2642. Computer Vision Foundation / IEEE, 2020. 3, 4

  108. [116]

    Yu, Fereshteh Forghani, Konstantinos G

    Jason J. Yu, Fereshteh Forghani, Konstantinos G. Derpanis, and Marcus A. Brubaker. Long-term photometric consistent novel view synthesis with diffusion models. In Proceed- ings of the International Conference on Computer Vision (ICCV), 2023. 3

  109. [117]

    Recent trends in 3d reconstruction of general non-rigid scenes

    Raza Yunus, Jan Eric Lenssen, Michael Niemeyer, Yiyi Liao, Christian Rupprecht, Christian Theobalt, Gerard Pons-Moll, Jia-Bin Huang, Vladislav Golyanik, and Eddy Ilg. Recent trends in 3d reconstruction of general non-rigid scenes. Computer Graphics Forum, 43, 2024. 2

  110. [118]

    A large-scale study of represen- tation learning with the visual task adaptation benchmark,

    Xiaohua Zhai, Joan Puigcerver, Alexander Kolesnikov, Pierre Ruyssen, Carlos Riquelme, Mario Lucic, Josip Djo- longa, Andre Susano Pinto, Maxim Neumann, Alexey Dosovitskiy, Lucas Beyer, Olivier Bachem, Michael Tschannen, Marcin Michalski, Olivier Bousquet, Sylvain Gelly, and Ne...

  111. [119]

    Learning unsupervised world models for autonomous driving via discrete diffu- sion

    Lunjun Zhang, Yuwen Xiong, Ze Yang, Sergio Casas, Rui Hu, and Raquel Urtasun. Learning unsupervised world models for autonomous driving via discrete diffu- sion. In International Conference on Learning Represen- tations (ICLR), 2024. 1, 2

  112. [120]

    The unreasonable effectiveness of deep features as a perceptual metric

    Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shecht- man, and Oliver Wang. The unreasonable effectiveness of deep features as a perceptual metric. In CVPR, 2018. 7

  113. [121]

    Structured local radiance fields for human avatar modeling

    Zerong Zheng, Han Huang, Tao Yu, Hongwen Zhang, Yandong Guo, and Yebin Liu. Structured local radiance fields for human avatar modeling. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2022. 3

  114. [122]

    Hugs: Holistic urban 3d scene understanding via gaussian splatting

    Hongyu Zhou, Jiahao Shao, Lu Xu, Dongfeng Bai, We- ichao Qiu, Bingbing Liu, Yue Wang, Andreas Geiger, and Yiyi Liao. Hugs: Holistic urban 3d scene understanding via gaussian splatting. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR),...

  115. [123]

    Garchingsim: An autonomous driving simulator with pho- torealistic scenes and minimalist workflow

    Liguo Zhou, Yinglei Song, Yichao Gao, Zhou Yu, Michael Sodamin, Hongshen Liu, Liang Ma, Lian Liu, Hao Liu, Yang Liu, Haichuan Li, Guang Chen, and Alois Knoll. Garchingsim: An autonomous driving simulator with pho- torealistic scenes and minimalist workflow. In 2023 IEEE 26th I...

  116. [124]

    Thingi10k: A dataset of 10, 000 3d-printing models

    Qingnan Zhou and Alec Jacobson. Thingi10k: A dataset of 10, 000 3d-printing models. ArXiv, abs/1605.04797, 2016. 2, 3

  117. [125]

    Drivinggaussian: Composite gaussian splatting for surrounding dynamic au- tonomous driving scenes

    Xiaoyu Zhou, Zhiwei Lin, Xiaojun Shan, Yongtao Wang, Deqing Sun, and Ming-Hsuan Yang. Drivinggaussian: Composite gaussian splatting for surrounding dynamic au- tonomous driving scenes. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages...

  118. [126]

    Jun-Yan Zhu, Taesung Park, Phillip Isola, and Alexei A. Efros. Unpaired image-to-image translation using cycle- consistent adversarial networks. In 2017 IEEE Inter- national Conference on Computer Vision (ICCV) , pages 2242–2251, 2017. 2 14 Appendix

  119. [128]

    Neo360 is listed because we use its re-implementation of PixelNeRF

    Licenses Below in Table 5 the licenses of the code and assets we make use of are listed. Neo360 is listed because we use its re-implementation of PixelNeRF. Table 5. Licenses. Item License CARLA code MIT CARLA assets CC-BY NeRFStudio Apache-2.0 PixelNeRF BSD-2-Clause SplatterI...

  120. [129]

    Dataset Details 7.1. Extended Dataset Description The static (3D) dataset encompasses 212k inward—and outward-facing vehicle images, while our dynamic (4D) dataset contains 16.8M images from 10k trajectories, each sampled at 100 points in time with egocentric and exocen- tric ...

  121. [130]

    Qualitative Results Multi-view Novel View Synthesis

    Benchmark Details 8.1. Qualitative Results Multi-view Novel View Synthesis. Figure 7 compares the qualitative results of Splatfacto, Nerfacto, and K-Planes. Our analysis shows that K-planes generalizes best both quantitatively and qualitatively, as demonstrated by the minimal ...

  122. [131]

    We welcome con- tributions to one of the proposed benchmarks or other sub- missions using the datasets

    Leaderboard We will actively maintain a leaderboard on the project page accompanying our SEED4D paper. We welcome con- tributions to one of the proposed benchmarks or other sub- missions using the datasets. Submissions can be made by contacting the first author

  123. [132]

    To find the latest hosting information of our datasets please see our project page here

    Hosting, licensing, and maintenance plan Hosting. To find the latest hosting information of our datasets please see our project page here. Licensing. Below in Table 8 the licenses of the code and assets we are publishing are listed. Table 8. Own Licenses. Item License Data gen...

  124. [133]

    Data Generation Details 11.1. Carla Towns The towns available within Carla vary in scenery, road structure, and size, with key characteristics highlighted be- low: Town 1: Town 1 is a compact environment divided by a river with several small bridges. The road network includes ...

  125. [134]

    The results in Figure 8 are obtained using a CylceGan- based framework as proposed in [48]

    Style transfer We experimented with existing style transfer methods to reduce the domain gap between Carla and NuScenes’ im- ages. The results in Figure 8 are obtained using a CylceGan- based framework as proposed in [48]. The checkpoint of our trained model will be made available

  126. [135]

    All full RGB images are paired with depth maps, op- tical flow, segmentation maps, and instance segmentation images

    Dataset Visualization. All full RGB images are paired with depth maps, op- tical flow, segmentation maps, and instance segmentation images. Since all values are ground truth, they can, for ex- ample, be used to generate a colored 3D point cloud using the camera’s extrinsics an...

  127. [136]

    Figure 11 and Figure 12 the static ego–exo dataset is visualized

    The sensory setup for an egocentric view is visualized in 10. Figure 11 and Figure 12 the static ego–exo dataset is visualized. Figure 13 and Figure 14 display the dynamic ego–exo dataset. 4 Figure 8. Example style transfer results. ’C’ denotes Carla images, ’C-to-N’ indicates...

  128. [2004]

    ICPR 2004., volume 3, pages 32–36 V ol.3, 2004. 2, 3

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.