Pith. sign in

Virtual Worlds as Proxy for Multi-Object Tracking Analysis

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

Modern computer vision algorithms typically require expensive data acquisition and accurate manual labeling. In this work, we instead leverage the recent progress in computer graphics to generate fully labeled, dynamic, and photo-realistic proxy virtual worlds. We propose an efficient real-to-virtual world cloning method, and validate our approach by building and publicly releasing a new video dataset, called Virtual KITTI (see http://www.xrce.xerox.com/Research-Development/Computer-Vision/Proxy-Virtual-Worlds), automatically labeled with accurate ground truth for object detection, tracking, scene and instance segmentation, depth, and optical flow. We provide quantitative experimental evidence suggesting that (i) modern deep learning algorithms pre-trained on real data behave similarly in real and virtual worlds, and (ii) pre-training on virtual data improves performance. As the gap between real and virtual worlds is small, virtual worlds enable measuring the impact of various weather and imaging conditions on recognition performance, all other things being equal. We show these factors may affect drastically otherwise high-performing deep models for tracking.

fields

cs.CV 1

years

2026 1

verdicts

CONDITIONAL 1

representative citing papers

Vision as Unified Multimodal Generation

cs.CV · 2026-07-07 · conditional · novelty 7.0

A single unified multimodal model matches leading task-specialized vision systems across detection, segmentation, dense geometry, and multi-view 3D by casting all outputs as native text or image generation.

citing papers explorer

Showing 1 of 1 citing paper.

  • Vision as Unified Multimodal Generation cs.CV · 2026-07-07 · conditional · none · ref 45 · internal anchor

    A single unified multimodal model matches leading task-specialized vision systems across detection, segmentation, dense geometry, and multi-view 3D by casting all outputs as native text or image generation.