Pith. sign in

REVIEW 1 cited by

RGBTrack: Fast, Robust Depth-Free 6D Pose Estimation and Tracking

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2506.17119 v1 pith:IZMWLPQ4 submitted 2025-06-20 cs.CV cs.RO

classification cs.CVcs.RO
keywords rgbtrackposetrackingdepthobjectrobustdepth-freedynamic
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce a robust framework, RGBTrack, for real-time 6D pose estimation and tracking that operates solely on RGB data, thereby eliminating the need for depth input for such dynamic and precise object pose tracking tasks. Building on the FoundationPose architecture, we devise a novel binary search strategy combined with a render-and-compare mechanism to efficiently infer depth and generate robust pose hypotheses from true-scale CAD models. To maintain stable tracking in dynamic scenarios, including rapid movements and occlusions, RGBTrack integrates state-of-the-art 2D object tracking (XMem) with a Kalman filter and a state machine for proactive object pose recovery. In addition, RGBTrack's scale recovery module dynamically adapts CAD models of unknown scale using an initial depth estimate, enabling seamless integration with modern generative reconstruction techniques. Extensive evaluations on benchmark datasets demonstrate that RGBTrack's novel depth-free approach achieves competitive accuracy and real-time performance, making it a promising practical solution candidate for application areas including robotics, augmented reality, and computer vision. The source code for our implementation will be made publicly available at https://github.com/GreatenAnoymous/RGBTrack.git.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Robust 6-DoF Object Pose Tracking with Built-In Recovery under Occlusions and Rapid Object Motions

    cs.CV 2026-07 conditional novelty 5.0 of 10

    An ICG+-based RGB-D tracker with SuperPoint matching, a keyframe store, cycle-consistency failure detection, and TEASER++ recovery matches SOTA accuracy at 57.6 FPS and is most robust under occlusion and fast motion.

Reference graph

Works this paper leans on

39 extracted references · 26 canonical work pages · cited by 1 Pith paper

  1. [1]

    Foundationpose: Unified 6d pose estimation and tracking of novel objects,

    B. Wen, W. Yang, J. Kautz, and S. Birchfield, “Foundationpose: Unified 6d pose estimation and tracking of novel objects,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 17 868–17 879

  2. [2]

    One-2-3-45: Any single image to 3d mesh in 45 seconds without per-shape optimization,

    M. Liu, C. Xu, H. Jin, L. Chen, M. Varma T, Z. Xu, and H. Su, “One-2-3-45: Any single image to 3d mesh in 45 seconds without per-shape optimization,”Advances in Neural Information Processing Systems, vol. 36, 2024

  3. [3]

    Zero-1-to-3: Zero-shot one image to 3d object,

    R. Liu, R. Wu, B. Van Hoorick, P. Tokmakov, S. Zakharov, and C. V ondrick, “Zero-1-to-3: Zero-shot one image to 3d object,” in Proceedings of the IEEE/CVF international conference on computer vision, 2023, pp. 9298–9309

  4. [4]

    One-2- 3-45: Any single image to 3d mesh in 45 seconds without per-shape optimization,

    M. Liu, C. Xu, H. Jin, L. Chen, M. V . T, Z. Xu, and H. Su, “One-2- 3-45: Any single image to 3d mesh in 45 seconds without per-shape optimization,”arXiv preprint arXiv:2306.16928, 2023

  5. [5]

    Shap-e: Generating conditional 3d implicit functions,

    H. Jun and A. Nichol, “Shap-e: Generating conditional 3d implicit functions,”arXiv preprint arXiv:2305.02463, 2023

  6. [6]

    Xmem: Long-term video object segmentation with an atkinson-shiffrin memory model,

    H. K. Cheng and A. G. Schwing, “Xmem: Long-term video object segmentation with an atkinson-shiffrin memory model,” inEuropean Conference on Computer Vision. Springer, 2022, pp. 640–658

  7. [7]

    Pvnet: Pixel- wise voting network for 6dof pose estimation,

    S. Peng, Y . Liu, Q. Huang, X. Zhou, and H. Bao, “Pvnet: Pixel- wise voting network for 6dof pose estimation,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2019, pp. 4561–4570

  8. [8]

    Single-stage 6d object pose estimation,

    Y . Hu, P. Fua, W. Wang, and M. Salzmann, “Single-stage 6d object pose estimation,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020, pp. 2930–2939

Show all 39 references
  1. [9]

    Segmentation-driven 6d object pose estimation,

    Y . Hu, J. Hugonot, P. Fua, and M. Salzmann, “Segmentation-driven 6d object pose estimation,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2019, pp. 3385–3394

  2. [10]

    Cdpn: Coordinates-based disentangled pose network for real-time rgb-based 6-dof object pose estimation,

    Z. Li, G. Wang, and X. Ji, “Cdpn: Coordinates-based disentangled pose network for real-time rgb-based 6-dof object pose estimation,” in Proceedings of the IEEE/CVF international conference on computer vision, 2019, pp. 7678–7687

  3. [11]

    Bb8: A scalable, accurate, robust to partial occlusion method for predicting the 3d poses of challenging objects without using depth,

    M. Rad and V . Lepetit, “Bb8: A scalable, accurate, robust to partial occlusion method for predicting the 3d poses of challenging objects without using depth,” inProceedings of the IEEE international conference on computer vision, 2017, pp. 3828–3836

  4. [12]

    Real-time seamless single shot 6d object pose prediction,

    B. Tekin, S. N. Sinha, and P. Fua, “Real-time seamless single shot 6d object pose prediction,” inProceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 292–301

  5. [13]

    Dpod: 6d pose object detector and refiner,

    S. Zakharov, I. Shugurov, and S. Ilic, “Dpod: 6d pose object detector and refiner,” inProceedings of the IEEE/CVF international conference on computer vision, 2019, pp. 1941–1950

  6. [14]

    se (3)-tracknet: Data- driven 6d pose tracking by calibrating image residuals in synthetic domains,

    B. Wen, C. Mitash, B. Ren, and K. E. Bekris, “se (3)-tracknet: Data- driven 6d pose tracking by calibrating image residuals in synthetic domains,” in2020 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). IEEE, 2020, pp. 10 367–10 373

  7. [15]

    Bundletrack: 6d pose tracking for novel objects without instance or category-level 3d models,

    B. Wen and K. Bekris, “Bundletrack: 6d pose tracking for novel objects without instance or category-level 3d models,” in2021 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). IEEE, 2021, pp. 8067–8074

  8. [16]

    Onepose: One-shot object pose estimation without cad models,

    J. Sun, Z. Wang, S. Zhang, X. He, H. Zhao, G. Zhang, and X. Zhou, “Onepose: One-shot object pose estimation without cad models,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 6825–6834

  9. [17]

    Onepose++: Keypoint-free one-shot object pose estimation without cad models,

    X. He, J. Sun, Y . Wang, D. Huang, H. Bao, and X. Zhou, “Onepose++: Keypoint-free one-shot object pose estimation without cad models,” Advances in Neural Information Processing Systems, vol. 35, pp. 35 103– 35 115, 2022

  10. [18]

    Accurate non-iterative o(n) solution to the pnp problem,

    F. Moreno-Noguer, V . Lepetit, and P. Fua, “Accurate non-iterative o(n) solution to the pnp problem,” in2007 IEEE 11th International Conference on Computer Vision. Ieee, 2007, pp. 1–8

  11. [19]

    A method for registration of 3-d shapes,

    P. J. Besl and N. D. McKay, “A method for registration of 3-d shapes,” IEEE Transactions on pattern analysis and machine intelligence, vol. 14, no. 2, pp. 239–256, 1992

  12. [20]

    Model globally, match locally: Efficient and robust 3d object recognition,

    B. Drost, M. Ulrich, N. Navab, and S. Ilic, “Model globally, match locally: Efficient and robust 3d object recognition,” in2010 IEEE computer society conference on computer vision and pattern recognition. Ieee, 2010, pp. 998–1005

  13. [21]

    Cdpn: Coordinates-based disentangled pose network for real-time rgb-based 6-dof object pose estimation,

    Y . Li, G. Wang, X. Ji, Y . Xiang, and D. Fox, “Cdpn: Coordinates-based disentangled pose network for real-time rgb-based 6-dof object pose estimation,” inProceedings of the IEEE International Conference on Computer Vision (ICCV). IEEE, 2020, pp. 7678–7687

  14. [22]

    Zs6d: Zero-shot 6d object pose estimation using vision transformers,

    P. Ausserlechner, D. Haberger, S. Thalhammer, J.-B. Weibel, and M. Vincze, “Zs6d: Zero-shot 6d object pose estimation using vision transformers,” in2024 IEEE International Conference on Robotics and Automation (ICRA). IEEE, 2024, pp. 463–469

  15. [23]

    Megapose: 6d pose estimation of novel objects via render & compare,

    Y . Labbé, L. Manuelli, A. Mousavian, S. Tyree, S. Birchfield, J. Trem- blay, J. Carpentier, M. Aubry, D. Fox, and J. Sivic, “Megapose: 6d pose estimation of novel objects via render & compare,”arXiv preprint arXiv:2212.06870, 2022

  16. [24]

    Ove6d: Object viewpoint encoding for depth-based 6d object pose estimation,

    D. Cai, J. Heikkilä, and E. Rahtu, “Ove6d: Object viewpoint encoding for depth-based 6d object pose estimation,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 6803–6813

  17. [25]

    Sam-6d: Segment anything model meets zero-shot 6d object pose estimation,

    J. Lin, L. Liu, D. Lu, and K. Jia, “Sam-6d: Segment anything model meets zero-shot 6d object pose estimation,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 27 906–27 916

  18. [26]

    Templates for 3d object pose estimation revisited: Generalization to new objects and robustness to occlusions,

    V . N. Nguyen, Y . Hu, Y . Xiao, M. Salzmann, and V . Lepetit, “Templates for 3d object pose estimation revisited: Generalization to new objects and robustness to occlusions,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022, pp. 6771–6780

  19. [27]

    Gigapose: Fast and robust novel object pose estimation via one correspondence,

    V . N. Nguyen, T. Groueix, M. Salzmann, and V . Lepetit, “Gigapose: Fast and robust novel object pose estimation via one correspondence,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 9903–9913

  20. [28]

    Zephyr: Zero-shot pose hypothesis rating,

    B. Okorn, Q. Gu, M. Hebert, and D. Held, “Zephyr: Zero-shot pose hypothesis rating,” in2021 IEEE International Conference on Robotics and Automation (ICRA). IEEE, 2021, pp. 14 141–14 148

  21. [29]

    Foundpose: Unseen object pose estimation with foundation features,

    E. P. Örnek, Y . Labbé, B. Tekin, L. Ma, C. Keskin, C. Forster, and T. Hodan, “Foundpose: Unseen object pose estimation with foundation features,” inEuropean Conference on Computer Vision. Springer, 2024, pp. 163–182

  22. [30]

    Posecnn: A convolutional neural network for 6d object pose estimation in cluttered scenes,

    Y . Xiang, T. Schmidt, V . Narayanan, and D. Fox, “Posecnn: A convolutional neural network for 6d object pose estimation in cluttered scenes,”arXiv preprint arXiv:1711.00199, 2017

  23. [31]

    Clearpose: Large-scale transparent object dataset and benchmark,

    X. Chen, H. Zhang, Z. Yu, A. Opipari, and O. Chadwicke Jenkins, “Clearpose: Large-scale transparent object dataset and benchmark,” in European conference on computer vision. Springer, 2022, pp. 381–396

  24. [32]

    Bundlesdf: Neural 6-dof tracking and 3d reconstruction of unknown objects,

    B. Wen, J. Tremblay, V . Blukis, S. Tyree, T. Müller, A. Evans, D. Fox, J. Kautz, and S. Birchfield, “Bundlesdf: Neural 6-dof tracking and 3d reconstruction of unknown objects,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 606–617

  25. [33]

    Zoedepth: Zero-shot transfer by combining relative and metric depth,

    S. F. Bhat, R. Birkl, D. Wofk, P. Wonka, and M. Müller, “Zoedepth: Zero-shot transfer by combining relative and metric depth,”arXiv preprint arXiv:2302.12288, 2023

  26. [34]

    Metric3d: Towards zero-shot metric 3d prediction from a single image,

    W. Yin, C. Zhang, H. Chen, Z. Cai, G. Yu, K. Wang, X. Chen, and C. Shen, “Metric3d: Towards zero-shot metric 3d prediction from a single image,” inProceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 9043–9053

  27. [35]

    Sam 2: Segment anything in images and videos,

    N. Ravi, V . Gabeur, Y .-T. Hu, R. Hu, C. Ryali, T. Ma, H. Khedr, R. Rädle, C. Rolland, L. Gustafsonet al., “Sam 2: Segment anything in images and videos,”arXiv preprint arXiv:2408.00714, 2024

  28. [36]

    Bop challenge 2020 on 6d object localization,

    T. Hoda ˇn, M. Sundermeyer, B. Drost, Y . Labbé, E. Brachmann, F. Michel, C. Rother, and J. Matas, “Bop challenge 2020 on 6d object localization,” inComputer Vision–ECCV 2020 Workshops: Glasgow, UK, August 23–28, 2020, Proceedings, Part II 16. Springer, 2020, pp. 577–594

  29. [37]

    Dinov2: Learning robust visual features without supervision,

    M. Oquab, T. Darcet, T. Moutakanni, H. V o, M. Szafraniec, V . Khalidov, P. Fernandez, D. Haziza, F. Massa, A. El-Noubyet al., “Dinov2: Learning robust visual features without supervision,”arXiv preprint arXiv:2304.07193, 2023

  30. [38]

    Knn model-based approach in classification,

    G. Guo, H. Wang, D. Bell, Y . Bi, and K. Greer, “Knn model-based approach in classification,” inOn The Move to Meaningful Internet Systems 2003: CoopIS, DOA, and ODBASE: OTM Confederated International Conferences, CoopIS, DOA, and ODBASE 2003, Catania, Sicily, Italy, November ...

  31. [39]

    Yolov11: An overview of the key architectural enhancements,

    R. Khanam and M. Hussain, “Yolov11: An overview of the key architectural enhancements,”arXiv preprint arXiv:2410.17725, 2024

Pith tools