REVIEW 4 major objections 5 minor 40 references
EgoTrack3D: A Modular Framework for Egocentric 3D Object Tracking
T0 review · 4 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read The paper claims that a modular pipeline lifting 2D segmentation masks into a global 3D frame, with a voxel-based duplicate-track merging heuristic, can persistently track all visible static and dynamic objects in egocentric video…
desk verdict A well-built egocentric 3D tracking system paper whose headline PCL gain depends on an unreported merging threshold; the evaluation needs hardening, but the contribution is real and deserves peer review. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central mechanism is the voxel-based track-merging heuristic (Eq. 6), which counts the occupied 3D voxels of two tracks' point clouds and merges them when the relative overlap exceeds a threshold τ; this is what prunes duplicate identities. Supporting it is the point-based motion score computed from tracked point trajectories: per-object averages of per-point 3D displacement and directional consistency decide whether an object is moving, so static observations can be accumulated while moving objects are replaced by current geometry.
What would settle it
Sweep the merging threshold τ in Eq. 6 across a range on the ADT sequences and compare merged track identities against ground-truth object IDs; if at any setting the number of overmerges exceeds the number of removed duplicate tracks, or if the PCL gain relative to 'no merge' disappears on a second dataset, the central claim that duplicate pruning is a reliable net win would fail.
Extended reading notes
Core claim
EgoTrack3D's central discovery is that the biggest obstacle to consistent egocentric 3D tracking is not association quality but duplicate tracks: when objects are partially occluded or re-observed, the same physical object is repeatedly re-initialized as a new track. The paper's voxel-based merging heuristic merges two tracks when the overlap of their occupied 3D voxel sets exceeds a threshold τ, and on ADT this heuristic lifts average PCL from 34.36 to 63.01, a relative improvement of 83.38%, while reducing duplicate tracks from 96 to 28 in one sequence. The framework combines a point-based motion scorer built from tracked point trajectories, an appearance embedding for visual similarity, and a Hungarian assignment over a cost combining 3D IoU, Chamfer distance, and feature similarity. It also introduces a sparse variant that replaces dense depth with learned 3D boxes and uses hand-object interaction cues to anchor manipulated objects, keeping their identities stable despite noisy geometry.
Load-bearing premise
The load-bearing premise is that the voxel-overlap merging heuristic rarely merges two different objects into one track, because the reported 83.38% gain comes almost entirely from that heuristic and the paper itself notes that occasional overmerges occur.
Editorial extensions
If this is right
- Persistent 3D tracking of all visible objects, static and dynamic, is achievable from egocentric RGB video without a prior scene model.
- Duplicate-track pruning is the dominant source of temporal consistency gain in egocentric 3D tracking, more than motion detection or appearance matching.
- When dense depth is unavailable, replacing mask lifting with learned 3D boxes and hand-object interaction cues keeps the framework viable, improving scene-level F1 from 29.88 to 40.79 over the strongest baseline on ADT.
- The pipeline transfers qualitatively to unconstrained egocentric video, preserving the identities of manipulated objects across large viewpoint changes.
- EgoTrack3D can serve as the perception backbone for constructing 3D scene graphs from egocentric observations.
Reading between the lines
- If the 83.38% gain is robust, it suggests that current egocentric trackers are bottlenecked by track initialization and termination discipline rather than by association features; a similar voxel-overlap merging step could be dropped into existing 2D-lifting pipelines to recover consistency.
- Because the merging threshold τ is not reported and the evaluation is on a single dataset, a practitioner should sweep τ on a validation split before deployment; the net effect may change with voxel size, camera motion, or object density.
- The sparse variant's interaction-guided association specifically protects hand-held objects, so it may fail for objects moved by tools, animals, or external agents; extending the dynamic cue beyond hands is a natural next step.
- The PCL metric counts a merged track as wrong if it fuses two distinct ground-truth objects, so the large net gain implies overmerging is rare on ADT; checking per-sequence overmerge counts would reveal where the trade-off point lies.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents EgoTrack3D, a modular framework for egocentric 3D object tracking from RGB video. The dense variant lifts 2D segmentation masks into 3D point clouds using depth, intrinsics, and poses; detects motion with CoTracker3; matches detections to tracks with a weighted IoU/Chamfer/feature cost; and merges duplicate tracks via voxel overlap. The sparse variant replaces dense depth with BoxerNet 3D box predictions and uses hand-object interaction plus SAM2 anchoring for dynamic objects. Experiments on ADT report an average PCL of 63.01 for the dense variant, an 11% relative improvement over Boxer (56.53), with the track-merging heuristic contributing an 83.38% relative gain over the no-merge ablation. Sparse-input experiments report improved AP/F1 over Boxer, and qualitative results on HD-EPIC illustrate tracking of manipulated objects.
Significance. If the evaluation is made robust, the contribution would be useful because it targets general persistent tracking of all visible objects with full point clouds, rather than only interacted or static objects. The paper's explicit ablation of the merge heuristic and its honest acknowledgment of overmerging are strengths, and the two instantiations of the modular pipeline make the framework concept credible. However, the headline numbers currently depend on unreported hyperparameters and weakened baseline adaptations, and no error bars, per-sequence breakdown, or code release are provided, so the practical significance is not yet established.
major comments (4)
- [§5.1, Eq. (6), Table 1] The central quantitative claim depends entirely on the voxel track-merging heuristic, but its threshold τ is never reported. Table 1 shows EgoTrack3D-Dense (merge) at 63.01 average PCL versus 34.36 without merging, and the text in §5.1 attributes an 83.38% relative improvement to the heuristic; without merging the method falls below Boxer (56.53). The same section acknowledges that overmerges occur when near-overlapping point clouds are incorrectly matched, and Appendix B repeats this failure mode. Because τ is a free parameter and the lookahead window W=10 in §5.1 is selected using ADT ground-truth trajectories, the reported gain may be a selected maximum rather than a predictive result. Please report τ, provide a sensitivity analysis over τ, and evaluate with the threshold chosen on a validation split or via cross-validation over the eight sequences; a per-sequence breakdown is also needed.
- [§4.3 and Appendix A] The comparison against the strongest baseline is weakened by the adaptation choices. IT3DEgo is evaluated without its 3D-guided Kalman filter because the authors' implementation was not released, and Boxer is evaluated only as a post-processed scene-level object set and is omitted from temporal tracking curves because it is static-only. Given that the abstract claims an 11% improvement over the strongest baseline, these choices make the relative gain difficult to interpret. Please either include an equivalent temporal-smoothing mechanism for IT3DEgo or clearly report how much of the gap is attributable to the missing Kalman filter, and provide the temporal PCL curve for Boxer if a per-frame output can be obtained from its online mode.
- [Table 1, Figure 2, Appendix B] The evaluation lacks variance reporting and the supporting duplicate-count statistics in Appendix B are internally inconsistent. Eight sequences are averaged without error bars or per-sequence tables, so it is unclear whether the 63.01 average PCL is driven by a few favorable sequences. In Appendix B, the Meal 132 scene is described as containing 267 predicted objects with 96 duplicates without merging, implying 171 distinct predictions, and 182 objects with 28 duplicates after merging, implying 154 distinct predictions; these numbers do not demonstrate that merging removes duplicates without sacrificing correct identities. Please clarify the counting convention (tracks vs. distinct object IDs) and report precision/recall on unique object identities before and after merging.
- [§5.2, Table 2] The sparse-input variant's robustness claims are not supported by the current numbers. The full method achieves AP 24.49 and F1 40.79, while removing 2D dynamic anchoring gives 24.11/39.28 and removing interaction-guided dynamics gives 23.13/40.43; these differences are small and no error bars or significance tests are provided. In addition, §4.2 states that 'all detections are valid and assumed to be above a fixed confidence threshold,' which is not a standard detection-evaluation protocol. Please report confidence-based AP curves or provide per-sequence variance, and temper the robustness claims accordingly.
minor comments (5)
- [§3.1.4, Eq. (4)] The matching cost weights λ_iou, λ_chamfer, and λ_feat are never reported; please include their values or state that they were chosen by a validation procedure. The normalization of the Chamfer and feature terms also needs a specification.
- [§3.1.1, Eq. (2)] The notation P(M_j) is not defined; please clarify that it denotes the set of CoTracker3 points falling inside mask M_j.
- [Table 1 and Figure 2] The table caption renders EgoSeg3D's α_c as 104, which should be 10^4, and Figure 2's y-axis is labeled PCL [%] while the axis values run from 0.0 to 1.0; please make the units consistent.
- [§4.2, PCL definition] The sentence 'the ratio between the number of correct predictions and the total number of unique ground-truth and predicted IDs' is ambiguous: please state whether the denominator is the size of the union of ground-truth and predicted IDs and how ID matching is established for objects that appear only in prediction or only in ground truth.
- [Appendix B] The qualitative figure captions refer to 'red boxes' as both duplicated tracks and false-positive duplicate predictions; please make the terminology consistent.
Assumptions & free parameters
free parameters (6)
- motion thresholds tau_m, tau_d =
0.1 (both)
- CoTracker3 grid size N, window Delta_t =
N=60, Delta_t=10
- voxel merge threshold tau =
not reported
- matching cost weights lambda_iou, lambda_chamfer, lambda_feat =
not reported
- appearance update weight alpha =
not reported
- sparse-variant thresholds (BoxerNet confidence, hand-overlap, SAM2 K) =
not reported
assumptions (4)
- domain assumption Ground-truth masks, depth, and poses on ADT are accurate and isolate tracking from perception errors.
- domain assumption Objects whose 3D point motion exceeds tau_m with consistent direction are genuinely dynamic.
- domain assumption Hand-object interaction is the main source of dynamic motion in egocentric scenes.
- ad hoc to paper The linear combination of IoU, Chamfer, and feature similarity is a valid matching cost.
Cite this review
Pith. "Pith review of EgoTrack3D: A Modular Framework for Egocentric 3D Object Tracking." pith.science (2026). https://pith.science/paper/R5M2XK6E
@misc{pith2026260808016,
author = {Pith},
title = {Pith review of: EgoTrack3D: A Modular Framework for Egocentric 3D Object Tracking},
year = {2026},
howpublished = {\url{https://pith.science/paper/R5M2XK6E}},
note = {Machine review of arXiv:2608.08016}
}
read the original abstract
Understanding 3D scenes from egocentric video is fundamental for robotics and autonomous navigation, yet rapid viewpoint changes and partial occlusions make building structured representations challenging. Existing 3D tracking and scene graph construction methods primarily address explicit interactions or assume static scenes, limiting their ability to capture complex dynamics. We introduce EgoTrack3D, a modular framework that reconstructs and maintains a dynamic 3D scene representation directly from egocentric RGB video. The framework lifts 2D segmentation masks into a global 3D coordinate frame, using a point-based motion scoring mechanism alongside a voxel-based merging heuristic to associate object tracks. EgoTrack3D maintains accurate representations over time, achieving an 11% improvement in percentage of correct locations (PCL) relative to the strongest baseline on the Aria Digital Twin (ADT) dataset, while addressing the more general setting of persistent 3D tracking for both static and dynamic objects. Furthermore, to demonstrate the system's robustness under degraded conditions that simulate real-world deployment constraints, we replace dense depth maps with sparse 3D bounding box estimation and integrate interaction-guided dynamic association, enabling EgoTrack3D to maintain accurate spatial representations despite noisy observations.
Figures
Figures from the paper (2 more)
Reference graph
Works this paper leans on
-
[1]
J. Johnson, R. Krishna, M. Stark, L.-J. Li, D. Shamma, M. Bernstein, and L. Fei-Fei. Image retrieval using scene graphs. InProceedings of the IEEE conference on computer vision and pattern recognition, pages 3668–3678, 2015
work page 2015
-
[2]
Krishna, Y
R. Krishna, Y . Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y . Kalantidis, L.-J. Li, D. A. Shamma, et al. Visual genome: Connecting language and vision using crowdsourced dense image annotations.International journal of computer vision, 123:32–73, 2017
2017
- [3]
-
[4]
Q. Gu, A. Kuwajerwala, S. Morin, K. M. Jatavallabhula, B. Sen, A. Agarwal, C. Rivera, W. Paul, K. Ellis, R. Chellappa, et al. Conceptgraphs: Open-vocabulary 3d scene graphs for perception and planning. In2024 IEEE International Conference on Robotics and Automation (ICRA), pages 5021–5028. IEEE, 2024
work page 2024
-
[5]
Long-Term Human Trajectory Prediction using 3D Dynamic Scene Graphs
N. Gorlo, L. Schmid, and L. Carlone. Long-term human trajectory prediction using 3d dynamic scene graphs.arXiv preprint arXiv:2405.00552, 2024
work page Pith review arXiv 2024
-
[6]
I. Armeni, Z.-Y . He, J. Gwak, A. R. Zamir, M. Fischer, J. Malik, and S. Savarese. 3d scene graph: A structure for unified semantics, 3d space, and camera. InProceedings of the IEEE/CVF international conference on computer vision, pages 5664–5673, 2019
work page 2019
- [7]
-
[8]
S.-C. Wu, J. Wald, K. Tateno, N. Navab, and F. Tombari. Scenegraphfusion: Incremental 3d scene graph prediction from rgb-d sequences. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 7515–7525, 2021
work page 2021
Show all 40 references
-
[9]
S. Koch, N. Vaskevicius, M. Colosi, P. Hermosilla, and T. Ropinski. Open3dsg: Open- vocabulary 3d scene graphs from point clouds with queryable objects and open-set relation- ships. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogni- tion, pages 14...
2024
-
[10]
Behrens, R
T. Behrens, R. Zurbr ¨ugg, M. Pollefeys, Z. Bauer, and H. Blum. Lost & found: Tracking changes from egocentric observations in 3d dynamic scene graphs.IEEE Robotics and Au- tomation Letters, 2025
2025
-
[11]
Y . Zhao, H. Ma, S. Kong, and C. Fowlkes. Instance tracking in 3d scenes from egocentric videos. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recog- nition, pages 21933–21944, 2024
2024
-
[12]
Plizzari, S
C. Plizzari, S. Goel, T. Perrett, J. Chalk, A. Kanazawa, and D. Damen. Spatial cognition from egocentric video: Out of sight, not out of mind. In2025 International Conference on 3D Vision (3DV), 2025
2025
-
[13]
Bhalgat, V
Y . Bhalgat, V . Tschernezki, I. Laina, J. F. Henriques, A. Vedaldi, and A. Zisserman. 3d- aware instance segmentation and tracking in egocentric videos. InProceedings of the Asian Conference on Computer Vision, pages 2562–2578, 2024
2024
-
[14]
DeTone, T
D. DeTone, T. Shen, F. Zhang, L. Ma, J. Straub, R. Newcombe, and J. Engel. Boxer: Robust lifting of open-world 2d bounding boxes to 3d.arXiv preprint arXiv:2604.05212, 2026
2026 arXiv
-
[15]
Karaev, I
N. Karaev, I. Makarov, J. Wang, N. Neverova, A. Vedaldi, and C. Rupprecht. Co- tracker3: Simpler and better point tracking by pseudo-labelling real videos.arXiv preprint arXiv:2410.11831, 2024
2024 arXiv
-
[16]
X. Pan, N. Charron, Y . Yang, S. Peters, T. Whelan, C. Kong, O. Parkhi, R. Newcombe, and Y . C. Ren. Aria digital twin: A new benchmark dataset for egocentric 3d machine perception. InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 20133– 20143, 2023
2023
-
[17]
Kirillov, E
A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y . Lo, et al. Segment anything. InProceedings of the IEEE/CVF international conference on computer vision, pages 4015–4026, 2023
2023
-
[18]
N. Ravi, V . Gabeur, Y .-T. Hu, R. Hu, C. Ryali, T. Ma, H. Khedr, R. R ¨adle, C. Rolland, L. Gustafson, et al. Sam 2: Segment anything in images and videos.arXiv preprint arXiv:2408.00714, 2024
2024 arXiv
-
[19]
Carion, L
N. Carion, L. Gustafson, Y .-T. Hu, S. Debnath, R. Hu, D. Suris, C. Ryali, K. V . Alwala, H. Khedr, A. Huang, J. Lei, T. Ma, B. Guo, A. Kalla, M. Marks, J. Greer, M. Wang, P. Sun, R. R¨adle, T. Afouras, E. Mavroudi, K. Xu, T.-H. Wu, Y . Zhou, L. Momeni, R. Hazra, S. Ding, S. V...
2025 arXiv
-
[20]
H. K. Cheng, S. W. Oh, B. Price, A. Schwing, and J.-Y . Lee. Tracking anything with decoupled video segmentation. InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 1316–1326, 2023
2023
-
[21]
J. Yang, M. Gao, Z. Li, S. Gao, F. Wang, and F. Zheng. Track anything: Segment anything meets videos.arXiv preprint arXiv:2304.11968, 2023
2023 arXiv
-
[22]
Raji ˇc, L
F. Raji ˇc, L. Ke, Y .-W. Tai, C.-K. Tang, M. Danelljan, and F. Yu. Segment anything meets point tracking. In2025 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), pages 9302–9311. IEEE, 2025
2025
-
[23]
S. F. Bhat, R. Birkl, D. Wofk, P. Wonka, and M. M ¨uller. Zoedepth: Zero-shot transfer by combining relative and metric depth.arXiv preprint arXiv:2302.12288, 2023
2023 arXiv
-
[24]
W. Yin, C. Zhang, H. Chen, Z. Cai, G. Yu, K. Wang, X. Chen, and C. Shen. Metric3d: Towards zero-shot metric 3d prediction from a single image. InProceedings of the IEEE/CVF international conference on computer vision, pages 9043–9053, 2023. 10
2023
-
[25]
Piccinelli, C
L. Piccinelli, C. Sakaridis, M. Segu, Y .-H. Yang, S. Li, W. Abbeloos, and L. Van Gool. Unik3d: Universal camera monocular 3d estimation. InProceedings of the Computer Vision and Pattern Recognition Conference, pages 1028–1039, 2025
2025
-
[26]
L. Yang, B. Kang, Z. Huang, X. Xu, J. Feng, and H. Zhao. Depth anything: Unleashing the power of large-scale unlabeled data. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10371–10381, 2024
2024
-
[27]
L. Yang, B. Kang, Z. Huang, Z. Zhao, X. Xu, J. Feng, and H. Zhao. Depth anything v2. Advances in Neural Information Processing Systems, 37:21875–21911, 2024
2024
-
[28]
B. Ke, A. Obukhov, S. Huang, N. Metzger, R. C. Daudt, and K. Schindler. Repurposing diffusion-based image generators for monocular depth estimation. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 9492–9502, 2024
2024
-
[29]
J. Shao, Y . Yang, H. Zhou, Y . Zhang, Y . Shen, V . Guizilini, Y . Wang, M. Poggi, and Y . Liao. Learning temporally consistent video depth from video diffusion priors. InProceedings of the Computer Vision and Pattern Recognition Conference, pages 22841–22852, 2025
2025
-
[30]
W. Hu, X. Gao, X. Li, S. Zhao, X. Cun, Y . Zhang, L. Quan, and Y . Shan. Depthcrafter: Generating consistent long depth sequences for open-world videos. InProceedings of the Computer Vision and Pattern Recognition Conference, pages 2005–2015, 2025
2005
-
[31]
S. Chen, H. Guo, S. Zhu, F. Zhang, Z. Huang, J. Feng, and B. Kang. Video depth anything: Consistent depth estimation for super-long videos. InProceedings of the Computer Vision and Pattern Recognition Conference, pages 22831–22840, 2025
2025
-
[32]
A. W. Harley, Z. Fang, and K. Fragkiadaki. Particle video revisited: Tracking through oc- clusions using point trajectories. InEuropean Conference on Computer Vision, pages 59–75. Springer, 2022
2022
-
[33]
Doersch, A
C. Doersch, A. Gupta, L. Markeeva, A. Recasens, L. Smaira, Y . Aytar, J. Carreira, A. Zis- serman, and Y . Yang. Tap-vid: A benchmark for tracking any point in a video.Advances in Neural Information Processing Systems, 35:13610–13626, 2022
2022
-
[34]
Karaev, I
N. Karaev, I. Rocco, B. Graham, N. Neverova, A. Vedaldi, and C. Rupprecht. Cotracker: It is better to track together. InEuropean conference on computer vision, pages 18–35. Springer, 2024
2024
-
[35]
Rosinol, A
A. Rosinol, A. Gupta, M. Abate, J. Shi, and L. Carlone. 3d dynamic scene graphs: Ac- tionable spatial perception with places, objects, and humans. arxiv 2020.arXiv preprint arXiv:2002.06289, 2020
2020 arXiv
-
[36]
S. Li, L. Ke, M. Danelljan, L. Piccinelli, M. Segu, L. Van Gool, and F. Yu. Matching anything by segmenting anything. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 18963–18973, 2024
2024
-
[37]
L. Qi, J. Kuen, W. Guo, T. Shen, J. Gu, J. Jia, Z. Lin, and M.-H. Yang. High-quality entity segmentation, 2023. URLhttps://arxiv.org/abs/2211.05776
2023 arXiv
-
[38]
Cheng, D
T. Cheng, D. Shan, A. S. Hassen, R. E. L. Higgins, and D. Fouhey. Towards a richer 2d un- derstanding of hands at scale. InThirty-seventh Conference on Neural Information Processing Systems, 2023
2023
-
[39]
Perrett, A
T. Perrett, A. Darkhalil, S. Sinha, O. Emara, S. Pollard, K. Parida, K. Liu, P. Gatti, S. Bansal, K. Flanagan, J. Chalk, Z. Zhu, R. Guerrier, F. Abdelazim, B. Zhu, D. Moltisanti, M. Wray, H. Doughty, and D. Damen. Hd-epic: A highly-detailed egocentric video dataset. InProceed-...
2025
-
[40]
Brazil, A
G. Brazil, A. Kumar, J. Straub, N. Ravi, J. Johnson, and G. Gkioxari. Omni3d: A large benchmark and model for 3d object detection in the wild. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 13154–13164, 2023. 12 A Baseline adaptation...
2023
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.