REVIEW 2 major objections 4 minor 51 references
Pre-training a Density-Aware Pose Transformer for Robust LiDAR-based 3D Human Pose Estimation
T0 review · 2 major / 4 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read DAPT, a density-aware pose transformer pre-trained on synthetic ray-cast LiDAR, achieves state-of-the-art single-frame LiDAR-only 3D human pose estimation, reducing MPJPE by 10.0 mm on Waymo and by 20.7 mm on SLOPER4D compared with prior…
desk verdict A solid architecture and pre-training pipeline with large gains on Waymo and HumanM3, but the SLOPER4D headline rests on a baseline comparison that isn't clearly apples-to-apples. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the Density-Aware Pose Transformer with its Multi-Density Exchange (MDE) modules. A UNet-style point transformer encodes the cloud into progressively sparser pooled features; at each decoding level, MDE exchanges information between point features and a fixed set of learnable joint anchors, so joints can draw on both dense near-body evidence and sparse global context. Joint positions are then decoded as the peak locations of three 1D heatmaps, one per coordinate axis, rather than regressed coordinates or point-segmentation votes. The other half of the machinery is the synthetic pre-training pipeline: SMPL meshes are rendered by ray casting onto a 64-line laser grid, randomly placed on a randomly oriented ground plane at distances of 4 to 20 meters, and portions of the laser grid are masked to simulate occlusions.
What would settle it
Train DAPT from scratch on each real dataset with the same fine-tuning budget and compare against the pre-train-then-fine-tune variant. If the from-scratch model matches or beats the pre-trained one on SLOPER4D's occluded urban sequences, or if increasing the laser-mask ratio (lowering rkeep below 0.6) improves real-world MPJPE, then the claimed transfer from synthetic pre-training is not doing the work the paper assigns to it.
Extended reading notes
Core claim
On its own terms, the paper establishes that the intrinsic density structure of a low-quality LiDAR point cloud carries enough information for 3D human pose estimation, provided the model is built to exploit it and pre-trained on realistic synthetic LiDAR. The proposed Density-Aware Pose Transformer (DAPT) introduces learnable joint anchors that are progressively updated by Multi-Density Exchange (MDE) modules while point features are decoded across pooling levels, and it represents each joint's location by the peaks of three 1D heatmaps over the XYZ axes. Before fine-tuning on real data, the model is pre-trained on ray-cast renders of SMPL meshes sampled from LiDARCap, with randomized positions, orientations, ground planes, and patchwise laser-grid masking that simulates occlusion. The paper reports that this recipe reduces average MPJPE by 10.0 mm over LPFormer on Waymo and by 20.7 mm over PRN on SLOPER4D, with consistent gains on LiDARHuman26M and HumanM3 and notably smaller variance on wrists and ankles.
Load-bearing premise
The key assumption is that point clouds synthesized by ray casting SMPL meshes on randomized ground scenes with laser-grid masks are realistic enough that a model pre-trained on them learns body priors that transfer to real LiDAR data; if the simulation-to-real gap is large, the reported pre-training gains would shrink or reverse.
Editorial extensions
If this is right
- Single-frame LiDAR-only 3D human pose estimation can reach state-of-the-art accuracy without temporal smoothing, multi-modal fusion, or SMPL optimization, simplifying deployment in autonomous driving and surveillance.
- Pre-training on synthetic ray-cast LiDAR with scene randomization and laser-level masking transfers to real outdoor datasets, improving MPJPE on Waymo by 7.5 mm compared to training from scratch.
- Joint anchors with multi-density exchange reduce reliance on correct point-to-body-part segmentation, so left-right ambiguity and noisy background points cause fewer joint errors.
- 1D heatmap decoding yields more stable predictions on end joints such as ankles and wrists, where point coverage is thinnest.
- The same architecture and synthesis pipeline can be re-trained on new LiDAR datasets with only heatmap supervision, since fine-tuning needs only joint annotations.
Reading between the lines
- The authors leave implicit that laser-level masking acts as a targeted occlusion-robustness regularizer; a direct test would be to vary mask ratio when fine-tuning on heavily occluded subsets and measure error on occluded frames.
- The joint-anchor plus multi-density exchange design is not human-specific and could be applied to other sparse point-cloud keypoint tasks, such as animal pose estimation or articulated object tracking.
- Because fine-tuning uses only heatmap loss and no segmentation labels, the pipeline could be adapted to new LiDAR sensors such as 32-line or 16-line units by re-running the ray-casting synthesizer at matching laser resolutions.
- The reported stability results under point jittering and noise clusters suggest the pre-trained model may be robust enough for downstream behavior understanding, but the paper does not evaluate that downstream task.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes DAPT, a density-aware pose transformer for single-frame LiDAR-based 3D human pose estimation, together with a synthetic pre-training pipeline. The model uses learnable joint anchors, multi-density exchange (MDE) modules, and 1D heatmap decoding, and is pre-trained on ray-cast SMPL meshes with randomized scene sampling, ground-plane modeling, and laser-level masking. The authors report state-of-the-art MPJPE on Waymo (51.59 mm, 10.0 mm lower than LPFormer), SLOPER4D (28.01 mm, 20.7 mm lower than PRN), HumanM3 (59.76 mm), and competitive results on LiDARHuman26M, with ablations and stability evaluations.
Significance. If the results hold, this is a practical contribution: the method is optimization-free and uses only single-frame LiDAR, avoiding temporal or multi-modal dependencies. The joint-anchor representation with 1D heatmaps is a reasonable design, and the synthetic pre-training pipeline is a useful data-augmentation idea. Strengths include extensive evaluation on four datasets, ablation studies isolating each component, stability analysis under point jittering and noise clusters, and public code. However, the comparison protocol for the headline SLOPER4D gain is not yet established, and the lack of a scratch baseline on SLOPER4D prevents attribution of the gain to pre-training versus architecture.
major comments (2)
- [§4.4, Table 1] The PRN baseline is modified by replacing its point cloud backbone with PTv3, and the paper does not state whether the LPFormer and PRN entries were retrained on the exact SLOPER4D split of Zhang et al. (2024) with the same evaluation code. This makes the advertised 20.7 mm MPJPE improvement over PRN not an apples-to-apples comparison. Please report a scratch DAPT (no pre-training) result on SLOPER4D and either rerun published baselines under the same protocol or clearly state which numbers are quoted.
- [§4.7, Table 4a] The pre-training ablation is reported only on Waymo (scratch 59.2 vs. full 51.7). Since the SLOPER4D result (28.01 mm) is the headline, it is currently unsupported that pre-training contributes to that gain; the gap versus PRN could be due to the DAPT architecture alone. Please provide a pre-training ablation on SLOPER4D.
minor comments (4)
- [§4.6, §5, §3.1] There are several typos: 'Statbility' in the Section 4.6 heading, 'Conclution' in Section 5, and 'SMLP' in Section 3.1 (should be SMPL). Also, the references list Weng et al. 2023a and 2023b for the same paper; please merge them.
- [Table 3] The checkmark layout in the ablation table is ambiguous; for example, row 4 shows two checkmarks (Anchor and Heatmaps) but the text says 'coordinate-based decoding is replaced with heatmap-based decoding,' which could be interpreted as Anchor+MDE+Heatmaps. Please clarify which components are enabled in each row, either by using explicit labels or a clearer table format.
- [§4.1, §3.1] Please specify the number of synthetic pre-training samples and the computational cost of pre-training, and clarify whether the voxelization grid size of 0.01 is in meters. Also, state the azimuth resolution of the simulated 64-line LiDAR with 2650 angles and whether it mimics a specific real sensor.
- [§4.2] The Waymo dataset description says it 'contains 10K human instances'; please clarify whether this is the number of annotated frames, instances, or subjects, and whether the evaluation subset is the same as in LPFormer. This is important for understanding the comparison basis.
Circularity Check
No circularity: the architecture and pre-training results are empirical, benchmarked against held-out datasets, and do not reduce to their own inputs.
full rationale
The paper's central claims are empirical measurements supported by ablations and held-out benchmarks rather than derivations from fitted parameters. The proposed components (joint anchors, MDE modules, 1D heatmap decoding) are architectural choices whose contributions are evaluated incrementally on the Waymo dataset (Table 3), and none of them is defined in terms of the final MPJPE or in terms of the baseline methods. The pre-training pipeline uses synthetic ray-cast SMPL meshes with random scene sampling, ground planes, and laser-level masking, and its benefit is tested both within-domain and via transfer to Waymo, HumanM3, and SLOPER4D; Table 4a shows that pre-training alone improves Waymo from 59.2 to 52.9 mm, an externally falsifiable effect that does not depend on the LiDARHuman26M evaluation set. Although the synthetic SMPL database originates from the LiDARCap/LiDARHuman26M source, the evaluation on Waymo and other datasets shows the pre-training gains are not forced by same-dataset leakage. The paper does not import any load-bearing result from the authors' own prior work; the self-citation to SHaRPose appears only in related-work context. The noted concern about replacing PRN's backbone with PTv3 and using a non-official SLOPER4D split is a comparison-protocol and reproducibility issue, not a circularity, because the baselines remain externally defined and DAPT is not constructed from them. No equation, loss, or reported number reduces to its own input by construction.
Assumptions & free parameters
free parameters (3)
- rkeep (laser-level masking retention ratio) =
0.6
- lambda_reg (pre-training joint regression loss weight) =
0.5
- lambda_seg (pre-training segmentation loss weight) =
1.0
assumptions (4)
- domain assumption SMPL model and LiDARCap SMPL database provide realistic human body geometry and pose distribution for pre-training.
- domain assumption Ray casting with a 64-line, 2650-angle LiDAR model accurately simulates the noise, sparsity, and occlusion characteristics of real LiDAR sensors.
- domain assumption 1D heatmaps on the XYZ axes with KL divergence are an adequate supervision signal for 3D joint localization.
- domain assumption Point Transformer V3 (PTv3) provides suitable point cloud features for the task when used as the backbone.
Cite this review
Pith. "Pith review of Pre-training a Density-Aware Pose Transformer for Robust LiDAR-based 3D Human Pose Estimation." pith.science (2026). https://pith.science/paper/EG2H4BJE
@misc{pith2026241213454,
author = {Pith},
title = {Pith review of: Pre-training a Density-Aware Pose Transformer for Robust LiDAR-based 3D Human Pose Estimation},
year = {2026},
howpublished = {\url{https://pith.science/paper/EG2H4BJE}},
note = {Machine review of arXiv:2412.13454}
}
abstract
With the rapid development of autonomous driving, LiDAR-based 3D Human Pose Estimation (3D HPE) is becoming a research focus. However, due to the noise and sparsity of LiDAR-captured point clouds, robust human pose estimation remains challenging. Most of the existing methods use temporal information, multi-modal fusion, or SMPL optimization to correct biased results. In this work, we try to obtain sufficient information for 3D HPE only by modeling the intrinsic properties of low-quality point clouds. Hence, a simple yet powerful method is proposed, which provides insights both on modeling and augmentation of point clouds. Specifically, we first propose a concise and effective density-aware pose transformer (DAPT) to get stable keypoint representations. By using a set of joint anchors and a carefully designed exchange module, valid information is extracted from point clouds with different densities. Then 1D heatmaps are utilized to represent the precise locations of the keypoints. Secondly, a comprehensive LiDAR human synthesis and augmentation method is proposed to pre-train the model, enabling it to acquire a better human body prior. We increase the diversity of point clouds by randomly sampling human positions and orientations and by simulating occlusions through the addition of laser-level masks. Extensive experiments have been conducted on multiple datasets, including IMU-annotated LidarHuman26M, SLOPER4D, and manually annotated Waymo Open Dataset v2.0 (Waymo), HumanM3. Our method demonstrates SOTA performance in all scenarios. In particular, compared with LPFormer on Waymo, we reduce the average MPJPE by $10.0mm$. Compared with PRN on SLOPER4D, we notably reduce the average MPJPE by $20.7mm$.
Figures
Figures from the paper (5 more)
Reference graph
Works this paper leans on
-
[1]
An, X.; Zhao, L.; Gong, C.; Wang, N.; Wang, D.; and Yang, J. 2024. SHaRPose : Sparse High-Resolution Representation for Human Pose Estimation . Proceedings of the AAAI Conference on Artificial Intelligence, 38(2): 691--699
2024
-
[2]
Cong, P.; Xu, Y.; Ren, Y.; Zhang, J.; Xu, L.; Wang, J.; Yu, J.; and Ma, Y. 2023. Weakly Supervised 3D Multi-Person Pose Estimation for Large-Scale Scenes Based on Monocular Camera and Single LiDAR . Proceedings of the AAAI Conference on Artificial Intelligence, 37(1): 461--469
work page 2023
-
[3]
Cong, P.; Zhu, X.; Qiao, F.; Ren, Y.; Peng, X.; Hou, Y.; Xu, L.; Yang, R.; Manocha, D.; and Ma, Y. 2022. STCrowd : A Multimodal Dataset for Pedestrian Perception in Crowded Scenes . In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition , 19608--19617
work page 2022
-
[4]
Dai, Y.; Lin, Y.; Lin, X.; Wen, C.; Xu, L.; Yi, H.; Shen, S.; Ma, Y.; and Wang, C. 2023. SLOPER4D : A Scene-Aware Dataset for Global 4D Human Pose Estimation in Urban Environments . In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition , 682--692
work page 2023
-
[5]
Devlin, J.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2019. BERT : Pre-training of Deep Bidirectional Transformers for Language Understanding . arXiv:1810.04805
arXiv 2019
-
[6]
Fan, B.; Wang, S.; Guo, W.; Zheng, W.; Feng, J.; and Zhou, J. 2023 a . Human- M3 : A Multi-view Multi-modal Dataset for 3D Human Pose Estimation in Outdoor Scenes . arXiv:2308.00628
arXiv 2023
- [7]
-
[8]
Fan, L.; Zhu, Y.; Zhu, J.; Liu, Z.; Zeng, O.; Gupta, A.; Creus-Costa , J.; Savarese, S.; and Fei-Fei , L. 2018. SURREAL : Open-Source Reinforcement Learning Framework and Robot Manipulation Benchmark . In Proceedings of The 2nd Conference on Robot Learning , 767--782
work page 2018
Show all 51 references
-
[9]
u rst, M.; Gupta, S. T. P.; Schuster, R.; Wasenm \
F \"u rst, M.; Gupta, S. T. P.; Schuster, R.; Wasenm \"u ller, O.; and Stricker, D. 2021. HPERL : 3D Human Pose Estimation from RGB and LiDAR . In Proceedings of International Conference on Pattern Recognition , 7321--7327
2021
-
[10]
Gong, K.; Zhang, J.; and Feng, J. 2021. PoseAug : A Differentiable Pose Augmentation Framework for 3D Human Pose Estimation . In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition , 8575--8584
2021
-
[11]
He, K.; Chen, X.; Xie, S.; Li, Y.; Doll \'a r, P.; and Girshick, R. 2022. Masked Autoencoders Are Scalable Vision Learners . In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition , 15979--15988
2022
-
[12]
He, K.; Fan, H.; Wu, Y.; Xie, S.; and Girshick, R. 2020. Momentum Contrast for Unsupervised Visual Representation Learning . In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition , 9729--9738
2020
-
[13]
Hong, S.; and Kim, Y. 2018. Dynamic Pose Estimation Using Multiple RGB-D Cameras . Sensors, 18(11): 3865
2018
-
[14]
Hu, X.; Zhong, B.; Liang, Q.; Zhang, S.; Li, N.; and Li, X. 2024. Towards Modalities Correlation for RGB-T Tracking. IEEE Transactions on Circuits and Systems for Video Technology
2024
-
[15]
Ionescu, C.; Papava, D.; Olaru, V.; and Sminchisescu, C. 2014. Human3. 6M : Large Scale Datasets and Predictive Methods for 3D Human Sensing in Natural Environments . IEEE Transactions on Pattern Analysis and Machine Intelligence, 36(7): 1325--1339
2014
-
[16]
Iskakov, K.; Burkov, E.; Lempitsky, V.; and Malkov, Y. 2019. Learnable Triangulation of Human Pose . In Proceedings of the IEEE / CVF International Conference on Computer Vision , 7718--7727
2019
-
[17]
Kang, Y.; Liu, Y.; Yao, A.; Wang, S.; and Wu, E. 2023. 3D Human Pose Lifting with Grid Convolution . Proceedings of the AAAI Conference on Artificial Intelligence, 37(1): 1105--1113
2023
-
[18]
Li, J.; Zhang, J.; Wang, Z.; Shen, S.; Wen, C.; Ma, Y.; Xu, L.; Yu, J.; and Wang, C. 2022 a . LiDARCap : Long-range Markerless 3D Human Motion Capture with LiDAR Point Clouds . In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition , 20470--20480
2022
-
[19]
Li, X.; Wang, W.; Yang, L.; and Yang, J. 2022 b . Uniform Masking : Enabling MAE Pre-training for Pyramid-based Vision Transformers with Locality . arXiv:2205.10063
2022 arXiv
-
[20]
Li, Y.; Yang, S.; Liu, P.; Zhang, S.; Wang, Y.; Wang, Z.; Yang, W.; and Xia, S.-T. 2022 c . SimCC : A Simple Coordinate Classification Perspective for Human Pose Estimation . In ECCV, 89--106
2022
-
[21]
Li, Z.; Ye, J.; Song, M.; Huang, Y.; and Pan, Z. 2021. Online Knowledge Distillation for Efficient Pose Estimation . In Proceedings of the IEEE / CVF International Conference on Computer Vision , 11740--11750
2021
-
[22]
Lian, J.; Wang, D.; Zhu, S.; Wu, Y.; and Li, C. 2022. Transformer-Based Attention Network for Vehicle Re-Identification. Electronics, 11(7)
2022
-
[23]
Lin, K.; Lin, C.-C.; Liang, L.; Liu, Z.; and Wang, L. 2024. MPT : Mesh Pre-Training With Transformers for Human Pose and Mesh Reconstruction . In Proceedings of the IEEE / CVF Winter Conference on Applications of Computer Vision , 3415--3425
2024
-
[24]
Loper, M.; Mahmood, N.; Romero, J.; Pons-Moll , G.; and Black, M. J. 2015. SMPL : A Skinned Multi-Person Linear Model. ACM Transactions on Graphics, 34(6): 248:1--248:16
2015
-
[25]
Loshchilov, I.; and Hutter, F. 2019. Decoupled Weight Decay Regularization . arXiv:1711.05101
2019 arXiv
-
[26]
F.; Pons-Moll , G.; and Black, M
Mahmood, N.; Ghorbani, N.; Troje, N. F.; Pons-Moll , G.; and Black, M. J. 2019. AMASS : Archive of Motion Capture As Surface Shapes . In Proceedings of the IEEE / CVF International Conference on Computer Vision , 5442--5451
2019
-
[27]
Z.; Yan, X.; Yan, H.; Qiao, S.; Zhu, Y.; Chen, L.-C.; Kretzschmar, H.; and Anguelov, D
Mei, J.; Zhu, A. Z.; Yan, X.; Yan, H.; Qiao, S.; Zhu, Y.; Chen, L.-C.; Kretzschmar, H.; and Anguelov, D. 2022. Waymo Open Dataset : Panoramic Video Panoptic Segmentation . arXiv:2206.07704
2022 arXiv
-
[28]
Qiu, Z.; Qiu, K.; Fu, J.; and Fu, D. 2023. Weakly-Supervised Pre-Training for 3D Human Pose Estimation via Perspective Knowledge. Pattern Recognition, 139: 109497
2023
-
[29]
Ren, Y.; Han, X.; Yao, Y.; Long, X.; Sun, Y.; and Ma, Y. 2024 a . LiveHPS ++: Robust and Coherent Motion Capture in Dynamic Free Environment . In ECCV, 127--144
2024
-
[30]
Ren, Y.; Han, X.; Zhao, C.; Wang, J.; Xu, L.; Yu, J.; and Ma, Y. 2024 b . LiveHPS : LiDAR-based Scene-level Human Pose and Shape Estimation in Free Environment . In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition , 1281--1291
2024
-
[31]
Ren, Y.; Zhao, C.; He, Y.; Cong, P.; Liang, H.; Yu, J.; Xu, L.; and Ma, Y. 2023. LiDAR-aid Inertial Poser : Large-scale Human Motion Capture by Sparse Inertial and LiDAR Sensors . IEEE Transactions on Visualization and Computer Graphics, 29(5): 2337--2347
2023
-
[32]
Shan, W.; Liu, Z.; Zhang, X.; Wang, S.; Ma, S.; and Gao, W. 2022. P- STMO : Pre-trained Spatial Temporal Many-to-One Model for 3D Human Pose Estimation . In ECCV, 461--478
2022
-
[33]
Simon, T.; Joo, H.; Matthews, I.; and Sheikh, Y. 2017. Hand Keypoint Detection in Single Images Using Multiview Bootstrapping . In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition , 4645--4653
2017
-
[34]
Su, Z.; Xu, L.; Zheng, Z.; Yu, T.; Liu, Y.; and Fang, L. 2020. RobustFusion : Human Volumetric Capture with Data-Driven Visual Cues Using a RGBD Camera . In ECCV, 246--264
2020
-
[35]
Sun, P.; Kretzschmar, H.; Dotiwalla, X.; Chouard, A.; Patnaik, V.; Tsui, P.; Guo, J.; Zhou, Y.; Chai, Y.; Caine, B.; Vasudevan, V.; Han, W.; Ngiam, J.; Zhao, H.; Timofeev, A.; Ettinger, S.; Krivokon, M.; Gao, A.; Joshi, A.; Zhang, Y.; Shlens, J.; Chen, Z.; and Anguelov, D. 202...
2020
-
[36]
Tu, H.; Wang, C.; and Zeng, W. 2020. VoxelPose : Towards Multi-camera 3D Human Pose Estimation in Wild Environment . In ECCV, 197--212
2020
-
[37]
Wang, H.; Liu, Q.; Yue, X.; Lasenby, J.; and Kusner, M. J. 2021. Unsupervised Point Cloud Pre-Training via Occlusion Completion . In Proceedings of the IEEE / CVF International Conference on Computer Vision , 9782--9792
2021
-
[38]
S.; Ji, J.; Najibi, M.; Zhou, Y.; and Anguelov, D
Weng, Z.; Gorban, A. S.; Ji, J.; Najibi, M.; Zhou, Y.; and Anguelov, D. 2023 a . 3D Human Keypoints Estimation From Point Clouds in the Wild Without Human Labels . In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition , 1158--1167
2023
-
[39]
S.; Ji, J.; Najibi, M.; Zhou, Y.; and Anguelov, D
Weng, Z.; Gorban, A. S.; Ji, J.; Najibi, M.; Zhou, Y.; and Anguelov, D. 2023 b . 3D Human Keypoints Estimation From Point Clouds in the Wild Without Human Labels . In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition , 1158--1167
2023
-
[40]
Wu, X.; Jiang, L.; Wang, P.-S.; Liu, Z.; Liu, X.; Qiao, Y.; Ouyang, W.; He, T.; and Zhao, H. 2024. Point Transformer V3 : Simpler Faster Stronger . In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition , 4840--4851
2024
-
[41]
Xu, Y.; Zhang, J.; ZHANG, Q.; and Tao, D. 2022. ViTPose : Simple Vision Transformer Baselines for Human Pose Estimation. In Advances in Neural Information Processing Systems, volume 35, 38571--38584
2022
-
[42]
Yan, M.; Wang, X.; Dai, Y.; Shen, S.; Wen, C.; Xu, L.; Ma, Y.; and Wang, C. 2023. CIMI4D : A Large Multimodal Climbing Motion Dataset under Human-scene Interactions . In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition , 12977--12988
2023
-
[43]
Yan, M.; Zhang, Y.; Cai, S.; Fan, S.; Lin, X.; Dai, Y.; Shen, S.; Wen, C.; Xu, L.; Ma, Y.; and Wang, C. 2024. RELI11D : A Comprehensive Multimodal Human Motion Dataset and Method . In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition , 2250--2262
2024
-
[44]
Ye, D.; Xie, Y.; Chen, W.; Zhou, Z.; Ge, L.; and Foroosh, H. 2024. LPFormer : LiDAR Pose Estimation Transformer with Multi-Task Network . In IEEE International Conference on Robotics and Automation , 16432--16438
2024
-
[45]
Ye, H.; Zhu, W.; Wang, C.; Wu, R.; and Wang, Y. 2022. Faster VoxelPose : Real-time 3D Human Pose Estimation by Orthographic Projection . In ECCV, 142--159
2022
-
[46]
Ying, J.; and Zhao, X. 2021. Rgb- D Fusion For Point-Cloud-Based 3d Human Pose Estimation . In Proceedings of IEEE International Conference on Image Processing , 3108--3112
2021
-
[47]
Zhang, J.; Mao, Q.; Hu, G.; Shen, S.; and Wang, C. 2024. Neighborhood- Enhanced 3D Human Pose Estimation with Monocular LiDAR in Long-Range Outdoor Scenes . Proceedings of the AAAI Conference on Artificial Intelligence, 38(7): 7169--7177
2024
-
[48]
Zhao, Z.; Luo, L.; Pan, S.; Zhang, C.; and Gong, C. 2024. Graph Stochastic Neural Process for Inductive Few-shot Knowledge Graph Completion. arXiv preprint arXiv:2408.01784
2024 arXiv
-
[49]
R.; Liu, T.; Chari, V.; Cornman, A.; Zhou, Y.; Li, C.; and Anguelov, D
Zheng, J.; Shi, X.; Gorban, A.; Mao, J.; Song, Y.; Qi, C. R.; Liu, T.; Chari, V.; Cornman, A.; Zhou, Y.; Li, C.; and Anguelov, D. 2022. Multi-Modal 3D Human Pose Estimation with 2D Weak Supervision in Autonomous Driving . In Proceedings of IEEE / CVF Conference on Computer Vis...
2022
-
[50]
, " * write output.state after.block = add.period write newline
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint howpublished institution isbn journal key month note number organization pages publisher school series title type volume year label extra.label sort.label short.list INTEGERS output.state before.all...
-
[51]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.