{"paper":{"title":"MAEPose: Self-Supervised Spatiotemporal Learning for Human Pose Estimation on mmWave Video","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"MAEPose shows masked autoencoding on unlabeled mmWave videos produces representations for accurate human pose estimation.","cross_cats":["cs.AI"],"primary_cat":"cs.CV","authors_text":"Kevin Chetty, Nadia Bianchi-Berthouze, Xijia Wei, Youngjun Cho, Yuan Fang","submitted_at":"2026-04-30T21:23:03Z","abstract_excerpt":"Millimetre-wave (mmWave) radar offers a more privacy-preserving alternative to RGB-based human pose estimation. However, existing methods typically rely on pre-extracted intermediate representations such as sparse point clouds or spectrogram images, where the rich spatiotemporal information naturally present in radar video streams is discarded for model learning, while such signal processing adds system complexity. In addition, existing solutions are mainly conducted in an end-to-end supervised manner without leveraging unlabelled raw video streams to learn generalized representations. In this"},"claims":{"count":4,"items":[{"kind":"strongest_claim","text":"MAEPose consistently outperforms state-of-the-art baselines by up to 22.1% in MPJPE p<0.05, and maintains robust accuracy under zero-shot bystander interference with only a 6.5% error increase.","source":"verdict.strongest_claim","status":"machine_extracted","claim_id":"C1","attestation":"unclaimed"},{"kind":"weakest_assumption","text":"The assumption that pre-training with masked autoencoding on unlabelled mmWave spectrogram videos learns representations that generalize to accurate multi-frame pose estimation via the heatmap decoder, particularly across different datasets and interference conditions.","source":"verdict.weakest_assumption","status":"machine_extracted","claim_id":"C2","attestation":"unclaimed"},{"kind":"one_line_summary","text":"MAEPose is a masked autoencoder that learns spatiotemporal representations from unlabeled mmWave radar videos to estimate human poses, outperforming baselines by up to 22.1% in MPJPE.","source":"verdict.one_line_summary","status":"machine_extracted","claim_id":"C3","attestation":"unclaimed"},{"kind":"headline","text":"MAEPose shows masked autoencoding on unlabeled mmWave videos produces representations for accurate human pose estimation.","source":"verdict.pith_extraction.headline","status":"machine_extracted","claim_id":"C4","attestation":"unclaimed"}],"snapshot_sha256":"d5c3bf5d7859c4e85283b43a9ab40e493e513855ab8c062affb83757ccce2a7c"},"source":{"id":"2605.00242","kind":"arxiv","version":2},"verdict":{"id":"3a189cba-b4ae-422f-85cd-adb8a7fac08d","model_set":{"reader":"grok-4.3"},"created_at":"2026-05-09T19:52:48.650959Z","strongest_claim":"MAEPose consistently outperforms state-of-the-art baselines by up to 22.1% in MPJPE p<0.05, and maintains robust accuracy under zero-shot bystander interference with only a 6.5% error increase.","one_line_summary":"MAEPose is a masked autoencoder that learns spatiotemporal representations from unlabeled mmWave radar videos to estimate human poses, outperforming baselines by up to 22.1% in MPJPE.","pipeline_version":"pith-pipeline@v0.9.0","weakest_assumption":"The assumption that pre-training with masked autoencoding on unlabelled mmWave spectrogram videos learns representations that generalize to accurate multi-frame pose estimation via the heatmap decoder, particularly across different datasets and interference conditions.","pith_extraction_headline":"MAEPose shows masked autoencoding on unlabeled mmWave videos produces representations for accurate human pose estimation."},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2605.00242/integrity.json","findings":[],"available":true,"detectors_run":[{"name":"ai_meta_artifact","ran_at":"2026-05-20T20:35:54.164874Z","status":"completed","version":"1.0.0","findings_count":0},{"name":"doi_compliance","ran_at":"2026-05-19T18:22:56.641746Z","status":"completed","version":"1.0.0","findings_count":0}],"snapshot_sha256":"0425fa93e6b894442a41036aada1d9f5e64c70ddf1c5781d762f63a56d4564ea"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"}