{"id":"d7627072-e330-4966-b457-691793f7eadb","arxiv_id":"2411.10582","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"DiffOpt integrates a motion diffusion prior via score distillation into a neural motion field optimization with joint human-camera updates, improving global trajectory recovery for long dynamic-camera videos.","lead":"The paper proposes DiffOpt, a three-stage optimization method that uses a pretrained motion diffusion model as a prior to improve 3D human mesh and trajectory reconstruction from monocular video captured by a moving camera. It reports better global trajectory accuracy than several existing methods on the EMDB and Egobody benchmarks, especially for long untrimmed sequences.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The long-video superiority claim rests on an incomplete baseline matrix: untrimmed EMDB omits WHAM and TRACE, Egobody omits GLAMR and TRACE, so 'superior over other SOTA' is not established until all named baselines are run on both benchmarks.","rationale":"I read the paper in good faith and accept that DiffOpt is a plausible new use of motion-diffusion priors for global HMR. The proposed optimization is coherent, the ablations show the MDM stage matters, and the paper's own limitations section honestly names the distribution-shift risk in the pretrained MDM prior. However, the headline claim is an empirical superiority claim, and the experiments as reported do not yet establish it. The reader's weakest assumption targets the MDM prior's distribution match; that is a real concern and is supported by the paper's own limitations. But the more immediate load-bearing weakness is that the long-video evidence is missing key baselines: the untrimmed EMDB result (Table 3) omits WHAM and TRACE, and the Egobody result (Table 4) omits GLAMR and TRACE. Since the claim explicitly says 'other state-of-the-art global HMR methods' and emphasizes long videos, the absence of these baselines from the long-video benchmarks is not a minor reporting gap; it is the difference between a supported claim and an unsupported one. I also note the selective exclusion of WHAM's NaN failure from the Table 2 mean, which makes the comparison more favorable to DiffOpt, and the lack of error bars across all tables. These are evidence issues, not accusations of misconduct, and they could be resolved by additional experiments. Because the reader already assigned CONDITIONAL with medium correctness risk, my concern does not move the verdict; it reinforces the need for the authors to complete the baseline matrix before the superiority claim can be accepted.","tokens_in":13663,"tokens_out":5145,"duration_ms":55322,"concrete_test":"Run the released implementations of WHAM and TRACE on the same untrimmed EMDB sequences used for Table 3, and GLAMR and TRACE on the same Egobody test sequences used for Table 4, using the prescribed metric protocol; additionally recompute Table 2 means with per-sequence errors and bootstrap confidence intervals. If DiffOpt still has the best G-MPJPE/G-MPVPE on both benchmarks with the complete baseline matrix, the concern is resolved. If any omitted baseline surpasses DiffOpt on either benchmark, the superiority claim must be narrowed or withdrawn.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is empirical: DiffOpt achieves superior global HMR over other state-of-the-art GHMR methods, most prominently in long videos. The two experiments that support the long-video part of the claim are Table 3 (untrimmed EMDB) and Table 4 (Egobody). Table 3 compares DiffOpt only against GLAMR and SLAHMR; WHAM and TRACE, both evaluated in the trimmed experiment (Table 2), are absent from the long-sequence benchmark where the headline claim is strongest. Table 4 compares DiffOpt only against WHAM and SLAHMR; GLAMR and TRACE are absent. No single experiment in the paper includes all four baselines, so the superiority claim is not established for the full set of named SOTA methods. The issue is compounded by the trimmed-table handling of WHAM: WHAM's mean G-MPJPE/G-MPVPE is computed after removing the 'soccer warmup' sequence, where WHAM produced NaN. This exclusion makes the comparison to WHAM in Table 2 favorable to the claim, and no error bars or per-sequence variance are reported for any table. If, for example, WHAM or TRACE performs on untrimmed EMDB as it does on trimmed EMDB, the reported 16% G-MPJPE advantage could disappear. This is an evidence-completeness problem rather than an internal inconsistency, but it is decisive for the paper's headline claim.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes DiffOpt, an optimization-based global 3D human mesh recovery (GHMR) method for monocular videos captured by dynamic cameras. The method represents human motion with neural motion fields, initializes from off-the-shelf HMR2.0 pose estimates and DROID-SLAM camera parameters, and uses a pretrained motion diffusion model (MDM) as a score-distillation prior in a three-stage optimization. Experiments on trimmed and untrimmed EMDB and on Egobody compare against GLAMR, SLAHMR, WHAM, and TRACE, with a focus on global metrics (G-MPJPE and G-MPVPE). The paper reports improved global motion recovery on long sequences and includes ablations showing the importance of the neural motion field and the multi-stage optimization scheme.","tokens_in":13943,"tokens_out":7258,"duration_ms":72612,"significance":"If the central claim holds, DiffOpt would be a meaningful contribution to global HMR, demonstrating that a generative motion prior can regularize root trajectory estimation and camera-human disentanglement. The paper has clear strengths: it leverages a strong pretrained motion prior (MDM) through score distillation, uses an implicit neural motion representation for temporal consistency, evaluates on two established benchmarks, and provides ablations that isolate the motion representation and optimization scheme. However, the headline claim of 'superior global human motion recovery over other state-of-the-art' is only partly supported because the long-video experiments compare against an incomplete baseline set and because the handling of WHAM's failure on one trimmed sequence skews the reported averages. The results are promising but require additional comparisons and more careful statistical reporting before the central claim is fully established.","major_comments":[{"comment":"The SDS objective in Eq. (3) is written as E[ w(t) || eps_phi(alpha_t x + sigma_t epsilon, t) - epsilon ||^2 ], which is the standard noise-prediction form of score distillation. However, Eq. (2) defines eps_phi as an x0-prediction network trained with || x0 - eps_phi(x_t, t) ||^2, as in MDM. If eps_phi outputs x0 rather than the noise epsilon, then subtracting the sampled noise epsilon from the network output is not a valid score-distillation target; for an x0-predictor, the SDS target should be the implied noise (x_t - alpha_t x0)/sigma_t or an x0-based objective. This discrepancy is load-bearing because stage 2 and stage 3 both use L_Diff as the motion prior. Please clarify which network is actually used and provide the exact gradient expression that is implemented.","section":"Tables 3 and 4, Sections 4.1.2 and 4.2"},{"comment":"The long-video claim is supported only against a partial baseline set. Table 3 compares DiffOpt to GLAMR and SLAHMR but omits WHAM and TRACE, both of which are named in the introduction and evaluated in Table 2. Table 4 compares DiffOpt only to WHAM and SLAHMR, omitting GLAMR and TRACE. Thus no experiment establishes superiority over all four baselines in the long-video regime where the headline claim is strongest. Given WHAM's strong trimmed global metrics (mean G-MPJPE 216.5 mm after exclusion), the untrimmed comparison with WHAM is essential. The paper should either run the missing baselines on both benchmarks or restrict the claim to the methods actually compared in each setting.","section":"Tables 3 and 4, Sections 4.1.2 and 4.2"},{"comment":"The mean G-MPJPE and G-MPVPE for WHAM are computed after excluding the 'soccer warmup' sequence because WHAM fails completely on that sequence. This exclusion is not neutral: it raises WHAM's mean and makes the statement that DiffOpt is 'marginally trailing behind WHAM' more favorable to the paper's narrative. For a fair comparison, report the mean including all seven sequences, with an explicit failure handling policy (e.g., NaN or a large penalty), and report per-sequence values in the main text. Without this, the claim that DiffOpt 'consistently outperforms most other methods' is not supported by the table as presented.","section":"Table 2 and Section 4.1.1"},{"comment":"No per-sequence variance, confidence intervals, or significance tests are reported anywhere in the quantitative evaluation. With only seven trimmed EMDB sequences and a single untrimmed mean per method, the claimed improvements (16% on untrimmed EMDB, 24.6% on Egobody) could be driven by one or two outlier sequences, notably SLAHMR's numerical breakdown on untrimmed EMDB. Please report per-sequence errors or at least standard deviations/confidence intervals, and identify which sequences drive the average improvement.","section":"Tables 2-4"},{"comment":"The paper's own limitations state that DiffOpt degrades on static-leg postures and on motions with external forces beyond ground contact, and the 'outdoor warmup' performance drop is attributed to contact with rigid objects being out-of-distribution for the AMASS-trained MDM. These are in-the-wild scenarios that the abstract explicitly invokes, so the superiority claim should be scoped to motions consistent with the MDM training distribution, or the method should be evaluated on a broader set of such interactions. This is a scope issue rather than an internal inconsistency, but it should be acknowledged in the abstract and conclusion.","section":"Section 5 and Section 4.1.1"}],"minor_comments":[{"comment":"There is a typo: 'the the Skinned Multi-Person Linear' should read 'the Skinned Multi-Person Linear'.","section":"Section 2.1"},{"comment":"The caption says 'comparing its performance against GLAMR and SLAHMR' but the table actually includes WHAM, TRACE, and DiffOpt as well; please update the caption to reflect the full set of methods.","section":"Table 2 caption"},{"comment":"The symbol t is used both for the diffusion timestep in Eq. (3) and for the frame index in Eqs. (8)-(10), which is confusing; please use distinct symbols, for example s for the diffusion step and i for the frame index.","section":"Section 3.2 and Section 3.3.2"},{"comment":"The reference 'Mahmood et al.' is incomplete; the full author list should be given, and the in-text citation 'et al. (2019)' should be replaced with a proper citation key such as 'Mahmood et al. (2019)'.","section":"References"},{"comment":"The initial translation heuristic is said to be described in the Supplementary Material, but no supplementary material is included with the arXiv version; please provide the details or a reference to a public supplement.","section":"Section 3.3.1"},{"comment":"The conclusion says 'including GLAMR and SLAHMR' when the paper actually compares against four baselines; please update the wording to reflect the full set of methods evaluated.","section":"Section 6"},{"comment":"The qualitative example shows the 'soccer warmup' sequence; consider also showing a known failure case such as 'outdoor warmup' to give readers a balanced view of the method's limitations.","section":"Figure 2"}],"recommendation":"major_revision","confidential_remarks":"The paper presents a promising optimization-based approach, but the empirical claims need to be tightened before publication. In particular, the authors should not claim superiority over all four named state-of-the-art methods until the full baseline matrix is run on both benchmarks and the WHAM failure sequence is handled transparently. The inconsistency between the x0-prediction formulation in Eq. (2) and the noise-target SDS in Eq. (3) should also be resolved; if the implemented loss differs from the written one, the paper should state the exact objective. The lack of error bars is a further concern given the small number of sequences."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line: DiffOpt is a legitimate new combination—SDS from a pretrained motion diffusion model applied to a neural motion field for global HMR, with alternating human/camera updates. The ablations show the prior and the field representation each matter a lot. But the paper's central claim of superior global recovery is only partially supported by the experiments as reported.\n\nWhat's new: the specific framework of using MDM-SDS inside a neural motion field, with a warm-up stage, an MDM-guidance stage that alternates between human and camera updates, and a final joint fine-tuning. The method is clearly described, and the ablations are useful. On trimmed EMDB, DiffOpt beats GLAMR, SLAHMR, and TRACE on G-MPJPE, and removing the MDM step raises G-MPJPE from 323 to 468. That's real evidence the prior does work.\n\nThe soft spots are in the evaluation, not the method. On trimmed EMDB, WHAM's mean G-MPJPE (216.5) is a lot better than DiffOpt's (322.6); calling that 'marginally trailing' is misleading. The untrimmed EMDB table omits WHAM and TRACE, and the Egobody table omits GLAMR and TRACE, so no single experiment includes all four baselines. The long-video superiority claim is exactly where the baseline matrix is thinnest. WHAM's average is also computed after dropping the soccer-warmup sequence where WHAM NaNs; the text discloses this, but it still tilts the comparison. No error bars or per-sequence variance are reported. The limitations section names two cases where the MDM prior hurts (static legs, external forces); that's honest and consistent with the ablations, but it means the prior is not universally reliable.\n\nWho should read it: people working on global HMR with dynamic cameras, especially those interested in diffusion-based priors. A serious referee should engage with this, but the authors should be asked to run the missing baselines or narrow their claim.","headline":"A promising combination of diffusion priors and neural motion fields for global HMR, but the evaluation as reported doesn't fully support the 'superior over SOTA' claim.","tokens_in":14519,"tokens_out":3041,"would_cite":true,"duration_ms":29629,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"DiffOpt shows that a pretrained motion diffusion model, used as a score-distillation prior inside a multi-stage optimization, recovers globally coherent human motion from monocular moving-camera video, substantially outperforming prior…","keywords":["global human mesh recovery","motion diffusion model","score distillation sampling","dynamic camera","neural motion fields","monocular video","human motion prior","trajectory optimization"],"falsifier":"Run DiffOpt on a lengthy dynamic-camera sequence where the subject climbs or pushes against objects (external forces), and check whether the recovered global trajectory stays close to ground truth; the paper's claimed limitation predicts a measurable degradation in G-MPJPE on such sequences compared to flat-ground locomotion, so a sequence where the method fails to beat GLAMR would contradict the primary claim of superior robustness.","tokens_in":13430,"feed_emoji":"🏃","tokens_out":5412,"duration_ms":45855,"temperature":0.7,"pith_summary":"DiffOpt is a method for recovering global 3D human motion from monocular video shot with a moving camera. The paper's central claim is that a pretrained motion diffusion model, used as a score-distillation prior, can push an initial pose-and-trajectory estimate toward temporally coherent human motion, while a learnable camera update keeps the reprojection error low. On the EMDB and Egobody benchmarks, DiffOpt reports the best global trajectory metrics among compared state-of-the-art methods, with the largest gains on long, untrimmed sequences where competitor optimization often collapses. If true, this makes markerless motion capture more practical for long in-the-wild footage, because the method does not need multi-view setups or static cameras.","feed_headline":"DiffOpt tops global human-motion baselines on long videos","feed_subtitle":"A motion diffusion prior disentangles human and camera movement, cutting global trajectory error on long clips.","key_machinery":"The central mechanism is the motion diffusion model (MDM) used as a prior through score distillation sampling (SDS) loss, jointly with a reprojection loss. The human motion is parameterized by neural motion fields (three MLPs for pose, root orientation, and translation), and camera motion is parameterized by learnable rotation bias, translation scale and bias, and focal length scale on top of DROID-SLAM estimates. A three-stage optimization first warms up the fields to an off-the-shelf HMR estimate, then alternates human-motion updates (MDM-SDS plus pose loss) and camera updates (2D reprojection), then fine-tunes all parameters together.","core_discovery":"The paper discovers that a motion diffusion model (MDM) trained on AMASS provides a strong enough prior over coherent human motion to guide global human mesh recovery: by representing the human motion with neural motion fields and optimizing them with the MDM score distillation sampling (SDS) loss against a reprojection loss, DiffOpt correctly disentangles human root translation from camera motion. The result is that the global root trajectory, not just per-frame pose, is recovered accurately, and this advantage grows with sequence length. On untrimmed EMDB sequences, DiffOpt achieves a G-MPJPE of 1776.2 mm versus 2113.5 mm for GLAMR and 5595.8 mm for SLAHMR, a margin the paper attributes to the diffusion prior's temporal coherence.","pith_inferences":["The same MDM-SDS recipe could serve as a generic temporal regularizer for other optimization-based reconstruction tasks, such as multi-person or object-interaction scenes, where the prior is replaced with a scene-aware motion model.","Because the prior is what supplies long-range temporal coherence, swapping in a diffusion model trained on more diverse data (including external forces and static poses) should directly extend DiffOpt to the failure cases the paper names.","The camera-motion parameterization (bias, scale, focal length) suggests an implicit calibration step; DiffOpt might also be used to refine camera estimates from SLAM in human-centric video.","The paper's limitation that static legs and external-force interactions degrade performance implies that the current prior is too narrow for full in-the-wild deployment, and a prior conditioned on scene or contact information would be a natural next step."],"forward_implications":["Long untrimmed videos become tractable: DiffOpt's global trajectory error on EMDB untrimmed sequences is 1776.2 mm versus 2113.5 mm for GLAMR, so applications that need full-length motion capture no longer require splitting videos into short segments.","The multi-stage optimization scheme is essential: ablations show single-stage optimization degrades G-MPJPE by 625 mm, meaning the staged warm-up, MDM, and camera alternation is what makes the diffusion prior effective.","Better camera estimation can plug in: replacing DROID-SLAM with TRAM's masked-SLAM keeps DiffOpt competitive, so future improvements in camera parameter estimation should directly improve global human motion recovery.","The method inherits biases from its components: HMR2.0, ViTPose, and SLAM failures propagate into the final motion, and the motion prior limits the diversity of recoverable motions.","The reported gains are largest on sequences with significant global root translation, suggesting the diffusion prior primarily corrects trajectory-level drift rather than per-frame pose."],"supporting_citations":[{"why":"Provides the motion diffusion model whose SDS prior is the core of the optimization.","marker":"Tevet et al. (2022)"},{"why":"Formulates score distillation sampling, the loss that turns the diffusion model into a gradient signal.","marker":"Poole et al. (2022)"},{"why":"Introduces neural motion fields, the representation DiffOpt uses for pose, root orientation, and translation.","marker":"Wang et al. (2022)"},{"why":"HMR2.0 supplies the initial per-frame articulations that warm up and regularize the optimization.","marker":"Goel et al. (2023)"},{"why":"DROID-SLAM provides the initial camera poses and focal lengths that DiffOpt refines.","marker":"Teed & Deng (2022)"},{"why":"ViTPose detects the 2D keypoints used in the reprojection loss for the camera update.","marker":"Xu et al. (2022)"},{"why":"AMASS is the large-scale motion capture dataset used to train the MDM prior.","marker":"Mahmood et al. (2019)"},{"why":"EMDB is the evaluation dataset with ground-truth global 3D motion in the wild.","marker":"Kaufmann et al. (2023)"},{"why":"Egobody supplies the second long-video test set with interacting people from head-mounted devices.","marker":"Zhang et al. (2022b)"},{"why":"SLAHMR is the primary strong baseline for comparison, also using DROID-SLAM and a motion prior.","marker":"Ye et al. (2023)"}],"fun_headline_variants":["Diffusion prior steers 3D human motion from moving cameras","DiffOpt: diffusion prior decouples human and camera motion","Long-clip human motion accuracy improves via diffusion prior","DiffOpt uses motion diffusion to fix human root trajectories","Diffusion prior boosts 3D human motion reconstruction from dynamic cameras"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The pretrained motion diffusion prior, trained on ground-contact motions in AMASS, accurately represents the kinds of motion in the test videos; if the true motion involves static legs or external forces beyond gravity, the SDS loss can pull the trajectory toward plausible but wrong paths, and the paper's own limitations section says performance drops in exactly those cases.","fun_headline_variants_meta":{"raw":{"variants":["Diffusion prior steers 3D human motion from moving cameras","DiffOpt: diffusion prior decouples human and camera motion","Long-clip human motion accuracy improves via diffusion prior","DiffOpt uses motion diffusion to fix human root trajectories","Diffusion prior boosts 3D human motion reconstruction from dynamic cameras"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000523,"raw_usage":{"total_tokens":2546,"prompt_tokens":979,"completion_tokens":1567,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":595,"completion_tokens_details":{"reasoning_tokens":1483}},"tokens_in":595,"tokens_out":1567,"duration_ms":26964,"temperature":1.0,"reasoning_tokens":1483,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T19:32:36.709195+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run DiffOpt on a lengthy dynamic-camera sequence where the subject climbs or pushes against objects (external forces), and check whether the recovered global trajectory stays close to ground truth; the paper's claimed limitation predicts a measurable degradation in G-MPJPE on such sequences compared to flat-ground locomotion, so a sequence where the method fails to beat GLAMR would contradict the primary claim of superior robustness.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the motion diffusion model whose SDS prior is the core of the optimization."},{"cited_title":"Barron, and Ben Mildenhall","cited_arxiv_id":null,"evidence_quote":"Formulates score distillation sampling, the loss that turns the diffusion model into a gradient signal."},{"cited_title":"Humans in 4d: Reconstructing and tracking humans with transformers, 2023","cited_arxiv_id":null,"evidence_quote":"HMR2.0 supplies the initial per-frame articulations that warm up and regularize the optimization."},{"cited_title":"Droid-slam: Deep visual slam for monocular, stereo, and rgb-d cameras, 2022","cited_arxiv_id":null,"evidence_quote":"DROID-SLAM provides the initial camera poses and focal lengths that DiffOpt refines."},{"cited_title":"Amass: Archive of motion capture as surface shapes","cited_arxiv_id":null,"evidence_quote":"AMASS is the large-scale motion capture dataset used to train the MDM prior."},{"cited_title":"Emdb: The electromagnetic database of global 3d human pose and shape in the wild, 2023","cited_arxiv_id":null,"evidence_quote":"EMDB is the evaluation dataset with ground-truth global 3D motion in the wild."}],"review_version":1}