REVIEW 4 major objections 5 minor 100 references
Deblur-Avatar: Animatable Avatars from Motion-Blurred Monocular Videos
T0 review · 4 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read Modeling the body's pose trajectory during exposure lets 3D Gaussian avatars be trained on motion-blurred monocular video and still render sharp, re-poseable humans.
desk verdict A plausible first solution to a real problem, but the sharp-avatar claim needs external validation before it fully lands. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the Slerp-based human motion trajectory model. Given a blurred frame with an estimated body pose, the method learns two poses $\theta_{\text{start}}$ and $\theta_{\text{end}}$ at the exposure endpoints; intermediate poses follow spherical linear interpolation in quaternion space (Eqs. 8–9), so the body is assumed to move at constant angular velocity during exposure. This trajectory yields $n$ virtual sharp images via the standard deformable-Gaussian avatar pipeline (non-rigid cloth deformation, rigid skinning, rasterization), and Eq. (6) averages them into the blurred image. The second component is a pose-dependent fusion MLP that takes the pose latent code, view embedding, position encoding, and the rendered virtual image to output a per-pixel mask blending the blurred and sharp renders during training. At inference the mask and trajectory modules are dropped, leaving a standard sharp Gaussian avatar renderer.
What would settle it
Train on a synthetic sequence with a known ground-truth pose trajectory that includes strong acceleration (for example, a fast punch or a sudden stop) and compare the optimized start/end poses to the true ones. If the recovered trajectory deviates substantially from the ground truth while the rendered blur still matches the input, the trajectory model is absorbing error into the Gaussians rather than recovering true motion; alternatively, if residual ghosting remains along the accelerating limb, the constant-velocity Slerp prior is the bottleneck.
Extended reading notes
Core claim
The paper's central claim is that human-motion blur in monocular avatar capture can be inverted by explicitly modeling the exposure period with a human motion trajectory. For each blurred frame, the method optimizes two poses of the parametric body model SMPL, one at exposure start and one at end, interpolates between them with spherical linear interpolation to obtain $n$ virtual poses, deforms canonical 3D Gaussians to each virtual pose, rasterizes $n$ sharp images, and averages them to synthesize the blurred observation (Eqs. 5–9). A pose-dependent fusion network then predicts a per-pixel mask that blends the averaged blurred render with the sharp virtual render, so the loss supervises each pixel in the regime that produced it. After training, only the sharp branch is rendered, giving a crisp animatable avatar at real-time frame rates. The paper presents this as the first framework aimed at sharp animatable avatars from motion-blurred monocular videos, with synthetic and real-world experiments supporting the claim.
Load-bearing premise
The method rests on the assumption that every frame's blur is the average of sharp renders of the body moving at constant velocity between two endpoint poses during exposure, with camera shake, acceleration, and non-rigid cloth motion treated as negligible.
Editorial extensions
If this is right
- Monocular casual capture with fast movements or low-light long exposures becomes usable for avatar reconstruction without requiring sharp video.
- Pre-deblurring with 2D video deblurring networks becomes unnecessary; the paper's experiments show such preprocessing helps baselines only marginally, while the 3D trajectory model outperforms it.
- The avatar stays animatable to out-of-distribution poses because the sharp representation is in canonical Gaussian space, reposed via SMPL during inference.
- Real-time rendering is preserved: the added modules cost training time but no inference cost, keeping roughly 50 FPS.
- The number of virtual frames can stay small ($n=5$); more frames do not consistently improve quality and increase training cost.
Reading between the lines
- The same 'average a trajectory of articulated renders' recipe should transfer to other articulated subjects (quadrupeds, robots, hands) whenever a parametric skeleton and skinning are available; the core is the Slerp trajectory, not the specific body model.
- Because the only pose supervision comes from a single pose estimate on a blurred frame, the optimization may absorb trajectory error into Gaussian parameters; an interesting extension would be to add temporal consistency across frames or to train a blur-aware pose estimator so the trajectory is constrained before avatar fitting.
- The pose-dependent fusion mask is essentially a learned per-pixel blur map; it could be reused as a free motion-magnitude signal or to weight a confidence-aware loss, which the paper does not explore.
- The constant-velocity Slerp assumption is a strong prior; replacing it with per-joint quadratic or learned trajectories could address acceleration, at the cost of more unknowns, and the ablation's cubic B-spline result suggests modest headroom.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes Deblur-Avatar, a 3D Gaussian Splatting framework for reconstructing sharp, animatable human avatars from monocular videos containing motion blur caused by human movement. The method models the exposure interval of each frame by estimating two SMPL poses (θ_start, θ_end) at the exposure endpoints and interpolating them with spherical linear interpolation (Eq. 8) to produce n virtual poses. Each virtual pose is used to deform canonical Gaussians and render a sharp image; the n images are averaged (Eq. 6) to synthesize the blurred frame. A pose-dependent fusion MLP (Eqs. 16–17) blends the averaged blurred rendering with one virtual sharp image, supervised by the input blurred frame, and is discarded at inference. Experiments compare against six human-avatar baselines, with and without a video-deblurring front end, on a synthesized ZJU-MoCap-Blur dataset and on a self-captured real dataset with qualitative evaluation only.
Significance. If the central claim holds, the paper addresses a real gap: existing animatable avatar methods assume sharp inputs, while human-motion blur is common in casual monocular capture. The paper's strengths include a physically motivated blur formation model, a new synthetic dataset built from real high-frame-rate ZJU-MoCap footage, a self-captured real dataset, and an ablation study that examines the trajectory representation and the number of virtual frames. The authors also release code and models, and the method renders at 50 FPS, which is practically relevant. However, the quantitative evidence for the sharp-avatar claim is currently limited to the synthetic benchmark, and the fusion mechanism has a potential collapse mode that could allow the model to fit the blurred input without recovering sharp Gaussians. These gaps need to be addressed before the claim of recovering sharp avatars from real motion-blurred monocular videos is fully supported.
major comments (4)
- [§IV-A, Eq. (6)] The quantitative evaluation is carried out on ZJU-MoCap-Blur, where each blurred frame is created by averaging 17–49 high-fps sharp frames (Sec. IV-A). This is precisely the temporal-averaging formation model used in the paper's forward model (Eq. 6). As a result, high PSNR/SSIM/LPIPS on this benchmark demonstrate that the method can invert its own assumed blur formation process; they do not demonstrate generalization to real motion blur, which may include camera shake, acceleration, and complex non-rigid cloth motion. The Real-Human-Blur dataset has no sharp ground truth and is used only for qualitative comparison (Sec. IV-C), so the central claim of recovering sharp avatars from real motion-blurred monocular videos is not quantitatively supported. I recommend adding a quantitative real-world benchmark (e.g., a dataset with high-speed sharp ground truth and real optical blur, or a synthetic benchmark generated with a different, richer blur model including camera motion) to test the transfer.
- [§III-E, Eq. (17)] The pose-dependent fusion mask can, in principle, collapse to M(x,j) ≈ 1 everywhere: then Cout(x,j) = B(x,j), and the training loss is minimized by fitting the blurred rendering to the blurred input. Because the sharp virtual image C(x,j) is never supervised directly and the mask is discarded at inference, the optimization does not force the Gaussians to represent sharp content; it only forces the blended output to match Bgt. The paper reports no statistics or visualizations of the learned masks, so it is unclear whether the reported gains come from genuinely sharp Gaussians or from the mask copying the blurred image. I suggest adding a regularizer that discourages large mask values (e.g., a prior towards M=0 in regions with small motion), directly supervising C against a deblurred estimate, or reporting mask statistics and sharpness metrics on the reconstructed virtual image.
- [§III-C, Eq. (8)] The trajectory model is a strong assumption: it represents the entire human motion during exposure as a constant-velocity Slerp between two SMPL poses, with no camera motion, acceleration, or per-frame temporal consistency. The paper's Limitation (1) explicitly states that the method relies on a single input pose and does not explore inter-frame relationships. Since the synthetic benchmark is generated by averaging real high-fps frames, it does not test the Slerp assumption against camera shake or complex human accelerations; the real dataset is only qualitative. This is a correctness-risk concern, not an internal inconsistency. A concrete test would be to add synthetic camera shake or to use a real blurred sequence with a synchronized high-speed sharp camera to measure whether the Slerp model remains accurate enough to produce sharp avatars.
- [§IV-B, Table I] All quantitative results are single training runs with no error bars or significance tests. Several PSNR differences are small (e.g., sequence 377: 30.36 vs. 30.29 for GART; sequence 386: 33.75 vs. 33.68 for 3DGS-Avatar). The claim that the method 'significantly outperforms' baselines is therefore not statistically supported. Reporting multiple seeds or paired per-frame comparisons would strengthen the quantitative claims.
minor comments (5)
- [Header, Index Terms] The Index Terms line reads 'All-in-Focus synthesis, main/ultra-wide camera, occlusion-aware networks', which appears to be a copy-paste artifact from an unrelated paper and does not match the content of this manuscript; please correct it.
- [Figure 1 caption] The caption notation 'LPIPS ∗=LPIPS∗103' is confusing; it should read 'LPIPS × 10^3' and clearly define the axes, since the figure plots PSNR against LPIPS.
- [Figure 5] The panel labels cite [56]+[85], [56]+[62], [56]+[35], and [56]+[82] while the text and Table II refer to [84] as the video deblurring method; the figure references are inconsistent and need to be aligned.
- [§IV-A, Real-Human-Blur] The sentence 'We train and evaluate the Real-Human-Blur dataset at the resolution of 540 × 540, 960 × 540, and 360 × 640' is unclear because the Real-Human-Blur dataset is used only for qualitative evaluation; please clarify how these resolutions are used and whether any quantitative evaluation is performed.
- [Throughout] There are several typos and formatting inconsistencies, including 'Canoncial' in Figure 2, 'Arah' for ARAH in Table I, and 'T V' in the author affiliation line; a careful proofreading pass is needed.
Circularity Check
No significant circularity: the blur formation model is a physical forward model, the trajectory parameters are optimized rather than fitted to the target metric, and the only same-author citation is a non-load-bearing related-work contrast.
full rationale
The paper's derivation chain is self-contained. The claimed result is that jointly optimizing human-motion trajectories and 3D Gaussians, rendering virtual sharp frames, averaging them (Eq. 6), and blending with a pose-dependent mask recovers sharp avatars from blurred monocular video. The forward model in Eqs. (5)-(9) is the standard temporal integration of irradiance: B(x) is the exposure-time integral of sharp images C_t(x), discretized as an average of n virtual sharp renders. This is a physical image-formation equation, not a definition of the output in terms of the input loss. The optimized quantities are the SMPL trajectory endpoints theta_start and theta_end and the Gaussian parameters, all supervised by the blurred RGB input; the reported metrics are computed against held-out sharp ground-truth frames from ZJU-MoCap-Blur, so the sharp-view PSNR/SSIM/LPIPS numbers are not statistically forced by the training objective. The synthetic benchmark does use temporal averaging to create blur, matching the assumed physics, but it also uses external high-frame-rate frame interpolation and re-estimated SMPL poses, and it evaluates novel views against original sharp frames; any limitation here concerns generalization to real blur, not circularity. The Real-Human-Blur evaluation is qualitative only, which limits evidence strength but is not a circular step. The one self-citation, DyBlurRF (ref. [26], sharing author Zhiguo Cao), appears only as a related-work contrast stating that prior dynamic deblurring lacks specialized human motion representation; no derivation or numerical result depends on it, so it is not load-bearing. The pose-dependent fusion mask could in principle absorb trajectory errors during training, but it is ablated (Table III), optimized with a delayed schedule, and discarded at inference; this is a correctness consideration, not an equivalence between input and output. No equation in the paper reduces to its own input by construction, no fitted parameter is renamed as a prediction, and no uniqueness claim is imported from the authors' prior work.
Assumptions & free parameters
free parameters (3)
- per-frame SMPL start/end poses (theta_start, theta_end) =
learned per frame, 72-d pose vector each
- virtual frame count n =
5
- loss weights lambda1..lambda5 =
0.01, 0.1, 10 (decayed), 1, 100
assumptions (5)
- domain assumption The observed blurred image is the normalized integral of instantaneous sharp images over exposure (Eqs. 5-6).
- ad hoc to paper Human motion during exposure is well approximated by Slerp between two SMPL poses at exposure endpoints (Eq. 8).
- domain assumption SMPL parameters, camera calibration, and foreground masks from blurry video are sufficiently accurate.
- domain assumption Human motion is the dominant blur source; camera motion is negligible.
- domain assumption Linear blend skinning with learned skinning and non-rigid MLPs can represent clothed human deformation.
invented entities (1)
-
Virtual sharp image sequence, one per interpolated pose during exposure
Cite this review
Pith. "Pith review of Deblur-Avatar: Animatable Avatars from Motion-Blurred Monocular Videos." pith.science (2026). https://pith.science/paper/P3S4PWVV
@misc{pith2026250113335,
author = {Pith},
title = {Pith review of: Deblur-Avatar: Animatable Avatars from Motion-Blurred Monocular Videos},
year = {2026},
howpublished = {\url{https://pith.science/paper/P3S4PWVV}},
note = {Machine review of arXiv:2501.13335}
}
read the original abstract
We introduce a novel framework for modeling high-fidelity, animatable 3D human avatars from motion-blurred monocular video inputs. Motion blur is prevalent in real-world dynamic video capture, especially due to human movements in 3D human avatar modeling. Existing methods either (1) assume sharp image inputs, failing to address the detail loss introduced by motion blur, or (2) mainly consider blur by camera movements, neglecting the human motion blur which is more common in animatable avatars. Our proposed approach integrates a human movement-based motion blur model into 3D Gaussian Splatting (3DGS). By explicitly modeling human motion trajectories during exposure time, we jointly optimize the trajectories and 3D Gaussians to reconstruct sharp, high-quality human avatars. We employ a pose-dependent fusion mechanism to distinguish moving body regions, optimizing both blurred and sharp areas effectively. Extensive experiments on synthetic and real-world datasets demonstrate that our method significantly outperforms existing methods in rendering quality and quantitative metrics, producing sharp avatar reconstructions and enabling real-time rendering under challenging motion blur conditions. Code and models are available at https://github.com/xianrui-luo/deblur_avatar.
Figures
Figures from the paper (5 more)
Reference graph
Works this paper leans on
-
[1]
H-nerf: Neural radiance fields for rendering and temporal reconstruction of humans in motion,
H. Xu, T. Alldieck, and C. Sminchisescu, “H-nerf: Neural radiance fields for rendering and temporal reconstruction of humans in motion,” in NeurIPS, 2021
2021
-
[2]
A-nerf: Articulated neural radiance fields for learning human shape, appearance, and pose,
S.-Y . Su, F. Yu, M. Zollh ¨ofer, and H. Rhodin, “A-nerf: Articulated neural radiance fields for learning human shape, appearance, and pose,” in NeurIPS, 2021
2021
-
[3]
Neural articulated radiance field,
A. Noguchi, X. Sun, S. Lin, and T. Harada, “Neural articulated radiance field,” in ICCV, 2021
2021
-
[4]
Efficient neural radiance fields for interactive free-viewpoint video,
H. Lin, S. Peng, Z. Xu, Y . Yan, Q. Shuai, H. Bao, and X. Zhou, “Efficient neural radiance fields for interactive free-viewpoint video,” in SIGGRAPH Asia 2022 Conference Papers , 2022
2022
-
[5]
Vid2avatar: 3d avatar reconstruction from videos in the wild via self-supervised scene decomposition,
C. Guo, T. Jiang, X. Chen, J. Song, and O. Hilliges, “Vid2avatar: 3d avatar reconstruction from videos in the wild via self-supervised scene decomposition,” in CVPR, 2023
2023
-
[6]
Neuman: Neural human radiance field from a single video,
W. Jiang, K. M. Yi, G. Samei, O. Tuzel, and A. Ranjan, “Neuman: Neural human radiance field from a single video,” in ECCV, 2022
2022
-
[7]
Tava: Template-free animatable volumetric actors,
R. Li, J. Tanke, M. V o, M. Zollh ¨ofer, J. Gall, A. Kanazawa, and C. Lassner, “Tava: Template-free animatable volumetric actors,” in ECCV, 2022
2022
-
[8]
Animatable neural implicit surfaces for creating avatars from videos,
S. Peng, S. Zhang, Z. Xu, C. Geng, B. Jiang, H. Bao, and X. Zhou, “Animatable neural implicit surfaces for creating avatars from videos,” arXiv preprint arXiv:2203.08133 , vol. 4, no. 5, 2022
arXiv 2022
Show all 100 references
-
[9]
Neural actor: Neural free-view synthesis of human actors with pose control,
L. Liu, M. Habermann, V . Rudnev, K. Sarkar, J. Gu, and C. Theobalt, “Neural actor: Neural free-view synthesis of human actors with pose control,” ACM TOG, vol. 40, no. 6, pp. 1–16, 2021
2021
-
[10]
Arah: Animatable volume rendering of articulated human sdfs,
S. Wang, K. Schwarz, A. Geiger, and S. Tang, “Arah: Animatable volume rendering of articulated human sdfs,” in ECCV, 2022
2022
-
[11]
Hdhumans: A hybrid approach for high-fidelity digital humans,
M. Habermann, L. Liu, W. Xu, G. Pons-Moll, M. Zollhoefer, and C. Theobalt, “Hdhumans: A hybrid approach for high-fidelity digital humans,” Proceedings of the ACM on Computer Graphics and Inter- active Techniques, vol. 6, no. 3, pp. 1–23, 2023
2023
-
[12]
3d gaussian splatting for real-time radiance field rendering
B. Kerbl, G. Kopanas, T. Leimk ¨uhler, and G. Drettakis, “3d gaussian splatting for real-time radiance field rendering.” ACM TOG, vol. 42, no. 4, pp. 139–1, 2023
2023
-
[13]
Gaussianavatar: Towards realistic human avatar modeling from a single video via animatable 3d gaussians,
L. Hu, H. Zhang, Y . Zhang, B. Zhou, B. Liu, S. Zhang, and L. Nie, “Gaussianavatar: Towards realistic human avatar modeling from a single video via animatable 3d gaussians,” in CVPR, 2024
2024
-
[14]
Splatarmor: Articulated gaussian splatting for animatable humans from monocular rgb videos,
R. Jena, G. S. Iyer, S. Choudhary, B. Smith, P. Chaudhari, and J. Gee, “Splatarmor: Articulated gaussian splatting for animatable humans from monocular rgb videos,” arXiv preprint arXiv:2311.10812 , 2023
2023 arXiv
-
[15]
Hugs: Human gaussian splats,
M. Kocabas, J.-H. R. Chang, J. Gabriel, O. Tuzel, and A. Ranjan, “Hugs: Human gaussian splats,” in CVPR, 2024
2024
-
[16]
Gart: Gaussian articulated template models,
J. Lei, Y . Wang, G. Pavlakos, L. Liu, and K. Daniilidis, “Gart: Gaussian articulated template models,” in CVPR, 2024
2024
-
[17]
Animatable gaussians: Learning pose-dependent gaussian maps for high-fidelity human avatar model- ing,
Z. Li, Z. Zheng, L. Wang, and Y . Liu, “Animatable gaussians: Learning pose-dependent gaussian maps for high-fidelity human avatar model- ing,” in CVPR, 2024
2024
-
[18]
Animatable 3d gaussian: Fast and high-quality reconstruction of multiple human avatars,
Y . Liu, X. Huang, M. Qin, Q. Lin, and H. Wang, “Animatable 3d gaussian: Fast and high-quality reconstruction of multiple human avatars,” arXiv preprint arXiv:2311.16482 , 2023
2023 arXiv
-
[19]
Gva: Reconstructing vivid 3d gaussian avatars from monoc- ular videos,
X. Liu, C. Wu, J. Liu, X. Liu, C. Zhao, H. Feng, E. Ding, and J. Wang, “Gva: Reconstructing vivid 3d gaussian avatars from monoc- ular videos,” CoRR, 2024
2024
-
[20]
3dgs-avatar: Animatable avatars via deformable 3d gaussian splatting,
Z. Qian, S. Wang, M. Mihajlovic, A. Geiger, and S. Tang, “3dgs-avatar: Animatable avatars via deformable 3d gaussian splatting,” in CVPR, 2024. DEBLUR-A V ATAR: ANIMATABLE A V ATARS FROM MOTION-BLURRED MONOCULAR VIDEOS 11
2024
-
[21]
Deblur-nerf: Neural radiance fields from blurry images,
L. Ma, X. Li, J. Liao, Q. Zhang, X. Wang, J. Wang, and P. V . Sander, “Deblur-nerf: Neural radiance fields from blurry images,” in CVPR, 2022
2022
-
[22]
Dp-nerf: Deblurred neural radiance field with physical scene priors,
D. Lee, M. Lee, C. Shin, and S. Lee, “Dp-nerf: Deblurred neural radiance field with physical scene priors,” in CVPR, 2023
2023
-
[23]
Bags: Blur agnostic gaussian splatting through multi-scale kernel modeling,
C. Peng, Y . Tang, Y . Zhou, N. Wang, X. Liu, D. Li, and R. Chellappa, “Bags: Blur agnostic gaussian splatting through multi-scale kernel modeling,” arXiv preprint arXiv:2403.04926 , 2024
2024 arXiv
-
[24]
Bad-nerf: Bundle adjusted deblur neural radiance fields,
P. Wang, L. Zhao, R. Ma, and P. Liu, “Bad-nerf: Bundle adjusted deblur neural radiance fields,” in CVPR, 2023
2023
-
[25]
Bad-gaussians: Bundle adjusted deblur gaussian splatting,
L. Zhao, P. Wang, and P. Liu, “Bad-gaussians: Bundle adjusted deblur gaussian splatting,” arXiv preprint arXiv:2403.11831 , 2024
2024 arXiv
-
[26]
Dyblurf: Dynamic neural radiance fields from blurry monocular video,
H. Sun, X. Li, L. Shen, X. Ye, K. Xian, and Z. Cao, “Dyblurf: Dynamic neural radiance fields from blurry monocular video,” in CVPR, 2024
2024
-
[27]
Video-based characters: creating new human performances from a multi-view video database,
F. Xu, Y . Liu, C. Stoll, J. Tompkin, G. Bharaj, Q. Dai, H.-P. Seidel, J. Kautz, and C. Theobalt, “Video-based characters: creating new human performances from a multi-view video database,” in ACM SIGGRAPH 2011 papers , 2011, pp. 1–10
2011
-
[28]
The relightables: V olumetric performance capture of humans with realistic relighting,
K. Guo, P. Lincoln, P. Davidson, J. Busch, X. Yu, M. Whalen, G. Harvey, S. Orts-Escolano, R. Pandey, J. Dourgarian et al. , “The relightables: V olumetric performance capture of humans with realistic relighting,” ACM TOG, vol. 38, no. 6, pp. 1–19, 2019
2019
-
[29]
Photorealistic monocular 3d reconstruction of humans wearing clothing,
T. Alldieck, M. Zanfir, and C. Sminchisescu, “Photorealistic monocular 3d reconstruction of humans wearing clothing,” in CVPR, 2022
2022
-
[30]
Pifu: Pixel-aligned implicit function for high-resolution clothed human digitization,
S. Saito, Z. Huang, R. Natsume, S. Morishima, A. Kanazawa, and H. Li, “Pifu: Pixel-aligned implicit function for high-resolution clothed human digitization,” in ICCV, 2019
2019
-
[31]
Pifuhd: Multi-level pixel- aligned implicit function for high-resolution 3d human digitization,
S. Saito, T. Simon, J. Saragih, and H. Joo, “Pifuhd: Multi-level pixel- aligned implicit function for high-resolution 3d human digitization,” in CVPR, 2020
2020
-
[32]
High-quality streamable free- viewpoint video,
A. Collet, M. Chuang, P. Sweeney, D. Gillett, D. Evseev, D. Calabrese, H. Hoppe, A. Kirk, and S. Sullivan, “High-quality streamable free- viewpoint video,” ACM TOG, vol. 34, no. 4, pp. 1–13, 2015
2015
-
[33]
Dynamicfusion: Recon- struction and tracking of non-rigid scenes in real-time,
R. A. Newcombe, D. Fox, and S. M. Seitz, “Dynamicfusion: Recon- struction and tracking of non-rigid scenes in real-time,” in CVPR, 2015
2015
-
[34]
Rapid avatar capture and simulation using commodity depth sensors,
A. Feng, A. Shapiro, W. Ruizhe, M. Bolas, G. Medioni, and E. Suma, “Rapid avatar capture and simulation using commodity depth sensors,” in ACM SIGGRAPH 2014 Talks , 2014, pp. 1–1
2014
-
[35]
Flyfusion: Realtime dynamic scene reconstruction using a flying depth camera,
L. Xu, W. Cheng, K. Guo, L. Han, Y . Liu, and L. Fang, “Flyfusion: Realtime dynamic scene reconstruction using a flying depth camera,” TVCG, vol. 27, no. 1, pp. 68–82, 2019
2019
-
[36]
Monocular, one-stage, regression of multiple 3d people,
Y . Sun, Q. Bao, W. Liu, Y . Fu, M. J. Black, and T. Mei, “Monocular, one-stage, regression of multiple 3d people,” in ICCV, 2021
2021
-
[37]
Vibe: Video inference for human body pose and shape estimation,
M. Kocabas, N. Athanasiou, and M. J. Black, “Vibe: Video inference for human body pose and shape estimation,” in CVPR, 2020
2020
-
[38]
Expressive body capture: 3d hands, face, and body from a single image,
G. Pavlakos, V . Choutas, N. Ghorbani, T. Bolkart, A. A. Osman, D. Tzionas, and M. J. Black, “Expressive body capture: 3d hands, face, and body from a single image,” in CVPR, 2019
2019
-
[39]
Smpl: A skinned multi-person linear model,
M. Loper, N. Mahmood, J. Romero, G. Pons-Moll, and M. J. Black, “Smpl: A skinned multi-person linear model,” in Seminal Graphics Papers: Pushing the Boundaries, Volume 2 , 2023, pp. 851–866
2023
-
[40]
Nerf: Representing scenes as neural radiance fields for view synthesis,
B. Mildenhall, P. P. Srinivasan, M. Tancik, J. T. Barron, R. Ramamoor- thi, and R. Ng, “Nerf: Representing scenes as neural radiance fields for view synthesis,” Communications of the ACM , vol. 65, no. 1, pp. 99–106, 2021
2021
-
[41]
Selfrecon: Self reconstruc- tion your digital avatar from monocular video,
B. Jiang, Y . Hong, H. Bao, and J. Zhang, “Selfrecon: Self reconstruc- tion your digital avatar from monocular video,” in CVPR, 2022
2022
-
[42]
Humannerf: Free-viewpoint rendering of moving people from monocular video,
C.-Y . Weng, B. Curless, P. P. Srinivasan, J. T. Barron, and I. Kemelmacher-Shlizerman, “Humannerf: Free-viewpoint rendering of moving people from monocular video,” in CVPR, 2022
2022
-
[43]
Monohuman: Animatable human neural field from monocular video,
Z. Yu, W. Cheng, X. Liu, W. Wu, and K.-Y . Lin, “Monohuman: Animatable human neural field from monocular video,” in CVPR, 2023
2023
-
[44]
Animatable neural radiance fields from monocular rgb videos,
J. Chen, Y . Zhang, D. Kang, X. Zhe, L. Bao, X. Jia, and H. Lu, “Animatable neural radiance fields from monocular rgb videos,” arXiv preprint arXiv:2106.13629, 2021
2021 arXiv
-
[45]
High-fidelity clothed avatar reconstruction from a single image,
T. Liao, X. Zhang, Y . Xiu, H. Yi, X. Liu, G.-J. Qi, Y . Zhang, X. Wang, X. Zhu, and Z. Lei, “High-fidelity clothed avatar reconstruction from a single image,” in CVPR, 2023
2023
-
[46]
Tech: Text-guided reconstruction of lifelike clothed humans,
Y . Huang, H. Yi, Y . Xiu, T. Liao, J. Tang, D. Cai, and J. Thies, “Tech: Text-guided reconstruction of lifelike clothed humans,” in 2024 International Conference on 3D Vision (3DV) . IEEE, 2024
2024
-
[47]
Neural body: Implicit neural representations with structured latent codes for novel view synthesis of dynamic humans,
S. Peng, Y . Zhang, Y . Xu, Q. Wang, Q. Shuai, H. Bao, and X. Zhou, “Neural body: Implicit neural representations with structured latent codes for novel view synthesis of dynamic humans,” in CVPR, 2021
2021
-
[48]
Snarf: Differentiable forward skinning for animating non-rigid neural implicit shapes,
X. Chen, Y . Zheng, M. J. Black, O. Hilliges, and A. Geiger, “Snarf: Differentiable forward skinning for animating non-rigid neural implicit shapes,” in ICCV, 2021
2021
-
[49]
Leap: Learning articulated occupancy of people,
M. Mihajlovic, Y . Zhang, M. J. Black, and S. Tang, “Leap: Learning articulated occupancy of people,” in CVPR, 2021
2021
-
[50]
Arch: Animatable reconstruction of clothed humans,
Z. Huang, Y . Xu, C. Lassner, H. Li, and T. Tung, “Arch: Animatable reconstruction of clothed humans,” in CVPR, 2020
2020
-
[51]
Motion-aware 3d gaussian splatting for efficient dynamic scene reconstruction,
Z. Guo, W. Zhou, L. Li, M. Wang, and H. Li, “Motion-aware 3d gaussian splatting for efficient dynamic scene reconstruction,” IEEE Transactions on Circuits and Systems for Video Technology , 2024
2024
-
[52]
Drivable 3d gaussian avatars,
W. Zielonka, T. Bagautdinov, S. Saito, M. Zollh ¨ofer, J. Thies, and J. Romero, “Drivable 3d gaussian avatars,” arXiv preprint arXiv:2311.08581, 2023
2023 arXiv
-
[53]
Human gaussian splatting: Real-time rendering of animat- able avatars,
A. Moreau, J. Song, H. Dhamo, R. Shaw, Y . Zhou, and E. P ´erez- Pellitero, “Human gaussian splatting: Real-time rendering of animat- able avatars,” in CVPR, 2024
2024
-
[54]
Haha: Highly articulated gaussian human avatars with textured mesh prior,
D. Svitov, P. Morerio, L. Agapito, and A. Del Bue, “Haha: Highly articulated gaussian human avatars with textured mesh prior,” arXiv preprint arXiv:2404.01053, 2024
2024 arXiv
-
[55]
Tgavatar: Reconstructing 3d gaussian avatars with transformer-based tri-plane,
R. Hu, X. Wang, Y . Yan, and C. Zhao, “Tgavatar: Reconstructing 3d gaussian avatars with transformer-based tri-plane,” IEEE Transactions on Circuits and Systems for Video Technology , 2025
2025
-
[56]
Gauhuman: Articulated gaussian splatting from monocular human videos,
S. Hu, T. Hu, and Z. Liu, “Gauhuman: Articulated gaussian splatting from monocular human videos,” in CVPR, 2024
2024
-
[57]
Gaussianbody: Clothed human reconstruction via 3d gaussian splatting,
M. Li, S. Yao, Z. Xie, K. Chen, and Y .-G. Jiang, “Gaussianbody: Clothed human reconstruction via 3d gaussian splatting,”arXiv preprint arXiv:2401.09720, 2024
2024 arXiv
-
[58]
Moss: Motion-based 3d clothed human synthesis from monocular video,
H. Wang, X. Cai, X. Sun, J. Yue, S. Zhang, F. Lin, and F. Wu, “Moss: Motion-based 3d clothed human synthesis from monocular video,” arXiv preprint arXiv:2405.12806 , 2024
2024 arXiv
-
[59]
Gomavatar: Efficient animatable human modeling from monocular video using gaussians-on-mesh,
J. Wen, X. Zhao, Z. Ren, A. G. Schwing, and S. Wang, “Gomavatar: Efficient animatable human modeling from monocular video using gaussians-on-mesh,” in CVPR, 2024
2024
-
[60]
Splattingavatar: Realistic real-time human avatars with mesh-embedded gaussian splatting,
Z. Shao, Z. Wang, Z. Li, D. Wang, X. Lin, Y . Zhang, M. Fan, and Z. Wang, “Splattingavatar: Realistic real-time human avatars with mesh-embedded gaussian splatting,” in CVPR, 2024
2024
-
[61]
Chase: 3d-consistent human avatars with sparse inputs via gaussian splatting and contrastive learning,
H. Zhao, H. Wang, C. Yang, and W. Shen, “Chase: 3d-consistent human avatars with sparse inputs via gaussian splatting and contrastive learning,” arXiv preprint arXiv:2408.09663 , 2024
2024 arXiv
-
[62]
Sg-gs: Photo- realistic animatable human avatars with semantically-guided gaussian splatting,
H. Zhao, C. Yang, H. Wang, X. Zhao, and W. Shen, “Sg-gs: Photo- realistic animatable human avatars with semantically-guided gaussian splatting,” arXiv preprint arXiv:2408.09665 , 2024
2024 arXiv
-
[63]
Blind deconvolution using a normalized sparsity measure,
D. Krishnan, T. Tay, and R. Fergus, “Blind deconvolution using a normalized sparsity measure,” in CVPR, 2011
2011
-
[64]
Single- image blind deblurring using multi-scale latent structure prior,
Y . Bai, H. Jia, M. Jiang, X. Liu, X. Xie, and W. Gao, “Single- image blind deblurring using multi-scale latent structure prior,” IEEE Transactions on Circuits and Systems for Video Technology , vol. 30, no. 7, pp. 2033–2045, 2019
2019
-
[65]
High-quality motion deblurring from a single image,
Q. Shan, J. Jia, and A. Agarwala, “High-quality motion deblurring from a single image,” ACM TOG, vol. 27, no. 3, pp. 1–10, 2008
2008
-
[66]
Non-uniform deblur- ring for shaken images,
O. Whyte, J. Sivic, A. Zisserman, and J. Ponce, “Non-uniform deblur- ring for shaken images,” in CVPR, 2010
2010
-
[67]
Deep multi-scale convolutional neural network for dynamic scene deblurring,
S. Nah, T. Hyun Kim, and K. Mu Lee, “Deep multi-scale convolutional neural network for dynamic scene deblurring,” in CVPR, 2017
2017
-
[68]
Deblurgan: Blind motion deblurring using conditional adversarial networks,
O. Kupyn, V . Budzan, M. Mykhailych, D. Mishkin, and J. Matas, “Deblurgan: Blind motion deblurring using conditional adversarial networks,” in CVPR, 2018
2018
-
[69]
Scale-recurrent network for deep image deblurring,
X. Tao, H. Gao, X. Shen, J. Wang, and J. Jia, “Scale-recurrent network for deep image deblurring,” in CVPR, 2018
2018
-
[70]
Deblurgan-v2: Deblur- ring (orders-of-magnitude) faster and better,
O. Kupyn, T. Martyniuk, J. Wu, and Z. Wang, “Deblurgan-v2: Deblur- ring (orders-of-magnitude) faster and better,” in ICCV, 2019
2019
-
[71]
Human-aware motion deblurring,
Z. Shen, W. Wang, X. Lu, J. Shen, H. Ling, T. Xu, and L. Shao, “Human-aware motion deblurring,” in ICCV, 2019
2019
-
[72]
Restormer: Efficient transformer for high-resolution image restoration,
S. W. Zamir, A. Arora, S. Khan, M. Hayat, F. S. Khan, and M.- H. Yang, “Restormer: Efficient transformer for high-resolution image restoration,” in CVPR, 2022
2022
-
[73]
Strip- former: Strip transformer for fast image deblurring,
F.-J. Tsai, Y .-T. Peng, Y .-Y . Lin, C.-C. Tsai, and C.-W. Lin, “Strip- former: Strip transformer for fast image deblurring,” in ECCV, 2022
2022
-
[74]
Multiscale structure guided diffusion for image deblurring,
M. Ren, M. Delbracio, H. Talebi, G. Gerig, and P. Milanfar, “Multiscale structure guided diffusion for image deblurring,” in ICCV, 2023
2023
-
[75]
Multi-scale residual low-pass filter network for image deblurring,
J. Dong, J. Pan, Z. Yang, and J. Tang, “Multi-scale residual low-pass filter network for image deblurring,” in ICCV, 2023
2023
-
[76]
Generalized video deblurring for dynamic scenes,
T. Hyun Kim and K. Mu Lee, “Generalized video deblurring for dynamic scenes,” in CVPR, 2015
2015
-
[77]
Cascaded deep video deblurring using temporal sharpness prior,
J. Pan, H. Bai, and J. Tang, “Cascaded deep video deblurring using temporal sharpness prior,” in CVPR, 2020
2020
-
[78]
Deep video deblurring for hand-held cameras,
S. Su, M. Delbracio, J. Wang, G. Sapiro, W. Heidrich, and O. Wang, “Deep video deblurring for hand-held cameras,” in CVPR, 2017
2017
-
[79]
Efficient spatio-temporal recurrent neural network for video deblurring,
Z. Zhong, Y . Gao, Y . Zheng, and B. Zheng, “Efficient spatio-temporal recurrent neural network for video deblurring,” in ECCV, 2020. 12 MANUSCRIPT SUBMITTED TO IEEE TRANS. ON CIRCUIT SYST. VIDEO TECHNOL
2020
-
[80]
Mc-blur: A comprehensive benchmark for image deblurring,
K. Zhang, T. Wang, W. Luo, W. Ren, B. Stenger, W. Liu, H. Li, and M.- H. Yang, “Mc-blur: A comprehensive benchmark for image deblurring,” IEEE Transactions on Circuits and Systems for Video Technology , vol. 34, no. 5, pp. 3755–3767, 2023
2023
-
[81]
Spatio-temporal deformable attention network for video deblurring,
H. Zhang, H. Xie, and H. Yao, “Spatio-temporal deformable attention network for video deblurring,” in ECCV, 2022
2022
-
[82]
Multi-attention convolutional neural network for video deblurring,
X. Zhang, T. Wang, R. Jiang, L. Zhao, and Y . Xu, “Multi-attention convolutional neural network for video deblurring,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 32, no. 4, pp. 1986– 1997, 2021
1986
-
[83]
Efficient video deblurring guided by motion magnitude,
Y . Wang, Y . Lu, Y . Gao, L. Wang, Z. Zhong, Y . Zheng, and A. Ya- mashita, “Efficient video deblurring guided by motion magnitude,” in ECCV, 2022
2022
-
[84]
Deep discriminative spatial and temporal network for efficient video deblurring,
J. Pan, B. Xu, J. Dong, J. Ge, and J. Tang, “Deep discriminative spatial and temporal network for efficient video deblurring,” in CVPR, 2023
2023
-
[85]
Exblurf: Efficient radiance fields for extreme motion blurred images,
D. Lee, J. Oh, J. Rim, S. Cho, and K. M. Lee, “Exblurf: Efficient radiance fields for extreme motion blurred images,” in ICCV, 2023
2023
-
[86]
Deblurring 3d gaussian splatting,
B. Lee, H. Lee, X. Sun, U. Ali, and E. Park, “Deblurring 3d gaussian splatting,” arXiv preprint arXiv:2401.00834 , 2024
2024 arXiv
-
[87]
Robust gaussian splatting,
F. Darmon, L. Porzi, S. Rota-Bul `o, and P. Kontschieder, “Robust gaussian splatting,” arXiv preprint arXiv:2404.04211 , 2024
2024 arXiv
-
[88]
Deblur-gs: 3d gaussian splatting from camera motion blurred images,
W. Chen and L. Liu, “Deblur-gs: 3d gaussian splatting from camera motion blurred images,” Proceedings of the ACM on Computer Graph- ics and Interactive Techniques , vol. 7, no. 1, pp. 1–15, 2024
2024
-
[89]
Crim-gs: Continuous rigid motion-aware gaussian splatting from motion blur images,
J. Lee, D. Kim, D. Lee, S. Cho, and S. Lee, “Crim-gs: Continuous rigid motion-aware gaussian splatting from motion blur images,” arXiv preprint arXiv:2407.03923, 2024
2024 arXiv
-
[90]
Structure-from-motion revisited,
J. L. Sch ¨onberger and J.-M. Frahm, “Structure-from-motion revisited,” in CVPR, 2016
2016
-
[91]
Animating rotation with quaternion curves,
K. Shoemake, “Animating rotation with quaternion curves,” in Pro- ceedings of the 12th annual conference on Computer graphics and interactive techniques, 1985
1985
-
[92]
Adam: A method for stochastic optimization,
D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” arXiv preprint arXiv:1412.6980 , 2014
2014 arXiv
-
[93]
Davanet: Stereo deblurring with view aggregation,
S. Zhou, J. Zhang, W. Zuo, H. Xie, J. Pan, and J. S. Ren, “Davanet: Stereo deblurring with view aggregation,” in CVPR, 2019
2019
-
[94]
Ntire 2019 challenge on video deblurring and super- resolution: Dataset and study,
S. Nah, S. Baik, S. Hong, G. Moon, S. Son, R. Timofte, and K. Mu Lee, “Ntire 2019 challenge on video deblurring and super- resolution: Dataset and study,” in CVPRW, 2019
2019
-
[95]
Extracting motion and appearance via inter-frame attention for efficient video frame interpolation,
G. Zhang, Y . Zhu, H. Wang, Y . Chen, G. Wu, and L. Wang, “Extracting motion and appearance via inter-frame attention for efficient video frame interpolation,” in CVPR, 2023
2023
-
[96]
Motion capture from internet videos,
J. Dong, Q. Shuai, Y . Zhang, X. Liu, X. Zhou, and H. Bao, “Motion capture from internet videos,” in ECCV, 2020
2020
-
[97]
Learning to reconstruct 3d human pose and shape via model-fitting in the loop,
N. Kolotouros, G. Pavlakos, M. J. Black, and K. Daniilidis, “Learning to reconstruct 3d human pose and shape via model-fitting in the loop,” in ICCV, 2019
2019
-
[98]
Sam 2: Segment anything in images and videos,
N. Ravi, V . Gabeur, Y .-T. Hu, R. Hu, C. Ryali, T. Ma, H. Khedr, R. R ¨adle, C. Rolland, L. Gustafson, E. Mintun, J. Pan, K. V . Alwala, N. Carion, C.-Y . Wu, R. Girshick, P. Doll ´ar, and C. Feichtenhofer, “Sam 2: Segment anything in images and videos,” arXiv preprint arXiv:...
2024 arXiv
-
[99]
Amass: Archive of motion capture as surface shapes,
N. Mahmood, N. Ghorbani, N. F. Troje, G. Pons-Moll, and M. J. Black, “Amass: Archive of motion capture as surface shapes,” in ICCV, 2019
2019
-
[100]
Ai choreographer: Music conditioned 3d dance generation with aist++,
R. Li, S. Yang, D. A. Ross, and A. Kanazawa, “Ai choreographer: Music conditioned 3d dance generation with aist++,” in ICCV, 2021
2021
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.