REVIEW 4 major objections 6 minor 121 references
Camera motion in video diffusion is set in the high-noise stage, so large self-supervised clip pairs plus a little paired data only there can re-shoot videos without 3D.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-31 12:58 UTC pith:PRVKRQBQ
load-bearing objection Solid engineering recipe for 3D-free re-shooting: timestep-routed self-supervision plus a little high-noise pair data actually moves the needle, even if the routing story is not fully isolated from data scale. the 4 major comments →
TARS: Timestep-Aware Data Scaling for 3D-Free Video Re-Shooting
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Timestep-wise analysis shows camera motion and coarse spatiotemporal structure form mainly in high-noise denoising stages; therefore a 3D-free two-stage recipe—large-scale self-supervised clip splitting for camera, viewpoint, and appearance, plus minimal cross-pair supervision only in the high-noise regime—yields more accurate, temporally consistent re-shooting and text-driven control of shot scale, viewing angle, and perspective than 3D-prior or fully paired baselines.
What carries the argument
Timestep-aware data routing: self-supervised pairs from temporally split unlabeled clips train all timesteps for structure, camera grids, and text viewpoint; scarce cross-pair data is applied only in the high-noise band (empirically t in [0.95, 1.0]) so a shortcut under guidance elicits motion-synced subject dynamics.
Load-bearing premise
Camera motion and subject structure are concentrated enough in a narrow high-noise band that a little paired supervision there alone can force synchronized motion without true multi-view pairs at later timesteps.
What would settle it
Retrain or ablate with cross-pair data barred from high noise and allowed only mid/low noise (or widen the high-noise band): if camera accuracy and motion sync (R-Pre, T-Pre, V-MPGE) then collapse relative to the reported full model, the routing claim fails; if they hold, the concentration premise is wrong.
If this is right
- Re-shooting no longer needs full 3D/4D reconstruction or large multi-camera paired corpora for competitive trajectory control.
- Text can specify shot scale, viewing angle, and first-/third-person perspective jointly with a camera grid, including reverse-angle and large-motion views.
- Unlabeled video scale becomes the main lever for viewpoint generalization and plausible synthesis of regions outside the source frame.
- The same high-noise vs mid/low split can guide where expensive paired supervision is spent in other controllable video tasks.
Where Pith is reading between the lines
- If high-noise stages truly own global geometry, other sparse geometric controls (depth, pose, layout) may also train cheaply with self-supervision everywhere and paired data only early.
- The reported shortcut under classifier-free guidance suggests many video editors could bootstrap temporal sync from tiny paired sets once strong unpaired priors exist.
- Failure modes on extreme open-domain dynamics would show up first as motion desync rather than identity drift, pointing where to add the next paired data.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. TARS proposes a 3D-free video re-shooting framework that combines camera-grid conditioning with text-driven semantic viewpoint control (shot scale, viewing angle, first-/third-person perspective). Motivated by a timestep-wise analysis arguing that camera motion and coarse structure form mainly in high-noise stages, the method uses a two-stage data strategy: large-scale self-supervised pairs obtained by temporally splitting monocular videos (with estimated camera grids and MLLM viewpoint captions), plus a small amount of cross-pair data applied only in a narrow high-noise band (t∈[0.95,1.0]) to elicit motion synchronization via an asserted “shortcut effect” under CFG. Experiments on a 1021-sample set report gains over CamClone, TrajCrafter, and SD-2.0 on camera accuracy, spatio-temporal consistency, content expansion, and viewpoint metrics, with ablations of Stage 1/2 and high vs mid/low source injection.
Significance. If the claims hold, the work is a practically important contribution to controllable video generation: it attacks the paired-data bottleneck for re-shooting without 3D reconstruction, unifies geometric camera control with semantic viewpoint instructions, and shows plausible synthesis under large motions including reverse-angle and perspective switches. The timestep-aware data-routing idea is a useful engineering principle for diffusion training under scarce multi-view supervision. Strengths include a balanced multi-category evaluation set, quantitative comparison against three relevant baselines (Tables 1–2), and ablations that partially support the high-noise vs mid/low division (Tables 3–4, Fig. 3). The result would matter for cinematography tools and camera-controllable video models more broadly, provided the routing claim is isolated from raw data scale and the evaluation is made more reproducible.
major comments (4)
- [Method (Timestep-Aware Data Routing); Tables 3–4] Method, Stage 2 and Experimental Setups; Tables 3–4: The central claim that scarce cross-pair supervision need only be applied in t∈[0.95,1.0] (via a “shortcut effect”) is not isolated from data scale. Ablations remove Stage 1, remove Stage 2, or change where the source video is injected at inference, but there is no control that trains the same 1M self-supervised + 60K cross-pair mixture with the cross-pair objective on all timesteps (or on a wider high-noise band). Without that comparison, Table 1 gains cannot be attributed to timestep-aware routing rather than simply more and more diverse training data. This control is load-bearing for the paper’s title and abstract claim.
- [Experimental Setups; Method Stage 2; Figure 3] Experimental Setups defines the high-noise regime as t∈[0.95,1.0] empirically, and Stage 2 asserts that synchronized motion is “easier to optimize” under CFG without a counter-example or quantitative probe (e.g., multi-modal motion pairs where unsynced completions are equally valid, or sweeps of the cutoff). Fig. 3 and the High vs Mid&Low rows support a coarse frequency split but do not establish that a 5% noise band is necessary or sufficient for open-domain dynamics. A cutoff sensitivity study and at least one failure-mode analysis of the shortcut would substantially strengthen the mechanism.
- [Experiments (Evaluation Metrics; Comparisons)] Comparisons and metrics: The backbone is an unspecified in-house T2V model, while several primary metrics (viewpoint accuracy, CE, FDR, VDR, FSCS, and parts of consistency) are judged by Gemini 3.1 Pro. This combination weakens external validity and reproducibility of the SOTA claim in Table 1–2. Please (i) clarify backbone capacity/training relative to baselines or release comparable checkpoints/protocol, (ii) report inter-judge agreement or human studies on a subset for LLM-judged metrics, and (iii) add standard low-level video metrics where applicable so that camera and identity gains are not solely model-judged.
- [Method Stage 1; Eq. (6)] Self-supervised construction (Eq. 6): V1/V2 are temporal halves of one monocular video with G2 estimated from V2. Large-gap non-overlapping splits are said to teach hallucination of unseen regions, but the paper does not quantify pose-estimation error, temporal gap distributions, or how often “unseen” content is truly out-of-view versus merely later-in-time. Error in G2 or systematic bias in split gaps could inflate camera metrics (R-Pre/T-Pre) or content-expansion scores. A brief audit of camera-grid quality and gap statistics on the 1M set belongs in the main method or appendix and is needed to support the 3D-free scaling narrative.
minor comments (6)
- [Figure 2] Figure 2 caption compares to “SD-2.0” while the text also uses Seedance 2.0; keep naming consistent with the citation (Seedance et al. 2026) to avoid confusion with Stable Diffusion 2.0.
- [Method; Figure 4] Eq. (5)–(6): conditioning is written as vθ(zt,t|Vsrc,G,T) but architectural fusion of video, camera grid, and text (concatenation, cross-attention, channel stacking) is not specified beyond “MMDiTBlock” in Fig. 4. A short paragraph would aid reimplementation.
- [Related Work] Related Work should more clearly position concurrent camera-grid / re-shooting lines (including OmniDirector / Liu et al. 2026 cited for the grid) so readers can separate representation choice from the timestep-aware data contribution.
- [Table 1] Table 1: ArcFace for Ours (0.41) is slightly below SD-2.0 (0.43) while the text says “significantly outperforms existing re-shooting baselines… comparable to SD2.0”—wording is fine but flag the SD-2.0 identity edge explicitly.
- [Throughout] Typos and spacing artifacts from PDF extraction appear throughout (e.g., “Videore-shooting”, “unseenregions”, “cameramotion”); clean the camera-ready text.
- [Experimental Setups] First-person special case (“discard the camera trajectory and follow the first-person subject”) is important for Table 2’s perspective score; describe the training/inference rule more precisely.
Circularity Check
Empirical ML methods paper; no derivation reduces to its inputs by construction. Minor self-citation only for the camera-grid representation.
specific steps
-
self citation load bearing
[Method / Preliminary, Camera Grid; also Stage 1 data construction]
"Camera Grid is a video-format representation (Liu et al. 2026) of camera motion that encodes camera parameters as visual grid transformations within an empty 3D room. Given its universality and ease of injection into diffusion models, we adopt this camera representation in our method. ... Pairing the camera grid with the video enables self-supervised camera motion injection (Liu et al. 2026)."
The motion condition G is taken from prior work by overlapping authors rather than derived here. This is ordinary self-citation of a representation, not a uniqueness result that forbids alternatives or that algebraically forces R-Pre/T-Pre. It does not make the timestep-routing or SOTA claims true by construction; those rest on external baselines and ablations. Flagged only as minor self-reference.
full rationale
TARS is a standard conditional flow-matching video model with a two-stage data recipe (large-scale self-supervised clip splits + scarce cross-pair data restricted to high-noise t). The training objectives (Eqs. 4–6) are ordinary vector-field matching toward VAE(V_tgt); evaluation metrics (R-Pre, T-Pre, V-MPGE, ArcFace, Gemini-judged viewpoint/quality) are external and not algebraically fixed by the loss or by any fitted scalar. The timestep concentration claim is supported by injection ablations (Fig. 3; Tables 3–4 High vs Mid&Low) and stage ablations, not by defining the target as the fit. The only self-reference is adoption of the camera-grid encoding from Liu et al. 2026 (overlapping authors) as the motion condition G; that supplies a representation, not a uniqueness theorem or a quantity that forces the reported accuracy numbers. No self-definitional loop, no fitted-input-called-prediction, and no uniqueness/ansatz smuggling that collapses the central claim. Score 1 only for that non-load-bearing self-citation of the grid format.
Axiom & Free-Parameter Ledger
free parameters (4)
- high-noise regime cutoff t∈[0.95,1.0] =
t in [0.95, 1.0]
- self-supervised vs cross-pair data scale =
1M + 60K
- training hyperparameters =
8K steps, 5e-5
- viewpoint attribute taxonomy for text labels =
close-up/medium/long; front/side/back/high/low; OTS/1st/3rd
axioms (6)
- domain assumption Rectified-flow / conditional flow-matching training of a video diffusion transformer is a valid generative backbone for re-shooting.
- domain assumption High-noise denoising steps predominantly determine low-frequency structure and camera motion; mid/low-noise steps refine appearance.
- ad hoc to paper Temporally splitting a single monocular video into V1/V2 with estimated camera grid G2 yields useful self-supervision for camera and viewpoint change without true multi-view pairs.
- ad hoc to paper A small amount of cross-pair data in high noise is enough to elicit temporally synchronized subject motion via a ‘shortcut effect’ under CFG.
- domain assumption Camera Grid (projected ceiling/floor lattice) is a faithful, injectable encoding of target camera trajectory.
- domain assumption MLLM-generated viewpoint text (Qwen3-VL) is an adequate semantic condition for shot scale, angle, and perspective.
invented entities (2)
-
TARS timestep-aware two-stage data routing
no independent evidence
-
Text-driven semantic viewpoint specification (shot scale / angle / perspective attributes)
no independent evidence
read the original abstract
Video re-shooting aims to regenerate videos with controllable camera motion and viewpoint. Existing methods rely on explicit 3D priors, which are limited by reconstruction quality and often perform poorly when synthesizing previously unseen regions, or on paired videos with different camera trajectories, whose scarcity hinders generalization. We revisit video re-shooting through text-driven semantic viewpoint specification, enabling control over shot scale, viewing angle, and first-/third-person perspective. To this end, we propose TARS, a 3D-free video re-shooting paradigm. Timestep-wise sensitivity analysis reveals that camera motion is primarily established during high-noise stages, where coarse spatiotemporal structures are formed. Based on this insight, we introduce self-supervised training to learn camera dynamics and fundamental visual representations without paired re-shooting data or 3D reconstruction. Through data scaling and joint textual-camera conditioning, TARS supports robust camera and viewpoint control, plausibly synthesizing regions beyond the source view under large camera motions while enabling reverse-angle re-shooting and perspective switching. Extensive experiments show that TARS provides more accurate and temporally consistent camera control than prior methods. Project Page: https://ymlinfeng.github.io/TARS.github.io/
Figures
Reference graph
Works this paper leans on
-
[1]
2026 , eprint=
ARGUS: Stacked Multi-View Identity Mosaic Injection for Subject-Preserving Video Generation , author=. 2026 , eprint=
2026
-
[2]
Shen, Zehong and Pi, Huaijin and Xia, Yan and Cen, Zhi and Peng, Sida and Hu, Zechen and Bao, Hujun and Hu, Ruizhen and Zhou, Xiaowei , title =. 2024 , isbn =. doi:10.1145/3680528.3687565 , booktitle =
arXiv 2024
-
[3]
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , month =
Deng, Jiankang and Guo, Jia and Xue, Niannan and Zafeiriou, Stefanos , title =. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , month =
-
[4]
Lin, Haotong and Chen, Sili and Liew, Junhao and Chen, Donny Y and Li, Zhenyu and Shi, Guang and Feng, Jiashi and Kang, Bingyi , journal=
-
[5]
Ho, Jonathan and Chan, William and Saharia, Chitwan and Whang, Jay and Gao, Ruiqi and Gritsenko, Alexey and Kingma, Diederik P and Poole, Ben and Norouzi, Mohammad and Fleet, David J and others , journal=
-
[6]
Wang, Xiang and Yuan, Hangjie and Zhang, Shiwei and Chen, Dayou and Wang, Jiuniu and Zhang, Yingya and Shen, Yujun and Zhao, Deli and Zhou, Jingren , journal=
-
[7]
Blattmann, Andreas and Dockhorn, Tim and Kulal, Sumith and Mendelevitch, Daniel and Kilian, Maciej and Lorenz, Dominik and Levi, Yam and English, Zion and Voleti, Vikram and Letts, Adam and others , journal=
-
[8]
Zheng, Zangwei and Peng, Xiangyu and Yang, Tianji and Shen, Chenhui and Li, Shenggui and Liu, Hongxin and Zhou, Yukun and Li, Tianyi and You, Yang , journal=
-
[9]
Ma, Xin and Wang, Yaohui and Chen, Xinyuan and Jia, Gengyun and Liu, Ziwei and Li, Yuan-Fang and Chen, Cunjian and Qiao, Yu , journal=
-
[10]
ICASSP 2026-2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pages=
Make a Game: A Novel Paradigm for Interactive Game Rendering , author=. ICASSP 2026-2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pages=. 2026 , organization=
2026
-
[11]
Ma, Yue and He, Yingqing and Wang, Hongfa and Wang, Andong and Shen, Leqi and Qi, Chenyang and Ying, Jixuan and Cai, Chengfei and Li, Zhifeng and Shum, Heung-Yeung and others , booktitle=
-
[12]
Lin, Han and Zala, Abhay and Cho, Jaemin and Bansal, Mohit , booktitle=
-
[13]
Bar-Tal, Omer and Chefer, Hila and Tov, Omer and Herrmann, Charles and Paiss, Roni and Zada, Shiran and Ephrat, Ariel and Hur, Junhwa and Liu, Guanghui and Raj, Amit and others , booktitle=
-
[14]
Ren, Weiming and Yang, Huan and Zhang, Ge and Wei, Cong and Du, Xinrun and Huang, Wenhao and Chen, Wenhu , journal=
-
[15]
Chen, Xinyuan and Wang, Yaohui and Zhang, Lingjun and Zhuang, Shaobin and Ma, Xin and Yu, Jiashuo and Wang, Yali and Lin, Dahua and Qiao, Yu and Liu, Ziwei , booktitle=
-
[16]
2024 , organization=
Xing, Jinbo and Xia, Menghan and Zhang, Yong and Chen, Haoxin and Yu, Wangbo and Liu, Hanyuan and Liu, Gongye and Wang, Xintao and Shan, Ying and Wong, Tien-Tsin , booktitle=. 2024 , organization=
2024
-
[17]
Zhang, Shiwei and Wang, Jiayu and Zhang, Yingya and Zhao, Kang and Yuan, Hangjie and Qin, Zhiwu and Wang, Xiang and Zhao, Deli and Zhou, Jingren , journal=
-
[18]
Chen, Weifeng and Ji, Yatai and Wu, Jie and Wu, Hefeng and Xie, Pan and Li, Jiashi and Xia, Xin and Xiao, Xuefeng and Lin, Liang , journal=
-
[19]
Zhang, Yabo and Wei, Yuxiang and ZHANG, XIAOPENG and Zuo, Wangmeng and Tian, Qi and others , booktitle=
-
[20]
Mou, Chong and Wang, Xintao and Xie, Liangbin and Wu, Yanze and Zhang, Jian and Qi, Zhongang and Shan, Ying , booktitle=
-
[21]
Zhang, Lvmin and Rao, Anyi and Agrawala, Maneesh , booktitle=
-
[22]
Polyak, Adam and Zohar, Amit and Brown, Andrew and Tjandra, Andros and Sinha, Animesh and Lee, Ann and Vyas, Apoorv and Shi, Bowen and Ma, Chih-Yao and Chuang, Ching-Yao and others , journal=
-
[23]
Peebles, William and Xie, Saining , booktitle=
-
[24]
Forty-first international conference on machine learning , year=
Esser, Patrick and Kulal, Sumith and Blattmann, Andreas and Entezari, Rahim and M. Forty-first international conference on machine learning , year=
-
[25]
Tim Brooks and Bill Peebles and Connor Holmes and Will DePue and Yufei Guo and Li Jing and David Schnurr and Joe Taylor and Troy Luhman and Eric Luhman and Clarence Ng and Ricky Wang and Aditya Ramesh , year=
-
[26]
Luo, Yawen and Shi, Xiaoyu and Zhuang, Junhao and Chen, Yutian and Liu, Quande and Wang, Xintao and Wan, Pengfei and Xue, Tianfan , journal=
-
[27]
Wu, Xiaoxue and Gao, Bingjie and Qiao, Yu and Wang, Yaohui and Chen, Xinyuan , journal=
-
[28]
IEEE transactions on pattern analysis and machine intelligence , volume=
A generic camera model and calibration method for conventional, wide-angle, and fish-eye lenses , author=. IEEE transactions on pattern analysis and machine intelligence , volume=. 2006 , publisher=
2006
-
[29]
Villegas, R and Moraldo, H and Castro, S and Babaeizadeh, M and Zhang, H and Kunze, J and Kindermans, PJ and Saffar, MT and Erhan, D , booktitle=
-
[30]
Singer, Uriel and Polyak, Adam and Hayes, Thomas and Yin, Xi and An, Jie and Zhang, Songyang and Hu, Qiyuan and Yang, Harry and Ashual, Oron and Gafni, Oran and others , journal=
-
[31]
Ho, Jonathan and Salimans, Tim and Gritsenko, Alexey and Chan, William and Norouzi, Mohammad and Fleet, David J , journal=
-
[32]
HaCohen, Yoav and Brazowski, Benny and Chiprut, Nisan and Bitterman, Yaki and Kvochko, Andrew and Berkowitz, Avishai and Shalem, Daniel and Lifschitz, Daphna and Moshe, Dudu and Porat, Eitan and Richardson, Eitan and Guy Shiran and Itay Chachy and Jonathan Chetboun and Michael Finkelson and Michael Kupchick and Nir Zabari and Nitzan Guetta and Noa Kotler ...
-
[33]
Luo, Yawen and Shi, Xiaoyu and Bai, Jianhong and Xia, Menghan and Xue, Tianfan and Wang, Xintao and Wan, Pengfei and Zhang, Di and Gai, Kun , booktitle=
-
[34]
Bai, Jianhong and Xia, Menghan and Fu, Xiao and Wang, Xintao and Mu, Lianrui and Cao, Jinwen and Liu, Zuozhu and Hu, Haoji and Bai, Xiang and Wan, Pengfei and others , booktitle=
-
[35]
Ho, Jonathan and Jain, Ajay and Abbeel, Pieter , journal=
-
[36]
2026 International Conference on 3D Vision (3DV) , pages=
Keetha, Nikhil and M. 2026 International Conference on 3D Vision (3DV) , pages=. 2026 , organization=
2026
-
[37]
2025 , publisher=
Yu, Wangbo and Xing, Jinbo and Yuan, Li and Hu, Wenbo and Li, Xiaoyu and Huang, Zhipeng and Gao, Xiangjun and Wong, Tien-Tsin and Shan, Ying and Tian, Yonghong , journal=. 2025 , publisher=
2025
-
[38]
He, Hao and Yang, Ceyuan and Lin, Shanchuan and Xu, Yinghao and Wei, Meng and Gui, Liangke and Zhao, Qi and Wetzstein, Gordon and Jiang, Lu and Li, Hongsheng , booktitle=
-
[39]
Li, Xinyang and Lai, Zhangyu and Xu, Linning and Qu, Yansong and Cao, Liujuan and Zhang, Shengchuan and Dai, Bo and Ji, Rongrong , journal=
-
[40]
Bai, Shuai and Cai, Yuxuan and Chen, Ruizhe and Chen, Keqin and Chen, Xionghui and Cheng, Zesen and Deng, Lianghao and Ding, Wei and Gao, Chang and Ge, Chunjiang and others , journal=
-
[41]
Guo, Yuwei and Yang, Ceyuan and Rao, Anyi and Liang, Zhengyang and Wang, Yaohui and Qiao, Yu and Agrawala, Maneesh and Lin, Dahua and Dai, Bo , journal=
-
[42]
Kong, Weijie and Tian, Qi and Zhang, Zijian and Min, Rox and Dai, Zuozhuo and Zhou, Jin and Xiong, Jiangfeng and Li, Xin and Wu, Bo and Zhang, Jianwei and others , journal=
-
[43]
Zheng, Guangcong and Li, Teng and Jiang, Rui and Lu, Yehao and Wu, Tao and Li, Xi , journal=
-
[44]
Xu, Dejia and Nie, Weili and Liu, Chao and Liu, Sifei and Kautz, Jan and Wang, Zhangyang and Vahdat, Arash , journal=
-
[45]
2024 , organization=
Girdhar, Rohit and Singh, Mannat and Brown, Andrew and Duval, Quentin and Azadi, Samaneh and Rambhatla, Sai Saketh and Shah, Akbar and Yin, Xi and Parikh, Devi and Misra, Ishan , booktitle=. 2024 , organization=
2024
-
[46]
Chen, Haoxin and Zhang, Yong and Cun, Xiaodong and Xia, Menghan and Wang, Xintao and Weng, Chao and Shan, Ying , booktitle=
-
[47]
Yin, Shengming and Wu, Chenfei and Liang, Jian and Shi, Jie and Li, Houqiang and Ming, Gong and Duan, Nan , journal=
-
[48]
2024 , organization=
Zhao, Rui and Gu, Yuchao and Wu, Jay Zhangjie and Zhang, David Junhao and Liu, Jia-Wei and Wu, Weijia and Keppo, Jussi and Shou, Mike Zheng , booktitle=. 2024 , organization=
2024
-
[49]
Proceedings of the 32nd ACM International Conference on Multimedia , pages=
Soucek, Tom. Proceedings of the 32nd ACM International Conference on Multimedia , pages=
-
[50]
Seedance, Team and Chen, De and Chen, Liyang and Chen, Xin and Chen, Ying and Chen, Zhuo and Chen, Zhuowei and Cheng, Feng and Cheng, Tianheng and Cheng, Yufeng and others , journal=
-
[51]
Hu, Teng and Zhang, Jiangning and Yi, Ran and Wang, Yating and Huang, Hongrui and Weng, Jieyu and Wang, Yabiao and Ma, Lizhuang , journal=
-
[52]
Ling, Pengyang and Bu, Jiazi and Zhang, Pan and Dong, Xiaoyi and Zang, Yuhang and Wu, Tong and Chen, Huaian and Wang, Jiaqi and Jin, Yi , journal=
-
[53]
Bahmani, Sherwin and Skorokhodov, Ivan and Qian, Guocheng and Siarohin, Aliaksandr and Menapace, Willi and Tagliasacchi, Andrea and Lindell, David B and Tulyakov, Sergey , booktitle=
-
[54]
Wang, Zhouxia and Yuan, Ziyang and Wang, Xintao and Li, Yaowei and Chen, Tianshui and Xia, Menghan and Luo, Ping and Shan, Ying , booktitle=
-
[55]
Huang, Ziqi and He, Yinan and Yu, Jiashuo and Zhang, Fan and Si, Chenyang and Jiang, Yuming and Zhang, Yuanhan and Wu, Tianxing and Jin, Qingyang and Chanpaisit, Nattapol and others , booktitle=
-
[56]
2025 , publisher=
Huang, Ziqi and Zhang, Fan and Xu, Xiaojie and He, Yinan and Yu, Jiashuo and Dong, Ziyue and Ma, Qianli and Chanpaisit, Nattapol and Si, Chenyang and Jiang, Yuming and others , journal=. 2025 , publisher=
2025
-
[57]
He, Hao and Xu, Yinghao and Guo, Yuwei and Wetzstein, Gordon and Dai, Bo and Li, Hongsheng and Yang, Ceyuan , journal=
-
[58]
Wang, Qinghe and Shi, Xiaoyu and Li, Baolu and Bian, Weikang and Liu, Quande and Lu, Huchuan and Wang, Xintao and Wan, Pengfei and Gai, Kun and Jia, Xu , booktitle=
-
[59]
Lin, Kuan Heng and Liu, Zhizheng and Salamanca, Pablo and Kant, Yash and Burgert, Ryan and Xu, Yuancheng and Namekata, Koichi and Zhao, Yiwei and Zhou, Bolei and Goldblum, Micah and others , booktitle=
-
[60]
Liu, Jiwen and Li, Shujuan and Fang, Zhixue and Li, Xiaohan and Zhou, Yan and Meng, Zijie and Zhang, Zhimin and Luo, Yawen and Zhang, Guoxin and Liu, Yu-Shen and Wan, Pengfei , journal=
-
[61]
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , month =
Wang, Jianyuan and Chen, Minghao and Karaev, Nikita and Vedaldi, Andrea and Rupprecht, Christian and Novotny, David , title =. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , month =. 2025 , pages =
2025
-
[62]
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , month =
Li, Zhengqi and Tucker, Richard and Cole, Forrester and Wang, Qianqian and Jin, Linyi and Ye, Vickie and Kanazawa, Angjoo and Holynski, Aleksander and Snavely, Noah , title =. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , month =. 2025 , pages =
2025
-
[63]
Zhang, Junyi and Herrmann, Charles and Hur, Junhwa and Jampani, Varun and darrell, trevor and Cole, Forrester and Sun, Deqing and Yang, Ming-Hsuan , booktitle =
-
[64]
Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) , month =
Lin, Shanchuan and Liu, Bingchen and Li, Jiashi and Yang, Xiao , title =. Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) , month =. 2024 , pages =
2024
-
[65]
and Ben-Hamu, Heli and Nickel, Maximilian and Le, Matt , year =
Lipman, Yaron and Chen, Ricky T.Q. and Ben-Hamu, Heli and Nickel, Maximilian and Le, Matt , year =
-
[66]
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , month =
Chen, Haoxin and Zhang, Yong and Cun, Xiaodong and Xia, Menghan and Wang, Xintao and Weng, Chao and Shan, Ying , title =. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , month =. 2024 , pages =
2024
-
[67]
Chun-Han Yao and Yiming Xie and Vikram Voleti and Huaizu Jiang and Varun Jampani , journal=
-
[68]
Bahmani, Sherwin and Skorokhodov, Ivan and Siarohin, Aliaksandr and Menapace, Willi and Qian, Guocheng and Vasilkovsky, Michael and Lee, Hsin-Ying and Wang, Chaoyang and Zou, Jiaxu and Tagliasacchi, Andrea and Lindell, David and Tulyakov, Sergey , booktitle =
-
[69]
European Conference on Computer Vision (ECCV) , year=
Generative Camera Dolly: Extreme Monocular Dynamic Novel View Synthesis , author=. European Conference on Computer Vision (ECCV) , year=
-
[70]
Wang, Yifan and Zhou, Jianjun and Zhu, Haoyi and Chang, Wenzheng and Zhou, Yang and Li, Zizun and Chen, Junyi and Pang, Jiangmiao and Shen, Chunhua and He, Tong , journal=
-
[71]
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , month =
Ren, Xuanchi and Shen, Tianchang and Huang, Jiahui and Ling, Huan and Lu, Yifan and Nimier-David, Merlin and M\"uller, Thomas and Keller, Alexander and Fidler, Sanja and Gao, Jun , title =. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , month =. 2025 , pages =
2025
-
[72]
Ma, Yue and Feng, Kunyu and Hu, Zhongyuan and Wang, Xinyu and Wang, Yucheng and Zheng, Mingzhe and He, Xuanhua and Zhu, Chenyang and Liu, Hongyu and He, Yingqing and others , journal=
-
[73]
Si, Chenyang and Huang, Ziqi and Jiang, Yuming and Liu, Ziwei , booktitle=
-
[74]
Balaji, Yogesh and Nah, Seungjun and Huang, Xun and Vahdat, Arash and Song, Jiaming and Zhang, Qinsheng and Kreis, Karsten and Aittala, Miika and Aila, Timo and Laine, Samuli and others , journal=
-
[75]
Yu, Mark and Hu, Wenbo and Xing, Jinbo and Shan, Ying , booktitle=
-
[76]
Wang, Qinghe and Luo, Yawen and Shi, Xiaoyu and Jia, Xu and Lu, Huchuan and Xue, Tianfan and Wang, Xintao and Wan, Pengfei and Zhang, Di and Gai, Kun , booktitle=
-
[77]
Bahmani, S.; Skorokhodov, I.; Siarohin, A.; Menapace, W.; Qian, G.; Vasilkovsky, M.; Lee, H.-Y.; Wang, C.; Zou, J.; Tagliasacchi, A.; Lindell, D.; and Tulyakov, S. 2025. VD3D: Taming Large Video Diffusion Transformers for 3D Camera Control . In Yue, Y.; Garg, A.; Peng, N.; Sha, F.; and Yu, R., eds., International Conference on Learning Representations, 66...
2025
-
[78]
Bai, J.; Xia, M.; Fu, X.; Wang, X.; Mu, L.; Cao, J.; Liu, Z.; Hu, H.; Bai, X.; Wan, P.; et al. 2025 a . ReCamMaster: Camera-Controlled Generative Rendering from A Single Video . In Proceedings of the IEEE/CVF International Conference on Computer Vision, 14834--14844
2025
-
[79]
Bai, S.; Cai, Y.; Chen, R.; Chen, K.; Chen, X.; Cheng, Z.; Deng, L.; Ding, W.; Gao, C.; Ge, C.; et al. 2025 b . Qwen3-VL Technical Report . arXiv preprint arXiv:2511.21631
Pith/arXiv arXiv 2025
-
[80]
Balaji, Y.; Nah, S.; Huang, X.; Vahdat, A.; Song, J.; Zhang, Q.; Kreis, K.; Aittala, M.; Aila, T.; Laine, S.; et al. 2022. eDiff-I: Text-to-Image Diffusion Models with An Ensemble of Expert Denoisers . arXiv preprint arXiv:2211.01324
Pith/arXiv arXiv 2022
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.