Pith. sign in

REVIEW 4 major objections 6 minor 121 references

Camera motion in video diffusion is set in the high-noise stage, so large self-supervised clip pairs plus a little paired data only there can re-shoot videos without 3D.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-31 12:58 UTC pith:PRVKRQBQ

load-bearing objection Solid engineering recipe for 3D-free re-shooting: timestep-routed self-supervision plus a little high-noise pair data actually moves the needle, even if the routing story is not fully isolated from data scale. the 4 major comments →

arxiv 2607.28261 v1 pith:PRVKRQBQ submitted 2026-07-30 cs.CV

TARS: Timestep-Aware Data Scaling for 3D-Free Video Re-Shooting

classification cs.CV
keywords video re-shootingcamera controldiffusion timestepsself-supervised learningviewpoint control3D-free generationcamera gridtext-driven perspective
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Video re-shooting means regenerating a clip under a new camera path and viewpoint while keeping subjects, actions, and scene content. Prior approaches either rebuild the scene in 3D and re-render it—brittle when geometry fails or unseen regions appear—or train on scarce paired videos shot from different trajectories, which limits generalization and semantic control. TARS argues that diffusion models lock in coarse structure and camera motion early, in the high-noise timesteps, and only refine appearance later. That split lets the authors train mostly on self-supervised pairs cut from ordinary unlabeled videos (plus text that names shot scale, angle, and first-/third-person view), and route only a small amount of true cross-pair data into the high-noise band to lock subject motion in sync. The result is 3D-free re-shooting that follows target trajectories more accurately, switches perspective on command, and invents plausible content outside the original view.

Core claim

Timestep-wise analysis shows camera motion and coarse spatiotemporal structure form mainly in high-noise denoising stages; therefore a 3D-free two-stage recipe—large-scale self-supervised clip splitting for camera, viewpoint, and appearance, plus minimal cross-pair supervision only in the high-noise regime—yields more accurate, temporally consistent re-shooting and text-driven control of shot scale, viewing angle, and perspective than 3D-prior or fully paired baselines.

What carries the argument

Timestep-aware data routing: self-supervised pairs from temporally split unlabeled clips train all timesteps for structure, camera grids, and text viewpoint; scarce cross-pair data is applied only in the high-noise band (empirically t in [0.95, 1.0]) so a shortcut under guidance elicits motion-synced subject dynamics.

Load-bearing premise

Camera motion and subject structure are concentrated enough in a narrow high-noise band that a little paired supervision there alone can force synchronized motion without true multi-view pairs at later timesteps.

What would settle it

Retrain or ablate with cross-pair data barred from high noise and allowed only mid/low noise (or widen the high-noise band): if camera accuracy and motion sync (R-Pre, T-Pre, V-MPGE) then collapse relative to the reported full model, the routing claim fails; if they hold, the concentration premise is wrong.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Re-shooting no longer needs full 3D/4D reconstruction or large multi-camera paired corpora for competitive trajectory control.
  • Text can specify shot scale, viewing angle, and first-/third-person perspective jointly with a camera grid, including reverse-angle and large-motion views.
  • Unlabeled video scale becomes the main lever for viewpoint generalization and plausible synthesis of regions outside the source frame.
  • The same high-noise vs mid/low split can guide where expensive paired supervision is spent in other controllable video tasks.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If high-noise stages truly own global geometry, other sparse geometric controls (depth, pose, layout) may also train cheaply with self-supervision everywhere and paired data only early.
  • The reported shortcut under classifier-free guidance suggests many video editors could bootstrap temporal sync from tiny paired sets once strong unpaired priors exist.
  • Failure modes on extreme open-domain dynamics would show up first as motion desync rather than identity drift, pointing where to add the next paired data.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. TARS proposes a 3D-free video re-shooting framework that combines camera-grid conditioning with text-driven semantic viewpoint control (shot scale, viewing angle, first-/third-person perspective). Motivated by a timestep-wise analysis arguing that camera motion and coarse structure form mainly in high-noise stages, the method uses a two-stage data strategy: large-scale self-supervised pairs obtained by temporally splitting monocular videos (with estimated camera grids and MLLM viewpoint captions), plus a small amount of cross-pair data applied only in a narrow high-noise band (t∈[0.95,1.0]) to elicit motion synchronization via an asserted “shortcut effect” under CFG. Experiments on a 1021-sample set report gains over CamClone, TrajCrafter, and SD-2.0 on camera accuracy, spatio-temporal consistency, content expansion, and viewpoint metrics, with ablations of Stage 1/2 and high vs mid/low source injection.

Significance. If the claims hold, the work is a practically important contribution to controllable video generation: it attacks the paired-data bottleneck for re-shooting without 3D reconstruction, unifies geometric camera control with semantic viewpoint instructions, and shows plausible synthesis under large motions including reverse-angle and perspective switches. The timestep-aware data-routing idea is a useful engineering principle for diffusion training under scarce multi-view supervision. Strengths include a balanced multi-category evaluation set, quantitative comparison against three relevant baselines (Tables 1–2), and ablations that partially support the high-noise vs mid/low division (Tables 3–4, Fig. 3). The result would matter for cinematography tools and camera-controllable video models more broadly, provided the routing claim is isolated from raw data scale and the evaluation is made more reproducible.

major comments (4)
  1. [Method (Timestep-Aware Data Routing); Tables 3–4] Method, Stage 2 and Experimental Setups; Tables 3–4: The central claim that scarce cross-pair supervision need only be applied in t∈[0.95,1.0] (via a “shortcut effect”) is not isolated from data scale. Ablations remove Stage 1, remove Stage 2, or change where the source video is injected at inference, but there is no control that trains the same 1M self-supervised + 60K cross-pair mixture with the cross-pair objective on all timesteps (or on a wider high-noise band). Without that comparison, Table 1 gains cannot be attributed to timestep-aware routing rather than simply more and more diverse training data. This control is load-bearing for the paper’s title and abstract claim.
  2. [Experimental Setups; Method Stage 2; Figure 3] Experimental Setups defines the high-noise regime as t∈[0.95,1.0] empirically, and Stage 2 asserts that synchronized motion is “easier to optimize” under CFG without a counter-example or quantitative probe (e.g., multi-modal motion pairs where unsynced completions are equally valid, or sweeps of the cutoff). Fig. 3 and the High vs Mid&Low rows support a coarse frequency split but do not establish that a 5% noise band is necessary or sufficient for open-domain dynamics. A cutoff sensitivity study and at least one failure-mode analysis of the shortcut would substantially strengthen the mechanism.
  3. [Experiments (Evaluation Metrics; Comparisons)] Comparisons and metrics: The backbone is an unspecified in-house T2V model, while several primary metrics (viewpoint accuracy, CE, FDR, VDR, FSCS, and parts of consistency) are judged by Gemini 3.1 Pro. This combination weakens external validity and reproducibility of the SOTA claim in Table 1–2. Please (i) clarify backbone capacity/training relative to baselines or release comparable checkpoints/protocol, (ii) report inter-judge agreement or human studies on a subset for LLM-judged metrics, and (iii) add standard low-level video metrics where applicable so that camera and identity gains are not solely model-judged.
  4. [Method Stage 1; Eq. (6)] Self-supervised construction (Eq. 6): V1/V2 are temporal halves of one monocular video with G2 estimated from V2. Large-gap non-overlapping splits are said to teach hallucination of unseen regions, but the paper does not quantify pose-estimation error, temporal gap distributions, or how often “unseen” content is truly out-of-view versus merely later-in-time. Error in G2 or systematic bias in split gaps could inflate camera metrics (R-Pre/T-Pre) or content-expansion scores. A brief audit of camera-grid quality and gap statistics on the 1M set belongs in the main method or appendix and is needed to support the 3D-free scaling narrative.
minor comments (6)
  1. [Figure 2] Figure 2 caption compares to “SD-2.0” while the text also uses Seedance 2.0; keep naming consistent with the citation (Seedance et al. 2026) to avoid confusion with Stable Diffusion 2.0.
  2. [Method; Figure 4] Eq. (5)–(6): conditioning is written as vθ(zt,t|Vsrc,G,T) but architectural fusion of video, camera grid, and text (concatenation, cross-attention, channel stacking) is not specified beyond “MMDiTBlock” in Fig. 4. A short paragraph would aid reimplementation.
  3. [Related Work] Related Work should more clearly position concurrent camera-grid / re-shooting lines (including OmniDirector / Liu et al. 2026 cited for the grid) so readers can separate representation choice from the timestep-aware data contribution.
  4. [Table 1] Table 1: ArcFace for Ours (0.41) is slightly below SD-2.0 (0.43) while the text says “significantly outperforms existing re-shooting baselines… comparable to SD2.0”—wording is fine but flag the SD-2.0 identity edge explicitly.
  5. [Throughout] Typos and spacing artifacts from PDF extraction appear throughout (e.g., “Videore-shooting”, “unseenregions”, “cameramotion”); clean the camera-ready text.
  6. [Experimental Setups] First-person special case (“discard the camera trajectory and follow the first-person subject”) is important for Table 2’s perspective score; describe the training/inference rule more precisely.

Circularity Check

1 steps flagged

Empirical ML methods paper; no derivation reduces to its inputs by construction. Minor self-citation only for the camera-grid representation.

specific steps
  1. self citation load bearing [Method / Preliminary, Camera Grid; also Stage 1 data construction]
    "Camera Grid is a video-format representation (Liu et al. 2026) of camera motion that encodes camera parameters as visual grid transformations within an empty 3D room. Given its universality and ease of injection into diffusion models, we adopt this camera representation in our method. ... Pairing the camera grid with the video enables self-supervised camera motion injection (Liu et al. 2026)."

    The motion condition G is taken from prior work by overlapping authors rather than derived here. This is ordinary self-citation of a representation, not a uniqueness result that forbids alternatives or that algebraically forces R-Pre/T-Pre. It does not make the timestep-routing or SOTA claims true by construction; those rest on external baselines and ablations. Flagged only as minor self-reference.

full rationale

TARS is a standard conditional flow-matching video model with a two-stage data recipe (large-scale self-supervised clip splits + scarce cross-pair data restricted to high-noise t). The training objectives (Eqs. 4–6) are ordinary vector-field matching toward VAE(V_tgt); evaluation metrics (R-Pre, T-Pre, V-MPGE, ArcFace, Gemini-judged viewpoint/quality) are external and not algebraically fixed by the loss or by any fitted scalar. The timestep concentration claim is supported by injection ablations (Fig. 3; Tables 3–4 High vs Mid&Low) and stage ablations, not by defining the target as the fit. The only self-reference is adoption of the camera-grid encoding from Liu et al. 2026 (overlapping authors) as the motion condition G; that supplies a representation, not a uniqueness theorem or a quantity that forces the reported accuracy numbers. No self-definitional loop, no fitted-input-called-prediction, and no uniqueness/ansatz smuggling that collapses the central claim. Score 1 only for that non-load-bearing self-citation of the grid format.

Axiom & Free-Parameter Ledger

4 free parameters · 6 axioms · 2 invented entities

The central claim rests on standard diffusion/flow-matching machinery, a borrowed camera-grid representation, an empirical coarse-to-fine timestep hypothesis, and large private data plus an in-house backbone. Free knobs (noise cutoff, data mix, LR) and the unproven ‘shortcut effect’ are load-bearing. No new physical entities; the invented pieces are methodological constructs.

free parameters (4)
  • high-noise regime cutoff t∈[0.95,1.0] = t in [0.95, 1.0]
    Empirically chosen band where cross-pair loss is applied; defines the entire data-routing claim. No sweep reported.
  • self-supervised vs cross-pair data scale = 1M + 60K
    1M unlabeled clips + 60K cross-pairs (10K real, 50K UE) chosen by authors; performance depends on this mix.
  • training hyperparameters = 8K steps, 5e-5
    8K iterations at LR 5e-5 on in-house backbone; affect final metrics but are standard fit knobs.
  • viewpoint attribute taxonomy for text labels = close-up/medium/long; front/side/back/high/low; OTS/1st/3rd
    Shot scale / viewing angle / perspective categories used to prompt Qwen3-VL; shapes what ‘semantic viewpoint control’ means.
axioms (6)
  • domain assumption Rectified-flow / conditional flow-matching training of a video diffusion transformer is a valid generative backbone for re-shooting.
    Preliminary and Eq. 4–5; standard in recent video generators, not re-derived.
  • domain assumption High-noise denoising steps predominantly determine low-frequency structure and camera motion; mid/low-noise steps refine appearance.
    Motivated by eDiff-I/FreeU and authors’ injection study (Fig. 3–4); treated as design law for data routing.
  • ad hoc to paper Temporally splitting a single monocular video into V1/V2 with estimated camera grid G2 yields useful self-supervision for camera and viewpoint change without true multi-view pairs.
    Stage 1 construction; core scalability claim. Overlapping vs non-overlapping splits asserted to help hallucination.
  • ad hoc to paper A small amount of cross-pair data in high noise is enough to elicit temporally synchronized subject motion via a ‘shortcut effect’ under CFG.
    Stage 2 paragraph; explains why 60K pairs suffice. Not independently proven outside this training setup.
  • domain assumption Camera Grid (projected ceiling/floor lattice) is a faithful, injectable encoding of target camera trajectory.
    Adopted from Liu et al. 2026; Eq. 1 and Preliminary.
  • domain assumption MLLM-generated viewpoint text (Qwen3-VL) is an adequate semantic condition for shot scale, angle, and perspective.
    Stage 1 text construction; evaluation also uses Gemini as judge.
invented entities (2)
  • TARS timestep-aware two-stage data routing no independent evidence
    purpose: Allocate self-supervised split data to most timesteps and scarce cross-pairs only to high-noise steps for re-shooting.
    Methodological construct, not a physical entity; existence is the training procedure itself.
  • Text-driven semantic viewpoint specification (shot scale / angle / perspective attributes) no independent evidence
    purpose: Control re-shoot viewpoint without explicit target extrinsics beyond camera grid.
    Interface abstraction built from MLLM labels; validated only inside this paper’s Gemini metrics.

pith-pipeline@v1.2.0-daily-grok45 · 18265 in / 4130 out tokens · 84477 ms · 2026-07-31T12:58:24.337416+00:00 · methodology

0 comments
read the original abstract

Video re-shooting aims to regenerate videos with controllable camera motion and viewpoint. Existing methods rely on explicit 3D priors, which are limited by reconstruction quality and often perform poorly when synthesizing previously unseen regions, or on paired videos with different camera trajectories, whose scarcity hinders generalization. We revisit video re-shooting through text-driven semantic viewpoint specification, enabling control over shot scale, viewing angle, and first-/third-person perspective. To this end, we propose TARS, a 3D-free video re-shooting paradigm. Timestep-wise sensitivity analysis reveals that camera motion is primarily established during high-noise stages, where coarse spatiotemporal structures are formed. Based on this insight, we introduce self-supervised training to learn camera dynamics and fundamental visual representations without paired re-shooting data or 3D reconstruction. Through data scaling and joint textual-camera conditioning, TARS supports robust camera and viewpoint control, plausibly synthesizing regions beyond the source view under large camera motions while enabling reverse-angle re-shooting and perspective switching. Extensive experiments show that TARS provides more accurate and temporally consistent camera control than prior methods. Project Page: https://ymlinfeng.github.io/TARS.github.io/

Figures

Figures reproduced from arXiv: 2607.28261 by Guoxin Zhang, Jiwen Liu, Shujuan Li, Xiaohan Li, Xinyue Liu, Yan Zhou, Yulong Xu, Zijie Meng.

Figure 1
Figure 1. Figure 1: TARS enables robust re-shooting and viewpoint control, plausibly synthesizing unseen regions under large camera [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 3
Figure 3. Figure 3: Visualization of source video injection at differ￾ent diffusion timesteps. High-noise timesteps capture low￾frequency structure and motion, while rest timesteps preserve high-frequency details such as subject identity. high-noise denoising stages play a dominant role in estab￾lishing coarse spatiotemporal structures, particularly camera motion and subject dynamics, whereas later stages mainly refine appear… view at source ↗
Figure 2
Figure 2. Figure 2: Comparison with SD-2.0 in viewpoint control. Our method enables accurate transition between the third￾person view and the first-person view. rectly from paired videos that capture the same scene with different camera trajectories. Prior-based methods (Lin et al. 2026; Yu et al. 2025; Ren et al. 2025) typically reconstruct the source video into explicit 3D/4D representations (Wang et al. 2025a; Lin et al. 2… view at source ↗
Figure 4
Figure 4. Figure 4: Timestep-aware self-supervised learning framework. Top: High-noise diffusion timesteps primarily learn global structure, viewpoint, and camera motion, while mid- and low-noise timesteps refine texture details. Bottom: Guided by this observation, we construct self-supervised training pairs by temporally splitting each video into two clips, enabling large-scale learning of camera transitions from unlabeled v… view at source ↗
Figure 5
Figure 5. Figure 5: Qualitative Evaluations. The results demonstrate that TARS can accurately reshoot the source video under novel camera trajectories and viewpoints. In real-world data collection, strictly paired data capturing the same scene with different camera trajectories (i.e., cross￾pair data) is extremely scarce, whereas data sharing identical high-frequency information (e.g., appearance and texture) is relatively ab… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

121 extracted references · 15 linked inside Pith

  1. [1]

    2026 , eprint=

    ARGUS: Stacked Multi-View Identity Mosaic Injection for Subject-Preserving Video Generation , author=. 2026 , eprint=

  2. [2]

    2024 , isbn =

    Shen, Zehong and Pi, Huaijin and Xia, Yan and Cen, Zhi and Peng, Sida and Hu, Zechen and Bao, Hujun and Hu, Ruizhen and Zhou, Xiaowei , title =. 2024 , isbn =. doi:10.1145/3680528.3687565 , booktitle =

  3. [3]

    Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , month =

    Deng, Jiankang and Guo, Jia and Xue, Niannan and Zafeiriou, Stefanos , title =. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , month =

  4. [4]

    Lin, Haotong and Chen, Sili and Liew, Junhao and Chen, Donny Y and Li, Zhenyu and Shi, Guang and Feng, Jiashi and Kang, Bingyi , journal=

  5. [5]

    Ho, Jonathan and Chan, William and Saharia, Chitwan and Whang, Jay and Gao, Ruiqi and Gritsenko, Alexey and Kingma, Diederik P and Poole, Ben and Norouzi, Mohammad and Fleet, David J and others , journal=

  6. [6]

    Wang, Xiang and Yuan, Hangjie and Zhang, Shiwei and Chen, Dayou and Wang, Jiuniu and Zhang, Yingya and Shen, Yujun and Zhao, Deli and Zhou, Jingren , journal=

  7. [7]

    Blattmann, Andreas and Dockhorn, Tim and Kulal, Sumith and Mendelevitch, Daniel and Kilian, Maciej and Lorenz, Dominik and Levi, Yam and English, Zion and Voleti, Vikram and Letts, Adam and others , journal=

  8. [8]

    Zheng, Zangwei and Peng, Xiangyu and Yang, Tianji and Shen, Chenhui and Li, Shenggui and Liu, Hongxin and Zhou, Yukun and Li, Tianyi and You, Yang , journal=

  9. [9]

    Ma, Xin and Wang, Yaohui and Chen, Xinyuan and Jia, Gengyun and Liu, Ziwei and Li, Yuan-Fang and Chen, Cunjian and Qiao, Yu , journal=

  10. [10]

    ICASSP 2026-2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pages=

    Make a Game: A Novel Paradigm for Interactive Game Rendering , author=. ICASSP 2026-2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pages=. 2026 , organization=

  11. [11]

    Ma, Yue and He, Yingqing and Wang, Hongfa and Wang, Andong and Shen, Leqi and Qi, Chenyang and Ying, Jixuan and Cai, Chengfei and Li, Zhifeng and Shum, Heung-Yeung and others , booktitle=

  12. [12]

    Lin, Han and Zala, Abhay and Cho, Jaemin and Bansal, Mohit , booktitle=

  13. [13]

    Bar-Tal, Omer and Chefer, Hila and Tov, Omer and Herrmann, Charles and Paiss, Roni and Zada, Shiran and Ephrat, Ariel and Hur, Junhwa and Liu, Guanghui and Raj, Amit and others , booktitle=

  14. [14]

    Ren, Weiming and Yang, Huan and Zhang, Ge and Wei, Cong and Du, Xinrun and Huang, Wenhao and Chen, Wenhu , journal=

  15. [15]

    Chen, Xinyuan and Wang, Yaohui and Zhang, Lingjun and Zhuang, Shaobin and Ma, Xin and Yu, Jiashuo and Wang, Yali and Lin, Dahua and Qiao, Yu and Liu, Ziwei , booktitle=

  16. [16]

    2024 , organization=

    Xing, Jinbo and Xia, Menghan and Zhang, Yong and Chen, Haoxin and Yu, Wangbo and Liu, Hanyuan and Liu, Gongye and Wang, Xintao and Shan, Ying and Wong, Tien-Tsin , booktitle=. 2024 , organization=

  17. [17]

    Zhang, Shiwei and Wang, Jiayu and Zhang, Yingya and Zhao, Kang and Yuan, Hangjie and Qin, Zhiwu and Wang, Xiang and Zhao, Deli and Zhou, Jingren , journal=

  18. [18]

    Chen, Weifeng and Ji, Yatai and Wu, Jie and Wu, Hefeng and Xie, Pan and Li, Jiashi and Xia, Xin and Xiao, Xuefeng and Lin, Liang , journal=

  19. [19]

    Zhang, Yabo and Wei, Yuxiang and ZHANG, XIAOPENG and Zuo, Wangmeng and Tian, Qi and others , booktitle=

  20. [20]

    Mou, Chong and Wang, Xintao and Xie, Liangbin and Wu, Yanze and Zhang, Jian and Qi, Zhongang and Shan, Ying , booktitle=

  21. [21]

    Zhang, Lvmin and Rao, Anyi and Agrawala, Maneesh , booktitle=

  22. [22]

    Polyak, Adam and Zohar, Amit and Brown, Andrew and Tjandra, Andros and Sinha, Animesh and Lee, Ann and Vyas, Apoorv and Shi, Bowen and Ma, Chih-Yao and Chuang, Ching-Yao and others , journal=

  23. [23]

    Peebles, William and Xie, Saining , booktitle=

  24. [24]

    Forty-first international conference on machine learning , year=

    Esser, Patrick and Kulal, Sumith and Blattmann, Andreas and Entezari, Rahim and M. Forty-first international conference on machine learning , year=

  25. [25]

    Tim Brooks and Bill Peebles and Connor Holmes and Will DePue and Yufei Guo and Li Jing and David Schnurr and Joe Taylor and Troy Luhman and Eric Luhman and Clarence Ng and Ricky Wang and Aditya Ramesh , year=

  26. [26]

    Luo, Yawen and Shi, Xiaoyu and Zhuang, Junhao and Chen, Yutian and Liu, Quande and Wang, Xintao and Wan, Pengfei and Xue, Tianfan , journal=

  27. [27]

    Wu, Xiaoxue and Gao, Bingjie and Qiao, Yu and Wang, Yaohui and Chen, Xinyuan , journal=

  28. [28]

    IEEE transactions on pattern analysis and machine intelligence , volume=

    A generic camera model and calibration method for conventional, wide-angle, and fish-eye lenses , author=. IEEE transactions on pattern analysis and machine intelligence , volume=. 2006 , publisher=

  29. [29]

    Villegas, R and Moraldo, H and Castro, S and Babaeizadeh, M and Zhang, H and Kunze, J and Kindermans, PJ and Saffar, MT and Erhan, D , booktitle=

  30. [30]

    Singer, Uriel and Polyak, Adam and Hayes, Thomas and Yin, Xi and An, Jie and Zhang, Songyang and Hu, Qiyuan and Yang, Harry and Ashual, Oron and Gafni, Oran and others , journal=

  31. [31]

    Ho, Jonathan and Salimans, Tim and Gritsenko, Alexey and Chan, William and Norouzi, Mohammad and Fleet, David J , journal=

  32. [32]

    HaCohen, Yoav and Brazowski, Benny and Chiprut, Nisan and Bitterman, Yaki and Kvochko, Andrew and Berkowitz, Avishai and Shalem, Daniel and Lifschitz, Daphna and Moshe, Dudu and Porat, Eitan and Richardson, Eitan and Guy Shiran and Itay Chachy and Jonathan Chetboun and Michael Finkelson and Michael Kupchick and Nir Zabari and Nitzan Guetta and Noa Kotler ...

  33. [33]

    Luo, Yawen and Shi, Xiaoyu and Bai, Jianhong and Xia, Menghan and Xue, Tianfan and Wang, Xintao and Wan, Pengfei and Zhang, Di and Gai, Kun , booktitle=

  34. [34]

    Bai, Jianhong and Xia, Menghan and Fu, Xiao and Wang, Xintao and Mu, Lianrui and Cao, Jinwen and Liu, Zuozhu and Hu, Haoji and Bai, Xiang and Wan, Pengfei and others , booktitle=

  35. [35]

    Ho, Jonathan and Jain, Ajay and Abbeel, Pieter , journal=

  36. [36]

    2026 International Conference on 3D Vision (3DV) , pages=

    Keetha, Nikhil and M. 2026 International Conference on 3D Vision (3DV) , pages=. 2026 , organization=

  37. [37]

    2025 , publisher=

    Yu, Wangbo and Xing, Jinbo and Yuan, Li and Hu, Wenbo and Li, Xiaoyu and Huang, Zhipeng and Gao, Xiangjun and Wong, Tien-Tsin and Shan, Ying and Tian, Yonghong , journal=. 2025 , publisher=

  38. [38]

    He, Hao and Yang, Ceyuan and Lin, Shanchuan and Xu, Yinghao and Wei, Meng and Gui, Liangke and Zhao, Qi and Wetzstein, Gordon and Jiang, Lu and Li, Hongsheng , booktitle=

  39. [39]

    Li, Xinyang and Lai, Zhangyu and Xu, Linning and Qu, Yansong and Cao, Liujuan and Zhang, Shengchuan and Dai, Bo and Ji, Rongrong , journal=

  40. [40]

    Bai, Shuai and Cai, Yuxuan and Chen, Ruizhe and Chen, Keqin and Chen, Xionghui and Cheng, Zesen and Deng, Lianghao and Ding, Wei and Gao, Chang and Ge, Chunjiang and others , journal=

  41. [41]

    Guo, Yuwei and Yang, Ceyuan and Rao, Anyi and Liang, Zhengyang and Wang, Yaohui and Qiao, Yu and Agrawala, Maneesh and Lin, Dahua and Dai, Bo , journal=

  42. [42]

    Kong, Weijie and Tian, Qi and Zhang, Zijian and Min, Rox and Dai, Zuozhuo and Zhou, Jin and Xiong, Jiangfeng and Li, Xin and Wu, Bo and Zhang, Jianwei and others , journal=

  43. [43]

    Zheng, Guangcong and Li, Teng and Jiang, Rui and Lu, Yehao and Wu, Tao and Li, Xi , journal=

  44. [44]

    Xu, Dejia and Nie, Weili and Liu, Chao and Liu, Sifei and Kautz, Jan and Wang, Zhangyang and Vahdat, Arash , journal=

  45. [45]

    2024 , organization=

    Girdhar, Rohit and Singh, Mannat and Brown, Andrew and Duval, Quentin and Azadi, Samaneh and Rambhatla, Sai Saketh and Shah, Akbar and Yin, Xi and Parikh, Devi and Misra, Ishan , booktitle=. 2024 , organization=

  46. [46]

    Chen, Haoxin and Zhang, Yong and Cun, Xiaodong and Xia, Menghan and Wang, Xintao and Weng, Chao and Shan, Ying , booktitle=

  47. [47]

    Yin, Shengming and Wu, Chenfei and Liang, Jian and Shi, Jie and Li, Houqiang and Ming, Gong and Duan, Nan , journal=

  48. [48]

    2024 , organization=

    Zhao, Rui and Gu, Yuchao and Wu, Jay Zhangjie and Zhang, David Junhao and Liu, Jia-Wei and Wu, Weijia and Keppo, Jussi and Shou, Mike Zheng , booktitle=. 2024 , organization=

  49. [49]

    Proceedings of the 32nd ACM International Conference on Multimedia , pages=

    Soucek, Tom. Proceedings of the 32nd ACM International Conference on Multimedia , pages=

  50. [50]

    Seedance, Team and Chen, De and Chen, Liyang and Chen, Xin and Chen, Ying and Chen, Zhuo and Chen, Zhuowei and Cheng, Feng and Cheng, Tianheng and Cheng, Yufeng and others , journal=

  51. [51]

    Hu, Teng and Zhang, Jiangning and Yi, Ran and Wang, Yating and Huang, Hongrui and Weng, Jieyu and Wang, Yabiao and Ma, Lizhuang , journal=

  52. [52]

    Ling, Pengyang and Bu, Jiazi and Zhang, Pan and Dong, Xiaoyi and Zang, Yuhang and Wu, Tong and Chen, Huaian and Wang, Jiaqi and Jin, Yi , journal=

  53. [53]

    Bahmani, Sherwin and Skorokhodov, Ivan and Qian, Guocheng and Siarohin, Aliaksandr and Menapace, Willi and Tagliasacchi, Andrea and Lindell, David B and Tulyakov, Sergey , booktitle=

  54. [54]

    Wang, Zhouxia and Yuan, Ziyang and Wang, Xintao and Li, Yaowei and Chen, Tianshui and Xia, Menghan and Luo, Ping and Shan, Ying , booktitle=

  55. [55]

    Huang, Ziqi and He, Yinan and Yu, Jiashuo and Zhang, Fan and Si, Chenyang and Jiang, Yuming and Zhang, Yuanhan and Wu, Tianxing and Jin, Qingyang and Chanpaisit, Nattapol and others , booktitle=

  56. [56]

    2025 , publisher=

    Huang, Ziqi and Zhang, Fan and Xu, Xiaojie and He, Yinan and Yu, Jiashuo and Dong, Ziyue and Ma, Qianli and Chanpaisit, Nattapol and Si, Chenyang and Jiang, Yuming and others , journal=. 2025 , publisher=

  57. [57]

    He, Hao and Xu, Yinghao and Guo, Yuwei and Wetzstein, Gordon and Dai, Bo and Li, Hongsheng and Yang, Ceyuan , journal=

  58. [58]

    Wang, Qinghe and Shi, Xiaoyu and Li, Baolu and Bian, Weikang and Liu, Quande and Lu, Huchuan and Wang, Xintao and Wan, Pengfei and Gai, Kun and Jia, Xu , booktitle=

  59. [59]

    Lin, Kuan Heng and Liu, Zhizheng and Salamanca, Pablo and Kant, Yash and Burgert, Ryan and Xu, Yuancheng and Namekata, Koichi and Zhao, Yiwei and Zhou, Bolei and Goldblum, Micah and others , booktitle=

  60. [60]

    Liu, Jiwen and Li, Shujuan and Fang, Zhixue and Li, Xiaohan and Zhou, Yan and Meng, Zijie and Zhang, Zhimin and Luo, Yawen and Zhang, Guoxin and Liu, Yu-Shen and Wan, Pengfei , journal=

  61. [61]

    Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , month =

    Wang, Jianyuan and Chen, Minghao and Karaev, Nikita and Vedaldi, Andrea and Rupprecht, Christian and Novotny, David , title =. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , month =. 2025 , pages =

  62. [62]

    Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , month =

    Li, Zhengqi and Tucker, Richard and Cole, Forrester and Wang, Qianqian and Jin, Linyi and Ye, Vickie and Kanazawa, Angjoo and Holynski, Aleksander and Snavely, Noah , title =. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , month =. 2025 , pages =

  63. [63]

    Zhang, Junyi and Herrmann, Charles and Hur, Junhwa and Jampani, Varun and darrell, trevor and Cole, Forrester and Sun, Deqing and Yang, Ming-Hsuan , booktitle =

  64. [64]

    Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) , month =

    Lin, Shanchuan and Liu, Bingchen and Li, Jiashi and Yang, Xiao , title =. Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) , month =. 2024 , pages =

  65. [65]

    and Ben-Hamu, Heli and Nickel, Maximilian and Le, Matt , year =

    Lipman, Yaron and Chen, Ricky T.Q. and Ben-Hamu, Heli and Nickel, Maximilian and Le, Matt , year =

  66. [66]

    Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , month =

    Chen, Haoxin and Zhang, Yong and Cun, Xiaodong and Xia, Menghan and Wang, Xintao and Weng, Chao and Shan, Ying , title =. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , month =. 2024 , pages =

  67. [67]

    Chun-Han Yao and Yiming Xie and Vikram Voleti and Huaizu Jiang and Varun Jampani , journal=

  68. [68]

    Bahmani, Sherwin and Skorokhodov, Ivan and Siarohin, Aliaksandr and Menapace, Willi and Qian, Guocheng and Vasilkovsky, Michael and Lee, Hsin-Ying and Wang, Chaoyang and Zou, Jiaxu and Tagliasacchi, Andrea and Lindell, David and Tulyakov, Sergey , booktitle =

  69. [69]

    European Conference on Computer Vision (ECCV) , year=

    Generative Camera Dolly: Extreme Monocular Dynamic Novel View Synthesis , author=. European Conference on Computer Vision (ECCV) , year=

  70. [70]

    Wang, Yifan and Zhou, Jianjun and Zhu, Haoyi and Chang, Wenzheng and Zhou, Yang and Li, Zizun and Chen, Junyi and Pang, Jiangmiao and Shen, Chunhua and He, Tong , journal=

  71. [71]

    Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , month =

    Ren, Xuanchi and Shen, Tianchang and Huang, Jiahui and Ling, Huan and Lu, Yifan and Nimier-David, Merlin and M\"uller, Thomas and Keller, Alexander and Fidler, Sanja and Gao, Jun , title =. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , month =. 2025 , pages =

  72. [72]

    Ma, Yue and Feng, Kunyu and Hu, Zhongyuan and Wang, Xinyu and Wang, Yucheng and Zheng, Mingzhe and He, Xuanhua and Zhu, Chenyang and Liu, Hongyu and He, Yingqing and others , journal=

  73. [73]

    Si, Chenyang and Huang, Ziqi and Jiang, Yuming and Liu, Ziwei , booktitle=

  74. [74]

    Balaji, Yogesh and Nah, Seungjun and Huang, Xun and Vahdat, Arash and Song, Jiaming and Zhang, Qinsheng and Kreis, Karsten and Aittala, Miika and Aila, Timo and Laine, Samuli and others , journal=

  75. [75]

    Yu, Mark and Hu, Wenbo and Xing, Jinbo and Shan, Ying , booktitle=

  76. [76]

    Wang, Qinghe and Luo, Yawen and Shi, Xiaoyu and Jia, Xu and Lu, Huchuan and Xue, Tianfan and Wang, Xintao and Wan, Pengfei and Zhang, Di and Gai, Kun , booktitle=

  77. [77]

    Bahmani, S.; Skorokhodov, I.; Siarohin, A.; Menapace, W.; Qian, G.; Vasilkovsky, M.; Lee, H.-Y.; Wang, C.; Zou, J.; Tagliasacchi, A.; Lindell, D.; and Tulyakov, S. 2025. VD3D: Taming Large Video Diffusion Transformers for 3D Camera Control . In Yue, Y.; Garg, A.; Peng, N.; Sha, F.; and Yu, R., eds., International Conference on Learning Representations, 66...

  78. [78]

    Bai, J.; Xia, M.; Fu, X.; Wang, X.; Mu, L.; Cao, J.; Liu, Z.; Hu, H.; Bai, X.; Wan, P.; et al. 2025 a . ReCamMaster: Camera-Controlled Generative Rendering from A Single Video . In Proceedings of the IEEE/CVF International Conference on Computer Vision, 14834--14844

  79. [79]

    Bai, S.; Cai, Y.; Chen, R.; Chen, K.; Chen, X.; Cheng, Z.; Deng, L.; Ding, W.; Gao, C.; Ge, C.; et al. 2025 b . Qwen3-VL Technical Report . arXiv preprint arXiv:2511.21631

  80. [80]

    Balaji, Y.; Nah, S.; Huang, X.; Vahdat, A.; Song, J.; Zhang, Q.; Kreis, K.; Aittala, M.; Aila, T.; Laine, S.; et al. 2022. eDiff-I: Text-to-Image Diffusion Models with An Ensemble of Expert Denoisers . arXiv preprint arXiv:2211.01324

Showing first 80 references.