{"id":"76a0853f-e2c2-4bb4-a68f-724131401b8c","arxiv_id":"2606.20233","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"An end-to-end diffusion model generates backgrounds, relights green-screen actors, and replaces or creates props in one pass, improving over cascaded inpainting plus relighting baselines in user preference and auto metrics.","lead":"Researchers built a video-generation system that composites green-screen actors into new scenes while relighting them to match the new environment and keeping props looking plausible. It trains one video-diffusion model to handle both background generation and character relighting instead of chaining two separate tools.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Relighting supervision and benchmark both use IC-Light; reported gains may be teacher-mimicry, not physical correctness.","rationale":"Following the paper's logic: the central claim is that a single diffusion model can simultaneously generate a dynamic background, relight the actor, and handle props better than existing cascades. What has to be true is that the relighting supervision is an accurate target for physical illumination, and that the evaluation measures that target. Both conditions fail if the only relighting signal comes from IC-Light. I checked the full text: Sec. 3.2.2 explicitly states V_relit is generated with IC-Light; Sec. 4.2 states the benchmark inputs are augmented with the same strategy. This is an in-distribution teacher-student setup. The baselines in Table 1 are either non-relighting (FLUX+animation, VACE) or a cascade that applies IC-Light only once (VACE+Relighting); none represents a state-of-the-art video relighting model (e.g., RelightVid, Light-A-Video, LightCtrl, UniLumos). The paper's own Related Work lists these methods, so their absence from the comparison is a gap. The ablation in Table 4 shows the lighting data improves metrics, but that is expected if the model is learning to mimic the teacher; it does not establish physical correctness. The user study is the only independent evidence, but with 10 videos and no significance testing, it is too weak to carry the headline claim alone. The conclusion (Sec. 5) also explicitly limits the work to below-4K and short clips; this is a scope limitation, not a correctness flaw, but it qualifies 'cinematic-quality' claims. Therefore the most load-bearing concern is the IC-Light circularity, and the recommended test is an external renderer-based or light-stage evaluation. This matches the reader's weakest assumption; I agree.","tokens_in":14372,"tokens_out":10265,"duration_ms":97844,"concrete_test":"Run the proposed model and the VACE+Relighting baseline on a benchmark with physically rendered relighting ground truth, e.g., RelightVid's game-engine-rendered test videos (time-varying environment maps) or a light-stage capture with known HDR environment. For each method, compute foreground relighting error (MSE/LPIPS against the rendered ground truth) and repeat the user study on a third-party set of at least 30 clips. If the proposed model does not significantly outperform VACE+Relighting and IC-Light on this external data, the synthetic benchmark's IC-Light augmentation is the source of the claimed gains and the physical-correctness claim is unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that the framework 'ensures physically consistent interactions' and 'significantly outperforms existing methods' rests on the relighting supervision being correct. However, both the training target V_relit (Sec. 3.2.2) and the synthetic benchmark inputs (Sec. 4.2) are generated with IC-Light [Zhang et al., 2025]. The model is therefore trained to reproduce a specific teacher model's relighting, and the benchmark evaluates how well it reproduces that same teacher on inputs that were also relit by that teacher. Any systematic bias of IC-Light — e.g., limited handling of dynamic lighting, specific color casts, per-frame temporal inconsistency — is baked into both the learning target and the evaluation. Consequently, the Foreground Similarity score (SSIM under the foreground mask) largely measures fidelity to IC-Light's outputs, not to physical ground-truth illumination. The measured advantage over VACE+Relighting, which uses IC-Light only as a post-hoc step, is exactly what one would expect from a student that has been directly optimized to match the teacher on the same augmentation distribution. The synthetic benchmark thus cannot support the claim of physical correctness or of superiority over stronger relighting methods. The only independent evidence is the user study on 10 real green-screen videos (Sec. 4.4), which lacks error bars and significance testing; with 10 videos and a few experts, the 60%+ preference margins are suggestive but not conclusive. This concern is load-bearing because it undermines the strongest quantitative evidence for the headline claim.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes an end-to-end video diffusion framework for green-screen compositing that models both character-to-environment (C2E) physical interaction and environment-to-character (E2C) lighting harmonization in a single model. The method introduces a tri-mask conditioning interface distinguishing preserve-and-relight, geometry-preserving replace, and full-generation regions; a joint RGB-D denoising strategy built on a DiT backbone; a prior-driven data curation pipeline that uses IC-Light relighting and AI-generated multi-illumination clips to construct training pairs; and a reference-conditioned mechanism for controllable environment and prop generation. Experiments are reported on a synthetic benchmark of 100 HOIGen1M clips and a user study on 10 real green-screen videos, with quantitative metrics showing gains over several cascaded baselines.","tokens_in":14682,"tokens_out":3674,"duration_ms":36550,"significance":"If the claims are substantiated, the paper makes a useful contribution by packaging C2E occlusion/interaction and E2C relighting into a single tri-mask-guided video diffusion interface, which is more unified than current cascaded pipelines. The RGB-D joint denoising is a technically sensible departure from RGB-only inpainting, and the data-curation pipeline is a reasonable attempt to avoid expensive rendered relighting pairs. The strongest parts are the problem decomposition and the coherent training/inference interface. However, the central quantitative claim of superior relighting and physical consistency rests on a benchmark that shares its relighting teacher with the training supervision and on a small, unmeasured user study; these issues must be addressed before the state-of-the-art claim is credible.","major_comments":[{"comment":"The synthetic benchmark is circular with respect to the main claim about E2C relighting. The training target V_relit is generated with IC-Light (Sec. 3.2.2), and the benchmark inputs are 'augmented using the same strategy introduced in Sec. 3.2' (Sec. 4.2). Therefore the Foreground Similarity metric primarily measures fidelity to IC-Light outputs, not physical or perceptually correct relighting. The gap over VACE+Relighting, which uses IC-Light only as a post-hoc step, is expected from a model directly trained to match IC-Light on the same augmentation distribution. The benchmark also samples from the same HOIGen1M base dataset with the same filtering criteria used in training, so the in-domain nature of the task further inflates the reported numbers. To support the claim of superior relighting and physical consistency, the authors should evaluate on independent ground truth (e.g., rende","section":"Sec. 4.2, Sec. 3.2.2, Table 1"},{"comment":"The user study on the real green-screen benchmark is too weak to support the strong preference claims. It uses 10 videos and reports percentage preferences without participant count, error bars, confidence intervals, or significance tests. The 61–69% preference ratios look suggestive, but with a small number of videos and evaluators, a few outliers can dominate. The current presentation does not rule out chance-level agreement or a strong ordering bias. At minimum, the paper should report the number of participants, number of comparisons per video, inter-rater agreement (e.g., Fleiss' kappa), and a statistical test (e.g., binomial test or Wilcoxon) for the preference margins. Without this, Table 2 should not be described as evidence that users 'consistently favor' the method.","section":"Sec. 4.4, Table 2"},{"comment":"The baseline comparison omits the most relevant task-specific methods. The paper compares against generic image-editing-plus-animation cascades and vanilla VACE, but the related work lists video relighting methods (RelightVid, Light-A-Video, LightCtrl), subject-aware background generation (ActAnywhere), and interactive character generation (Animate Anyone 2, MoCha). None of these are included as baselines. The claim of 'significantly outperforming existing methods' in cinematic compositing is therefore not established against the strongest prior work. The authors should compare against at least one state-of-the-art video relighting model and one subject-aware background synthesis model, even if adapted to the green-screen setting, or explicitly justify why such comparisons are not feasible.","section":"Sec. 2, Sec. 4.2"},{"comment":"The E2C supervision is generated by applying IC-Light to each frame independently. IC-Light is an image-based relighting model and is not designed to enforce temporal coherence across video frames. The paper does not analyze the temporal consistency of the relit training targets, nor does it report any temporal metrics on the generated videos. Given that the task is dynamic video compositing, unstable per-frame relighting supervision could teach the model to produce flickering illumination. The authors should either verify that their IC-Light-generated targets are temporally stable, or add a temporal consistency metric (e.g., warped-frame error or CLIP temporal consistency) on outputs.","section":"Sec. 3.2.2"},{"comment":"The objective metrics are reported without variance or statistical significance. The differences between the best and second-best methods are small (e.g., Foreground Similarity 0.642 vs 0.629, Background Similarity 0.704 vs 0.688), and the synthetic benchmark has only 100 clips. Without confidence intervals or a paired test, it is unclear whether these gaps are robust. This is especially important given the benchmark's dependence on the same relighting teacher, as a small in-domain advantage could be an artifact of the training/evaluation overlap.","section":"Sec. 4.2, Table 1"}],"minor_comments":[{"comment":"Typo: 'proprelighting' should be 'prop relighting'.","section":"Figure 1 caption"},{"comment":"The Aesthetic Score is computed with a LAION aesthetic predictor designed for images. The paper does not explain how it is aggregated over video frames or whether any temporal aggregation is used. Please clarify.","section":"Sec. 4.3"},{"comment":"The VACE+Relighting baseline removes the foreground with Attentive Eraser before relighting. This may introduce artifacts and unfairly disadvantage the baseline. Please report or discuss the impact of the eraser step, or use a cleaner foreground-removal method.","section":"Sec. 4.2"},{"comment":"The paper states that the model is fine-tuned on 65-frame clips, while Sec. 3.2.2 says the retained clips are trimmed to 81 frames. The relationship between these numbers is not explained; please clarify.","section":"Sec. 4.1"},{"comment":"The multi-illumination subset is generated with Wan 2.2 T2V and then used as 'ground truth' in the five-tuple. Calling AI-generated video 'ground truth' is misleading; consider calling it 'source video' or 'synthesized reference'.","section":"Sec. 3.2.2"}],"recommendation":"major_revision","confidential_remarks":"The paper presents a clean and interesting framework, but the evaluation is the main weakness. The synthetic benchmark is trained and tested on the same relighting teacher, so the quantitative superiority claim is not yet convincing. The authors should be given the opportunity to provide an independent evaluation (rendered or real ground-truth relighting) and a more rigorous user study. I would not reject the paper, as the architectural contributions and the tri-mask interface are valuable even if the current benchmark does not fully prove the state-of-the-art claim."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The short version: this is a solid systems paper with a genuinely new problem decomposition, but the quantitative relighting story is weakened by using IC-Light for both training supervision and benchmark augmentation.\n\nWhat's new: the tri-mask formulation (preserve-and-relight, depth-anchored regenerate, full generation), joint RGB-D denoising in a single video diffusion latent trajectory, and a prior-driven data curation pipeline that synthesizes multi-illumination pairs without rendering. Those are real contributions beyond stitching existing tools together. The reference injection with a canvas token is a nice practical touch for prop placement.\n\nWhat it does well: the ablations match the design narrative—adding depth helps background similarity and interaction realism, adding lighting data helps aesthetic and foreground similarity. The user study on 10 real green-screen videos shows strong preference margins, which is independent evidence that the pipeline works end-to-end. The writing is clear about what is being formulated.\n\nSoft spots: the big one is in Sections 3.2.2 and 4.2. Training pairs are built by relighting with IC-Light, and the synthetic benchmark augments inputs with the same IC-Light strategy. So the Foreground Similarity metric largely measures how well the model reproduces its teacher, and the gain over VACE+Relighting (which uses IC-Light only as a post-hoc step) is expected from direct optimization. That does not support a claim of physically consistent relighting. Identity and background metrics are less compromised, but the headline relighting claim rests on the circular part. The baseline set is thin—no direct comparison to RelightVid, Light-A-Video, or other recent relighters. The real benchmark is 10 videos with no error bars or significance testing; the 60-70% preferences are suggestive, not conclusive. No code or weights are released, which makes it hard to verify the engineering claims. The depth-cue reliance is a reasonable minor worry; errors in Video Depth Anything will propagate to the claimed interaction quality.\n\nWho it's for: anyone working on generative VFX, video editing, or compositing pipelines. The tri-mask idea alone is worth discussing. It deserves a serious referee, but the authors should be pushed to evaluate on an independently relit benchmark, add more baselines, and release code. The central idea holds up; the evidence is currently uneven.","headline":"Unified tri-mask RGB-D green-screen compositing is a genuine and useful contribution, but the relighting evaluation is partly student–teacher circularity; the paper deserves a serious referee but needs an external benchmark.","tokens_in":15209,"tokens_out":2248,"would_cite":false,"duration_ms":22958,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that green-screen video compositing can be solved by a single video diffusion model that jointly handles background synthesis, actor relighting, and three modes of prop handling, via a tri-mask-guided RGB-D architecture.","keywords":["video compositing","green screen","video diffusion model","relighting","human-object interaction","RGB-D joint denoising","tri-mask guidance","generative video editing"],"falsifier":"Take a set of real green-screen clips with measured environment lighting (for example, using a light probe), have the model composite each actor into that environment, and compare the output to ground-truth footage of the same actor shot in the same physical environment. If the model cannot reproduce the measured lighting direction and color within a defined tolerance, the claim of genuine environment-to-character harmonization is weakened. A complementary check is to train the model with physically rendered relighting targets instead of algorithm-generated ones and see whether the gap over th","tokens_in":1354,"feed_emoji":"🎬","tokens_out":2363,"duration_ms":55776,"temperature":0.7,"pith_summary":"The paper claims that cinematic green-screen compositing should be treated as one generative problem rather than a pipeline of separate matting, inpainting, and relighting tools. It proposes a video diffusion model that simultaneously synthesizes a new dynamic background, relights the foreground actor to match that background, and handles props in three ways: preserve-and-relight, replace-with-new-object, or generate-from-scratch. To make this possible, it introduces a tri-mask that assigns each pixel one of the three modes, and a joint RGB-D denoising strategy that reasons about geometry and appearance together. If this works, filmmakers could replace a cascade of static tools with a single prompt- and mask-driven model that keeps identity, lighting, and contact consistent. The paper backs this with quantitative benchmarks and a user study in which the proposed method receives the majority preference on all evaluation criteria.","feed_headline":"One video diffusion model replaces the green-screen VFX cascade","feed_subtitle":"A tri-mask steers preserve, replace, or generate per region while joint RGB-D denoising keeps lighting and contact consistent.","key_machinery":"The central mechanism is the tri-mask, which assigns each pixel one of three states: preserve-and-relight (mask value 1), geometry-preserving regeneration (mask value 0), or full generation (mask value -1). The tri-mask is converted into separate RGB-preservation and depth-preservation masks, and the masked RGB and depth sequences are jointly denoised in a single latent trajectory by a Diffusion Transformer. This lets the model explicitly use depth as a geometric anchor while regenerating appearance, and lets it adaptively choose the right conditioning per region. The training signal is organized as a five-tuple (ground-truth video, relit counterpart, depth, tri-mask, text prompt) built thro","core_discovery":"The central claim is that a unified video diffusion model can jointly model character-to-environment physical interaction (C2E) and environment-to-character lighting harmonization (E2C), including interactive props, more effectively than cascaded pipelines. The authors assert that a tri-mask-guided architecture with joint RGB-D denoising produces composite videos where the background responds to actor motion and geometry, the actor is relit coherently with the new environment, and props can be preserved, replaced, or generated according to three mask states. They report that the method outperforms first-frame editing plus animation, inpainting-only, and inpainting-plus-relighting baselines o","pith_inferences":["If the model's relighting quality is bounded by the relighting prior used to generate training targets, then training on physically rendered relighting pairs could reveal larger gains over the same baselines; the current benchmark may understate the method's true ceiling.","The tri-mask formulation is a general region-semantics interface that could be adopted by other generative video editing tasks that need to mix preservation, partial regeneration, and full generation in one pass.","Because the model trusts monocular depth as a geometric anchor, its interaction quality likely degrades when depth estimation is wrong; testing on foregrounds with known synthetic depth would isolate that sensitivity.","The reference canvas injection suggests a path toward interactive editing tools where artists compose a rough layout and the model generates temporally consistent video around it, potentially extendable to multiple reference objects with explicit occlusion ordering."],"forward_implications":["A single model instance can replace the conventional cascade of matting, background inpainting, and relighting for dynamic green-screen shots, reducing error accumulation between stages.","The tri-mask interface lets an artist decide per region whether to keep the captured actor's appearance (with relighting), keep only geometry (for prop replacement), or let the model generate content from scratch.","Joint RGB-D denoising gives the model an explicit geometric anchor, which should improve hand-object contact and body-environment occlusion plausibility in generated composite video.","The prior-driven data pipeline offers a way to build large relighting supervision sets without game-engine rendering, which could be reused for other video harmonization and editing tasks.","Reference-conditioned generation with a first-frame canvas enables spatially controlled prop placement and environment customization, supporting product placement and directed scene changes."],"fun_headline_variants":["Video diffusion model fuses actor and scene lighting","Tri-mask video compositing keeps props and relight true","One diffusion pass for green-screen character relighting","Video diffusion joins actor motion with environment light","C2E and E2C in one model: cinematic VFX simplified"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"The model learns relighting from training pairs in which the correct relit result is produced by an image relighting algorithm, so if that algorithm systematically biases or degrades illumination, the model's relighting ability is capped by that bias and part of the measured improvement over baselines could be imitation of the teacher rather than genuinely better harmonization.","fun_headline_variants_meta":{"raw":{"variants":["Video diffusion model fuses actor and scene lighting","Tri-mask video compositing keeps props and relight true","One diffusion pass for green-screen character relighting","Video diffusion joins actor motion with environment light","C2E and E2C in one model: cinematic VFX simplified"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":9.5e-05,"raw_usage":{"total_tokens":806,"prompt_tokens":681,"completion_tokens":125,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":425,"completion_tokens_details":{"reasoning_tokens":46}},"tokens_in":425,"tokens_out":125,"duration_ms":2236,"temperature":1.0,"reasoning_tokens":46,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T10:51:46.760950+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a set of real green-screen clips with measured environment lighting (for example, using a light probe), have the model composite each actor into that environment, and compare the output to ground-truth footage of the same actor shot in the same physical environment. If the model cannot reproduce the measured lighting direction and color within a defined tolerance, the claim of genuine environment-to-character harmonization is weakened. A complementary check is to train the model with physically rendered relighting targets instead of algorithm-generated ones and see whether the gap over th","supporting_citations":[],"review_version":2}