{"id":"54c66706-1f9b-412e-bc03-7e267db7e2b9","arxiv_id":"2506.15673","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Jointly predicting albedo and relit appearance with one video-diffusion pass improves relighting fidelity and generalization over two-stage inverse-plus-forward pipelines.","lead":"UniRelight trains a video-diffusion model to output an albedo map and a relit video in one pass, conditioned on a target HDR lighting probe, instead of splitting the task into separate inverse and forward rendering steps. The authors report that this joint formulation generalizes better to real scenes and outperforms existing two-stage and end-to-end relighting baselines on synthetic and real benchmarks.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The claimed generalization benefit of 150k auto-labeled real videos rests on pseudo-albedo labels from the authors' own synthetic-fine-tuned inverse renderer; the only reported ablation evidence is a user study within noise of chance (55%±8%), so the auto-labeling premise is not established.","rationale":"The reader's weakest assumption is the load-bearing point: the 150k real-world clips are the main ingredient claimed to provide domain generalization, but their labels come from the authors' own inverse renderer fine-tuned on synthetic data. If those labels are biased, the real-world training signal is not an independent source of realism; it is a loop through the same synthetic-render priors. The paper's only quantitative support for this data's contribution is Table 4's user study, which is statistically indistinguishable from chance (45%±8% preferred the version without auto-labeled data, i.e., 55%±8% for the version with it; about 1.25 standard deviations from 50%). The MIT multi-illumination benchmark result (20.76 vs 17.29 dB PSNR) is strong evidence for the full system against baselines, but it does not isolate the auto-labeled data's contribution. The proposed check uses the MIT test set's 25 illuminations to test the labeler's albedo invariance, a necessary property for unbiased albedo; if the pseudo-albedo varies with illumination, the labels bake shading, and the auto-labeled generalization claim is unsupported. If the check passes, the bias concern is resolved, though the temporal-consistency claim would still need a dedicated metric. I agree with the reader's assessment and see no reason to change the conditional verdict.","tokens_in":15629,"tokens_out":6599,"duration_ms":74048,"concrete_test":"On the held-out MIT multi-illumination test set (30 scenes, 25 illuminations), run the auto-labeler from Section 4.2 on each of the 25 illuminations independently (without averaging) and compute per-pixel albedo variance across illuminations, focusing on regions exhibiting cast shadows or specular highlights. Because true albedo is illumination-invariant, statistically significant variance that tracks illumination direction would demonstrate that the pseudo-albedo labels bake shading, confirming the bias concern and invalidating the auto-labeled training signal as unbiased real-world supervision.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 4.2 states that the 150k real-world clips are auto-labeled with albedo maps from a re-implemented inverse renderer fine-tuned on the authors' synthetic dataset. The only support that this data improves relighting is the StreetScenes user study in Table 4: the version with auto-labeled data is preferred 55%±8% (equivalently, the without-auto-labeled version is preferred 45%±8%), which is within roughly 1.25 standard deviations of chance. No quantitative comparison of the two ablated models on the MIT test set or any other held-out multi-illumination benchmark is provided. In the real-world training objective (Section 4.3), the pseudo-albedo is a conditioning input and the original RGB video is the target; if the pseudo-albedo bakes in shadows or synthetic shading, the model is trained to reproduce that shading from the albedo, reinforcing rather than correcting the synthetic bias. The generalization advantage of the full model is therefore potentially an artifact of the labeler rather than evidence for the joint decomposition-synthesis formulation. The abstract's temporal-consistency claim is additionally unsupported: no temporal metric is reported in Tables 1-4.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents UniRelight, a video diffusion framework that jointly denoises the latent of a relit video and its albedo in a single pass, conditioned on the input video and a target HDR environment map. Training combines a new 108k-clip synthetic dataset, the MIT multi-illumination dataset, and 150k automatically labeled real-world clips. The method is evaluated on a held-out synthetic set and the MIT test set, reporting higher PSNR/SSIM/LPIPS than DiLightNet, NeuralGaffer, DiffusionRenderer, and a Cosmos-backed re-implementation of DiffusionRenderer, together with user-study preferences.","tokens_in":15770,"tokens_out":4814,"duration_ms":52581,"significance":"The headline relighting result is externally grounded on the MIT multi-illumination benchmark with the light-probe masking protocol and on a held-out synthetic set whose assets are disjoint from training. The comparison against a re-implemented DiffusionRenderer on the same Cosmos backbone is a reasonable attempt to isolate the algorithmic contribution. The paper also demonstrates the practical value of joint albedo prediction for avoiding shadow baking. However, the evidence for the auto-labeled real-data contribution is not yet statistically supported, and the claimed temporal-consistency advantage is not measured. If the identified gaps are filled, the contribution would be a solid advance for video relighting.","major_comments":[{"comment":"The claim that auto-labeled real-world data improves generalization is not supported by the reported statistics. The only quantitative evidence is the StreetScenes user study, where the full model is preferred over the no-auto-labeled variant in 55%±8% of samples (Table 4); this is within roughly one standard deviation of chance. No PSNR/SSIM/LPIPS comparison between these two variants is given on the MIT test set or any other held-out multi-illumination benchmark. Because the pseudo-albedo labels are produced by an inverse renderer that was fine-tuned on the authors' own synthetic data, the auto-labeled training signal may reinforce synthetic-render biases rather than correct them. Please provide metric-based ablation evidence, or weaken the generalization claim accordingly.","section":"4.2, 4.3, Table 4"},{"comment":"The abstract and Section 5 claim that UniRelight surpasses previous methods in both visual fidelity and temporal consistency, but no temporal consistency metric is reported anywhere in Tables 1-4 or the appendix. The reported metrics (PSNR, SSIM, LPIPS) are per-frame, and the MIT user study is image-based; the StreetScenes user study asks about shadows and reflections, not temporal coherence. Please add a temporal consistency evaluation (e.g., warping error, temporal flicker metric, or a user study targeting temporal artifacts) or remove the temporal-consistency claim.","section":"Abstract, Section 5"},{"comment":"The joint-modeling ablation is evaluated only on SyntheticScenes, where the gain over the no-joint-modeling variant is small (26.97 vs 26.42 PSNR). The qualitative claim of improved out-of-distribution generalization (Figure 5) would be strengthened by a quantitative comparison on MIT or another real-world test set with ground truth. Without this, the paper's central claim that joint decomposition improves generalization across domains rests on a single qualitative figure.","section":"5.2, Table 3"}],"minor_comments":[{"comment":"There is a duplicated article in 'Since the the original input video and the environment map information are unavailable'; it should read 'Since the original input video...'.","section":"4.3"},{"comment":"The phrase 'performs on or par with previous work' contains a typo; it should be 'on par with previous work'.","section":"5.1"},{"comment":"The two-column layout of Table 2 merges entries such as 'IntrinsicImageDiffusion [32]16.41 0.543 0.395'; adding explicit line breaks or column separators would improve readability.","section":"Table 2"},{"comment":"The pseudo-albedo labeler is described as producing 'high-quality albedo estimates', but no quantitative validation of this labeler on independent albedo ground truth is provided; reporting such validation would strengthen the auto-labeling argument.","section":"4.2"},{"comment":"References [1] and [47] both cite the Cosmos world foundation model platform preprint; they should be consolidated to a single citation to avoid duplicate entries.","section":"References"},{"comment":"The text states that classifier-free guidance is not used at inference, while Section 4.3 describes a 10% condition-dropout probability; the relationship between these choices should be explained in one sentence for clarity.","section":"B.1"}],"recommendation":"major_revision","confidential_remarks":"The main barrier to acceptance is the unestablished benefit of the auto-labeled real-world data; the authors should be asked for a quantitative ablation on a held-out multi-illumination test set. The temporal-consistency claim also needs support. The core relighting result is promising and externally grounded, so this seems fixable within the manuscript's scope."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First the bottom line: this paper deserves a serious referee. The core idea—concatenating albedo and relit-video latents along the time dimension and denoising them jointly in a single DiT pass—is new relative to the cited two-stage pipelines (RGB↔X, DiffusionRenderer) and to the joint-appearance-motion work (VideoJAM). The headline result is externally grounded: on the MIT multi-illumination test set, the method reaches 20.76 dB PSNR vs 17.29 for the strongest baseline (DiffusionRenderer with Cosmos backbone), with light probes masked per protocol, and 84–96% user-study preference. Those are real numbers against held-out ground truth.\n\nThe joint-modeling ablation is also convincing qualitatively: without it, the relight model bakes shadows from the input (Figure 5), and the quantitative gap on SyntheticScenes (26.97 vs 26.42) is modest but in the right direction. I also credit the paper for being honest about the NeuralGaffer LPIPS outlier and for reporting that its albedo is only on par with, not better than, DiffusionRenderer (Cosmos).\n\nNow the soft spots. The biggest one is the auto-labeled real-data claim. The paper says 150k real clips auto-labeled with albedo from a re-implemented inverse renderer fine-tuned on their own synthetic data makes the model generalize. The only direct ablation evidence is a user study on StreetScenes where the version with auto-labeled data wins 55%±8%—within noise of chance. No quantitative comparison with/without auto-labeled data on MIT or any other held-out multi-illumination set is provided. That is a real gap, and the stress-test concern is fair: if the pseudo-albedo bakes in shadows or synthetic shading, the training objective (condition on pseudo-albedo, reconstruct RGB) would reinforce that bias. I would not call the central relighting result circular because the MIT benchmark is external, but the generalization claim to in-the-wild scenes is under-supported.\n\nSecond, the paper claims temporal consistency in the abstract and intro but never measures it. Table 1 uses only image metrics; there is no video consistency metric. For a video relighting paper that is a noticeable omission.\n\nThird, no code, data, or weights are released, which limits reproducibility. Not fatal, but noted.\n\nWho is this for? Anyone working on intrinsic decomposition, video diffusion, or relighting. The core formulation and the external benchmark are worth engaging with. It deserves a proper peer review with an expectation that the authors address the auto-labeling ablation and temporal-consistency measurement in revision.","headline":"A genuinely new joint-decomposition-and-synthesis formulation for video relighting, with a strong external benchmark win on MIT; the auto-labeled real-data generalization claim is the weak link.","tokens_in":16508,"tokens_out":5077,"would_cite":true,"duration_ms":46539,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Video relighting improves when albedo and relit appearance are predicted together in a single diffusion denoising pass, rather than through separate inverse and forward rendering stages.","keywords":["video relighting","intrinsic decomposition","albedo estimation","diffusion transformer","joint denoising","temporal consistency","HDR environment lighting","single-pass synthesis"],"falsifier":"Re-train the same pipeline on the same 150k real clips but with albedo labels produced by an independent inverse renderer or by averaging multi-illumination captures; if the multi-illumination benchmark scores and the street-scene user preference revert to parity with the two-stage baseline, then the reported generalization gain is carried by the auto-labeler's bias rather than by joint denoising.","tokens_in":15285,"feed_emoji":"💡","tokens_out":8794,"duration_ms":87873,"temperature":0.7,"pith_summary":"The paper argues that relighting a video is better learned as one generative step than as two: rather than first inverse-rendering G-buffers (albedo, normals, depth) and then forward-rendering the relit image, the proposed model denoises the relit video and its albedo together in a single diffusion pass. The central claim is that this joint formulation gives the model an implicit understanding of scene intrinsics, which improves generalization to real-world scenes and temporal consistency. The paper supports the claim with a synthetic dataset of 108k rendered clips and 150k auto-labeled real-world clips, reporting higher PSNR/SSIM/LPIPS than two-stage baselines on the multi-illumination benchmark and preferred by users 84–96% of the time. A skeptical reader should note that the real-world labels come from the authors' own inverse renderer, so the generalization result rests partly on the quality of that labeler.","feed_headline":"A single joint diffusion pass beats two-stage video relighting","feed_subtitle":"Joint albedo and relit-video denoising lifts real-scene PSNR by 3.5 dB over two-stage baselines","key_machinery":"The load-bearing mechanism is the concatenated-latent joint denoising pass: the diffusion transformer denoises a single token sequence formed by stacking the latent of the relit video (with HDR lighting features concatenated along the channel dimension) and the latent of the albedo along the temporal/frame dimension, with the input video as conditioning. Type embeddings and binary condition masks tell the transformer which tokens are input, albedo, or relit output, and the training objective sums an $\\ell^2$ loss on the relit latent with a ten-times-smaller weight $\\lambda_a = 0.1$ on the albedo latent. This single-pass cross-modal self-attention is what lets albedo demodulation act as a prior for relighting.","core_discovery":"The paper's central discovery is that relighting and albedo demodulation can be solved as one joint denoising problem rather than sequentially. The model concatenates the latent codes of the input video, the albedo, and the relit video along the temporal dimension, adds learnable type embeddings and condition masks, and fine-tunes a video diffusion transformer to simultaneously predict the relit video and the albedo. On the multi-illumination benchmark the method reaches PSNR 20.76 (SSIM 0.749, LPIPS 0.251), beating the strongest two-stage baseline at 17.29 (0.622, 0.355); on held-out synthetic scenes it reaches 26.97 (0.847, 0.190) versus 26.61 (0.841, 0.222). The paper interprets this as evidence that joint prediction makes the model learn an internal representation of scene structure, reducing the error accumulation that plagues inverse-plus-forward pipelines.","pith_inferences":["A testable extension the paper does not run: apply the same concatenated-latent joint denoising to other coupled inverse/synthesis pairs, such as depth or normals with novel-view synthesis, to see whether the generalization gain is specific to albedo or general to joint intrinsics.","The paper's own limitation section concedes that emitting objects, such as lights toggled inside a scene, are out of scope; that boundary follows from conditioning only on environment maps and marks the edge of the joint-denosing claim.","Because the real-world albedo labels come from the authors' own inverse renderer fine-tuned on their synthetic data, the generalization story is only as strong as that labeler; an independent albedo ground-truth check on a small real set would be a cheap decisive test.","The reported preference for the auto-labeled variant (55% vs 45%, within ±8%) sits inside the noise band, so the perceptual benefit of real-world data may be smaller than the qualitative figures suggest."],"forward_implications":["Relighting can be done from a single image or video in one generative pass, so the model no longer needs explicit G-buffer estimates and avoids inverse-to-forward error accumulation.","Because albedo demodulation is trained jointly, the model transfers to out-of-domain scenes without baking input shadows into the relit output, as shown on urban street scenes.","Adding 150k auto-labeled real-world RGB–albedo clips improves perceptual quality on natural scenes beyond what synthetic and multi-illumination data alone provide.","The same joint-trained model can be used without its albedo output at inference time, so the albedo head is a training-time prior rather than a runtime requirement.","On a 57-frame video the single pass runs in 445.5 seconds, less than the 566.6–780.0 seconds reported for two-stage baselines, because it replaces five inverse passes plus one forward pass."],"supporting_citations":[{"why":"the two-stage inverse-plus-forward video relighting pipeline that the paper extends and re-implements as its strongest baseline.","marker":"[39]"},{"why":"introduces joint denoising of two modalities in one DiT pass, the design seed for jointly denoising albedo and relit video.","marker":"[12]"},{"why":"supplies the pretrained video DiT and VAE that UniRelight fine-tunes for joint relighting and albedo prediction.","marker":"[1]"},{"why":"the video foundation model used to re-implement the inverse renderer that auto-labels 150k real-world clips.","marker":"[47]"},{"why":"the multi-illumination benchmark that provides real paired relighting data and the test set for the main evaluation.","marker":"[46]"},{"why":"a diffusion-based relighting baseline with explicit lighting conditioning, compared in Table 1 and the user study.","marker":"[26]"},{"why":"the RGB↔X two-stage decomposition-and-synthesis framework that typifies the separate inverse-then-forward approach the paper argues against.","marker":"[67]"},{"why":"an intrinsic-image diffusion baseline used to evaluate albedo estimation quality.","marker":"[32]"}],"fun_headline_variants":["UniRelight: joint denoising tops two-stage relighting","One-pass joint relighting beats inverse+forward pipelines","Joint albedo and relit prediction outperforms cascade","Single diffusion pass improves video relighting fidelity","Joint decomposition and synthesis lifts relighting quality"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The claim that the model generalizes to real-world scenes rests on the 150k auto-labeled real videos, whose albedo labels come from the authors' own inverse-rendering model fine-tuned on their synthetic data; if that labeler bakes in synthetic shading or shadows, the training signal reinforces rather than corrects domain bias.","fun_headline_variants_meta":{"raw":{"variants":["UniRelight: joint denoising tops two-stage relighting","One-pass joint relighting beats inverse+forward pipelines","Joint albedo and relit prediction outperforms cascade","Single diffusion pass improves video relighting fidelity","Joint decomposition and synthesis lifts relighting quality"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000226,"raw_usage":{"total_tokens":1456,"prompt_tokens":924,"completion_tokens":532,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":540,"completion_tokens_details":{"reasoning_tokens":456}},"tokens_in":540,"tokens_out":532,"duration_ms":6148,"temperature":1.0,"reasoning_tokens":456,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T23:54:10.581720+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-train the same pipeline on the same 150k real clips but with albedo labels produced by an independent inverse renderer or by averaging multi-illumination captures; if the multi-illumination benchmark scores and the street-scene user preference revert to parity with the two-stage baseline, then the reported generalization gain is carried by the auto-labeler's bias rather than by joint denoising.","supporting_citations":[{"cited_title":"A multi-illumination dataset of indoor object appearance","cited_arxiv_id":null,"evidence_quote":"the multi-illumination benchmark that provides real paired relighting data and the test set for the main evaluation."},{"cited_title":"Neural Gaffer: Relighting any object via diffusion","cited_arxiv_id":null,"evidence_quote":"a diffusion-based relighting baseline with explicit lighting conditioning, compared in Table 1 and the user study."}],"review_version":1}