A 1,000-pair real-world HDR benchmark with two new affine-invariant error scores shows the best image-editing models reproduce the relative structure of real light transport but degrade in dim regions, and that VLMs fail at pixel-level light checks.
arXiv preprint arXiv:2506.15673 , year=
13 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
roles
background 1polarities
background 1representative citing papers
BodyReLux achieves photorealistic, temporally consistent full-body video relighting via a diffusion model with token-based lighting conditioning trained on a hybrid static-dynamic capture dataset.
A relightable Gaussian Splatting method for virtual production decomposes scenes into fixed appearance and variable lighting by parameterizing primitives to directly sample high-resolution background textures, enabling controllable relighting without physically-based rendering or far-field maps.
A framework for consistent long-horizon video relighting that propagates target latents across chunks and trains continuation via masked target-domain self-conditioning plus warm-start prompting.
AlbedoEdit fine-tunes video foundation models to translate RGB videos into edited versions conditioned on user-edited first-frame albedo maps, trained on a new synthetic paired dataset for insertion, removal, and texture tasks.
RoHIL adapts human-in-the-loop RL policies to new illumination conditions offline by combining world-model image relighting, illumination-retention replay, and anchored Bellman regularisation, improving shifted-light performance while preserving source performance on four real-robot tasks.
A transformer-based neural renderer that transfers arbitrary PBR lighting to single images via shared intrinsic conditioning extracted from both multi-illumination photos and path-traced coarse 3D renders.
UniVidX unifies diverse video generation tasks into one conditional diffusion model using stochastic condition masking, decoupled gated LoRAs, and cross-modal self-attention.
RenderFlow replaces iterative diffusion with flow matching for deterministic single-step neural rendering that achieves near real-time photorealistic quality and extends to inverse rendering via an adapter module.
Hybrid system that uses ray-traced 3D Gaussians to supply radiometric guidance and material regularization to a neural renderer for editable, realistic output from captured scenes.
SRUG uses shadow-guided 3D completion and iterative LMM-based material decomposition to create relightable urban scenes from sparse views.
Cosmos-Predict2.5 unifies text-to-world, image-to-world, and video-to-world generation in one model trained on 200M clips with RL post-training, delivering improved quality and control for physical AI.
citing papers explorer
-
Do Image Editing Models Understand Lighting?
A 1,000-pair real-world HDR benchmark with two new affine-invariant error scores shows the best image-editing models reproduce the relative structure of real light transport but degrade in dim regions, and that VLMs fail at pixel-level light checks.
-
BodyReLux: Temporally Consistent Full-Body Video Relighting
BodyReLux achieves photorealistic, temporally consistent full-body video relighting via a diffusion model with token-based lighting conditioning trained on a hybrid static-dynamic capture dataset.
-
Relightable Gaussian Splatting for Virtual Production Using Image-Based Illumination
A relightable Gaussian Splatting method for virtual production decomposes scenes into fixed appearance and variable lighting by parameterizing primitives to directly sample high-resolution background textures, enabling controllable relighting without physically-based rendering or far-field maps.
-
HorizonRelight: Relighting Long-horizon Videos Consistently via Diffusion Transformers
A framework for consistent long-horizon video relighting that propagates target latents across chunks and trains continuation via masked target-domain self-conditioning plus warm-start prompting.
-
AlbedoEdit: Unified Instance-Level Video Editing with Albedo Guidance
AlbedoEdit fine-tunes video foundation models to translate RGB videos into edited versions conditioned on user-edited first-frame albedo maps, trained on a new synthetic paired dataset for insertion, removal, and texture tasks.
-
RoHIL: Robust Human-in-the-Loop Robotic Reinforcement Learning Against Illumination Variations
RoHIL adapts human-in-the-loop RL policies to new illumination conditions offline by combining world-model image relighting, illumination-retention replay, and anchored Bellman regularisation, improving shifted-light performance while preserving source performance on four real-robot tasks.
-
PIXLRelight: Controllable Relighting via Intrinsic Conditioning
A transformer-based neural renderer that transfers arbitrary PBR lighting to single images via shared intrinsic conditioning extracted from both multi-illumination photos and path-traced coarse 3D renders.
-
UniVidX: A Unified Multimodal Framework for Versatile Video Generation via Diffusion Priors
UniVidX unifies diverse video generation tasks into one conditional diffusion model using stochastic condition masking, decoupled gated LoRAs, and cross-modal self-attention.
-
RenderFlow: Single-Step Neural Rendering via Flow Matching
RenderFlow replaces iterative diffusion with flow matching for deterministic single-step neural rendering that achieves near real-time photorealistic quality and extends to inverse rendering via an adapter module.
-
TRON: Tracing Rays to Orchestrate a Neural Renderer for 3D Gaussian Reconstructions
Hybrid system that uses ray-traced 3D Gaussians to supply radiometric guidance and material regularization to a neural renderer for editable, realistic output from captured scenes.
-
SRUG: Shadow-Guided Relightable Urban Scene with Generation Model
SRUG uses shadow-guided 3D completion and iterative LMM-based material decomposition to create relightable urban scenes from sparse views.
-
World Simulation with Video Foundation Models for Physical AI
Cosmos-Predict2.5 unifies text-to-world, image-to-world, and video-to-world generation in one model trained on 200M clips with RL post-training, delivering improved quality and control for physical AI.
- FusionRelight: Relighting Portraits in Real Time via Hybrid Domain Knowledge Fusion