WildRelight supplies the first in-the-wild benchmark for single-image relighting and a physics-guided self-supervised adaptation technique that turns temporal lighting changes into a domain-alignment signal.
hub
RGB↔X: Image decomposition and synthesis using material- and lighting-aware diffusion models , year =
20 Pith papers cite this work, alongside 19 external citations. Polarity classification is still indexing.
hub tools
citation-role summary
citation-polarity summary
representative citing papers
An autoregressive diffusion model with a hybrid explicit-root/latent-body representation generates real-time, controllable 3D human motion from text and spatial constraints.
A 1,000-pair real-world HDR benchmark with two new affine-invariant error scores shows the best image-editing models reproduce the relative structure of real light transport but degrade in dim regions, and that VLMs fail at pixel-level light checks.
A technique for controllable diversity in text-to-image generation by inducing structured semantic variations at the prompt level via VLM and agentic workflow.
A feed-forward transformer estimates per-pixel normal, albedo, roughness, and metallicity from single-shot spectro-polarimetric measurements captured with a polarimetric display and augmented RGB polarization camera, using a generative manifold to expand limited BRDF training data.
BodyReLux achieves photorealistic, temporally consistent full-body video relighting via a diffusion model with token-based lighting conditioning trained on a hybrid static-dynamic capture dataset.
Materialist performs single-image inverse rendering via neural-initialized progressive differentiable rendering to enable physically consistent material editing, object insertion, relighting, and transparency edits without full scene geometry.
STREAM decouples text (via AdaLN) from music (via energy-based BEAM attention) to generate editable, musically aligned dance motions with a new annotated dataset and editability metric.
A framework for consistent long-horizon video relighting that propagates target latents across chunks and trains continuation via masked target-domain self-conditioning plus warm-start prompting.
A Diffusion Transformer framework applies coordinate-transformed RoPE and disjoint attention masks to achieve controllable, high-fidelity texture tiling that preserves reference structure and scene lighting.
PTIR-GS develops a splatting-free path-traced inverse rendering method for 3D Gaussian fields to achieve consistent optimization with global illumination and multi-bounce light transport.
AlbedoEdit fine-tunes video foundation models to translate RGB videos into edited versions conditioned on user-edited first-frame albedo maps, trained on a new synthetic paired dataset for insertion, removal, and texture tasks.
PhysEditBench is a protocol-conditioned benchmark evaluating image editors on dense prediction of depth, normal, albedo, roughness, and metallic maps from RGB images using curated data and fixed scoring rules.
Scaling transformer context with sparse attention and 3D-aware block routing improves feed-forward 3D reconstruction and inverse rendering, closing much of the quality gap with dense-view optimization.
A diffusion framework decomposes images into intrinsic maps via an inverse renderer and renders controllable weather changes via a forward renderer with CLIP prompt interpolation and map-aware attention, outperforming pixel-space baselines on new 38k synthetic and 18k real datasets.
Prior-first body-hand kinematic model with layered adapters for real-time, low-supervision hand motion completion conditioned on body and semantics.
Hybrid system that uses ray-traced 3D Gaussians to supply radiometric guidance and material regularization to a neural renderer for editable, realistic output from captured scenes.
A palette-based framework decomposes 2D Gaussian Splatting scenes into shared BRDF prototypes via a spatial material field for coherent editing and relighting under physical rendering.
SRUG uses shadow-guided 3D completion and iterative LMM-based material decomposition to create relightable urban scenes from sparse views.
citing papers explorer
-
WildRelight: A Real-World Benchmark and Physics-Guided Adaptation for Single-Image Relighting
WildRelight supplies the first in-the-wild benchmark for single-image relighting and a physics-guided self-supervised adaptation technique that turns temporal lighting changes into a domain-alignment signal.
-
ARDY: Autoregressive Diffusion with Hybrid Representation for Interactive Human Motion Generation
An autoregressive diffusion model with a hybrid explicit-root/latent-body representation generates real-time, controllable 3D human motion from text and spatial constraints.
-
Do Image Editing Models Understand Lighting?
A 1,000-pair real-world HDR benchmark with two new affine-invariant error scores shows the best image-editing models reproduce the relative structure of real light transport but degrade in dim regions, and that VLMs fail at pixel-level light checks.
-
Semantic Browsing: Controllable Diversity for Image Generation
A technique for controllable diversity in text-to-image generation by inducing structured semantic variations at the prompt level via VLM and agentic workflow.
-
Snapshot Polarimetric Display Inverse Rendering
A feed-forward transformer estimates per-pixel normal, albedo, roughness, and metallicity from single-shot spectro-polarimetric measurements captured with a polarimetric display and augmented RGB polarization camera, using a generative manifold to expand limited BRDF training data.
-
BodyReLux: Temporally Consistent Full-Body Video Relighting
BodyReLux achieves photorealistic, temporally consistent full-body video relighting via a diffusion model with token-based lighting conditioning trained on a hybrid static-dynamic capture dataset.
-
Materialist: Physically Based Editing Using Single-Image Inverse Rendering
Materialist performs single-image inverse rendering via neural-initialized progressive differentiable rendering to enable physically consistent material editing, object insertion, relighting, and transparency edits without full scene geometry.
-
Text Dictates, Music Decorates: Energy-based Attention for Editable Dance Motion Generation
STREAM decouples text (via AdaLN) from music (via energy-based BEAM attention) to generate editable, musically aligned dance motions with a new annotated dataset and editability metric.
-
HorizonRelight: Relighting Long-horizon Videos Consistently via Diffusion Transformers
A framework for consistent long-horizon video relighting that propagates target latents across chunks and trains continuation via masked target-domain self-conditioning plus warm-start prompting.
-
Controllable Texture Tiling with Transformed RoPE-Enhanced Diffusion Models
A Diffusion Transformer framework applies coordinate-transformed RoPE and disjoint attention masks to achieve controllable, high-fidelity texture tiling that preserves reference structure and scene lighting.
-
PTIR-GS: Path-Traced Inverse Rendering with Global Illumination in 3D Gaussian Fields
PTIR-GS develops a splatting-free path-traced inverse rendering method for 3D Gaussian fields to achieve consistent optimization with global illumination and multi-bounce light transport.
-
AlbedoEdit: Unified Instance-Level Video Editing with Albedo Guidance
AlbedoEdit fine-tunes video foundation models to translate RGB videos into edited versions conditioned on user-edited first-frame albedo maps, trained on a new synthetic paired dataset for insertion, removal, and texture tasks.
-
PhysEditBench: A Protocol-Conditioned Benchmark for Dense Physical-Map Prediction with Image Editors
PhysEditBench is a protocol-conditioned benchmark evaluating image editors on dense prediction of depth, normal, albedo, roughness, and metallic maps from RGB images using curated data and fixed scoring rules.
-
LSRM: High-Fidelity Object-Centric Reconstruction via Scaled Context Windows
Scaling transformer context with sparse attention and 3D-aware block routing improves feed-forward 3D reconstruction and inverse rendering, closing much of the quality gap with dense-view optimization.
-
IntrinsicWeather: Controllable Weather Editing in Intrinsic Space
A diffusion framework decomposes images into intrinsic maps via an inverse renderer and renders controllable weather changes via a forward renderer with CLIP prompt interpolation and map-aware attention, outperforming pixel-space baselines on new 38k synthetic and 18k real datasets.
-
Prior-First, Condition-Second: Scalable and Controllable Hand Motion Completion
Prior-first body-hand kinematic model with layered adapters for real-time, low-supervision hand motion completion conditioned on body and semantics.
-
TRON: Tracing Rays to Orchestrate a Neural Renderer for 3D Gaussian Reconstructions
Hybrid system that uses ray-traced 3D Gaussians to supply radiometric guidance and material regularization to a neural renderer for editable, realistic output from captured scenes.
-
MaterialClusterGS: Palette-Based Material Decomposition and Physically-Based Relighting with 2D Gaussian Splatting
A palette-based framework decomposes 2D Gaussian Splatting scenes into shared BRDF prototypes via a spatial material field for coherent editing and relighting under physical rendering.
-
SRUG: Shadow-Guided Relightable Urban Scene with Generation Model
SRUG uses shadow-guided 3D completion and iterative LMM-based material decomposition to create relightable urban scenes from sparse views.
- ExpertEdit: Learning Skill-Aware Motion Editing from Expert Videos