REVIEW 2 cited by
TimeRewind: Rewinding Time with Image-and-Events Video Diffusion
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
This paper addresses the novel challenge of ``rewinding'' time from a single captured image to recover the fleeting moments missed just before the shutter button is pressed. This problem poses a significant challenge in computer vision and computational photography, as it requires predicting plausible pre-capture motion from a single static frame, an inherently ill-posed task due to the high degree of freedom in potential pixel movements. We overcome this challenge by leveraging the emerging technology of neuromorphic event cameras, which capture motion information with high temporal resolution, and integrating this data with advanced image-to-video diffusion models. Our proposed framework introduces an event motion adaptor conditioned on event camera data, guiding the diffusion model to generate videos that are visually coherent and physically grounded in the captured events. Through extensive experimentation, we demonstrate the capability of our approach to synthesize high-quality videos that effectively ``rewind'' time, showcasing the potential of combining event camera technology with generative models. Our work opens new avenues for research at the intersection of computer vision, computational photography, and generative modeling, offering a forward-thinking solution to capturing missed moments and enhancing future consumer cameras and smartphones. Please see the project page at https://timerewind.github.io/ for video results and code release.
Forward citations
Cited by 2 Pith papers
-
Learning Normal Flow Directly From Event Neighborhoods
A point-based network learns per-event normal flow from raw event camera data and, with IMU data, estimates egomotion; it transfers across datasets better than frame-based optical flow methods.
-
Repurposing Pre-trained Video Diffusion Models for Event-based Video Interpolation
RE-VDM adapts the frozen Stable Video Diffusion model to event-based video frame interpolation via event-conditioned ControlNet blocks and two-sided latent fusion, reporting state of the art results on BS-ERGB, HQF, a...
Discussion (0). Continue with ORCID to comment.