REVIEW 1 cited by
Towards Online Real-Time Memory-based Video Inpainting Transformers
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Video inpainting tasks have seen significant improvements in recent years with the rise of deep neural networks and, in particular, vision transformers. Although these models show promising reconstruction quality and temporal consistency, they are still unsuitable for live videos, one of the last steps to make them completely convincing and usable. The main limitations are that these state-of-the-art models inpaint using the whole video (offline processing) and show an insufficient frame rate. In our approach, we propose a framework to adapt existing inpainting transformers to these constraints by memorizing and refining redundant computations while maintaining a decent inpainting quality. Using this framework with some of the most recent inpainting models, we show great online results with a consistent throughput above 20 frames per second. The code and pretrained models will be made available upon acceptance.
Forward citations
Cited by 1 Pith paper
-
A Real-Time Diminished Reality Approach to Privacy in MR Collaboration
A single-camera pipeline combining YOLOv11 detection with a real-time TensorRT-optimized DSTT inpainting model demonstrates object-level privacy redaction in MR collaboration at over 20 fps, with the caveat that depth...
Discussion (0). Sign in to comment.