Pith. sign in

REVIEW 2 cited by

Can SAM Boost Video Super-Resolution?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.06524 v2 pith:X2KAQDAR submitted 2023-05-11 cs.CV

classification cs.CV
keywords framesmethodsexistinginformationmodelmodulepriorseem
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

The primary challenge in video super-resolution (VSR) is to handle large motions in the input frames, which makes it difficult to accurately aggregate information from multiple frames. Existing works either adopt deformable convolutions or estimate optical flow as a prior to establish correspondences between frames for the effective alignment and fusion. However, they fail to take into account the valuable semantic information that can greatly enhance it; and flow-based methods heavily rely on the accuracy of a flow estimate model, which may not provide precise flows given two low-resolution frames. In this paper, we investigate a more robust and semantic-aware prior for enhanced VSR by utilizing the Segment Anything Model (SAM), a powerful foundational model that is less susceptible to image degradation. To use the SAM-based prior, we propose a simple yet effective module -- SAM-guidEd refinEment Module (SEEM), which can enhance both alignment and fusion procedures by the utilization of semantic information. This light-weight plug-in module is specifically designed to not only leverage the attention mechanism for the generation of semantic-aware feature but also be easily and seamlessly integrated into existing methods. Concretely, we apply our SEEM to two representative methods, EDVR and BasicVSR, resulting in consistently improved performance with minimal implementation effort, on three widely used VSR datasets: Vimeo-90K, REDS and Vid4. More importantly, we found that the proposed SEEM can advance the existing methods in an efficient tuning manner, providing increased flexibility in adjusting the balance between performance and the number of training parameters. Code will be open-source soon.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Semantic-Guided Diffusion Model for Single-Step Image Super-Resolution

    cs.CV 2025-05 conditional novelty 5.0 of 10

    SAMSR combines SAM segmentation masks with noise shaping and pixel-wise sampling to improve the perceptual quality of single-step diffusion super-resolution, with modest gains over SinSR.

  2. Mitigating Hallucinations on Object Attributes using Multiview Images and Negative Instructions

    cs.CV 2025-01 reject novelty 5.0 of 10

    The paper introduces a face-attribute hallucination benchmark and MIAVLM, a model that fuses multiview generated images and negative-instruction training, reporting higher balanced accuracy than zero-shot baselines.

Pith tools