REVIEW 2 cited by
Video Summarization Techniques: A Comprehensive Review
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
The rapid expansion of video content across a variety of industries, including social media, education, entertainment, and surveillance, has made video summarization an essential field of study. The current work is a survey that explores the various approaches and methods created for video summarizing, emphasizing both abstractive and extractive strategies. The process of extractive summarization involves the identification of key frames or segments from the source video, utilizing methods such as shot boundary recognition, and clustering. On the other hand, abstractive summarization creates new content by getting the essential content from the video, using machine learning models like deep neural networks and natural language processing, reinforcement learning, attention mechanisms, generative adversarial networks, and multi-modal learning. We also include approaches that incorporate the two methodologies, along with discussing the uses and difficulties encountered in real-world implementations. The paper also covers the datasets used to benchmark these techniques. This review attempts to provide a state-of-the-art thorough knowledge of the current state and future directions of video summarization research.
Forward citations
Cited by 2 Pith papers
-
Characterizing Collective Efforts in Content Sharing and Quality Control for ADHD-relevant Content on Video-sharing Platforms
A mixed-method study of 373 ADHD videos finds YouTube content skews professional and TikTok personal, with viewers and creators collectively policing quality and accessibility, though these efforts remain incomplete.
-
Multimodal Non-Semantic Feature Fusion for Predicting Segment Access Frequency in Lecture Archives
The paper claims that a lightweight fusion of gesture, audio, and slide features can predict segment access frequency in lecture archives, but the experiments supporting these numbers are not included in this preprint.
Discussion (0). Continue with ORCID to comment.