REVIEW 2 cited by
Temporal and Contextual Transformer for Multi-Camera Editing of TV Shows
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The ability to choose an appropriate camera view among multiple cameras plays a vital role in TV shows delivery. But it is hard to figure out the statistical pattern and apply intelligent processing due to the lack of high-quality training data. To solve this issue, we first collect a novel benchmark on this setting with four diverse scenarios including concerts, sports games, gala shows, and contests, where each scenario contains 6 synchronized tracks recorded by different cameras. It contains 88-hour raw videos that contribute to the 14-hour edited videos. Based on this benchmark, we further propose a new approach temporal and contextual transformer that utilizes clues from historical shots and other views to make shot transition decisions and predict which view to be used. Extensive experiments show that our method outperforms existing methods on the proposed multi-camera editing benchmark.
Forward citations
Cited by 2 Pith papers
-
Generative Timelines for Instructed Visual Assembly
A fine-tuned multimodal large language model that represents visual collections and timelines as token sequences can execute natural language timeline editing instructions more accurately than GPT-4o on synthetic benchmarks.
-
V-Trans4Style: Visual Transition Recommendation for Video Production Style Adaptation
A transformer encoder-decoder plus an inference-time style conditioning module recommends transition sequences that match a target production style, evaluated on a new style-labeled video dataset.
Discussion (0). Continue with ORCID to comment.