Pith. sign in

REVIEW 2 cited by

Temporal and Contextual Transformer for Multi-Camera Editing of TV Shows

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2210.08737 v1 pith:HKD73HIP submitted 2022-10-17 cs.CV cs.AIcs.MM

classification cs.CVcs.AIcs.MM
keywords benchmarkcamerascontainscontextualeditinghourmulti-cameratemporal
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The ability to choose an appropriate camera view among multiple cameras plays a vital role in TV shows delivery. But it is hard to figure out the statistical pattern and apply intelligent processing due to the lack of high-quality training data. To solve this issue, we first collect a novel benchmark on this setting with four diverse scenarios including concerts, sports games, gala shows, and contests, where each scenario contains 6 synchronized tracks recorded by different cameras. It contains 88-hour raw videos that contribute to the 14-hour edited videos. Based on this benchmark, we further propose a new approach temporal and contextual transformer that utilizes clues from historical shots and other views to make shot transition decisions and predict which view to be used. Extensive experiments show that our method outperforms existing methods on the proposed multi-camera editing benchmark.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Generative Timelines for Instructed Visual Assembly

    cs.CV 2024-11 conditional novelty 6.0 of 10

    A fine-tuned multimodal large language model that represents visual collections and timelines as token sequences can execute natural language timeline editing instructions more accurately than GPT-4o on synthetic benchmarks.

  2. V-Trans4Style: Visual Transition Recommendation for Video Production Style Adaptation

    cs.CV 2025-01 conditional novelty 5.0 of 10

    A transformer encoder-decoder plus an inference-time style conditioning module recommends transition sequences that match a target production style, evaluated on a new style-labeled video dataset.

Pith tools