Pith. sign in

REVIEW 2 cited by

Generative Timelines for Instructed Visual Assembly

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.12293 v1 pith:S5CIFK3Q submitted 2024-11-19 cs.CV cs.HCcs.MM

classification cs.CVcs.HCcs.MM
keywords visualtimelineassemblyinputcontentinstructedinstructionslanguage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The objective of this work is to manipulate visual timelines (e.g. a video) through natural language instructions, making complex timeline editing tasks accessible to non-expert or potentially even disabled users. We call this task Instructed visual assembly. This task is challenging as it requires (i) identifying relevant visual content in the input timeline as well as retrieving relevant visual content in a given input (video) collection, (ii) understanding the input natural language instruction, and (iii) performing the desired edits of the input visual timeline to produce an output timeline. To address these challenges, we propose the Timeline Assembler, a generative model trained to perform instructed visual assembly tasks. The contributions of this work are three-fold. First, we develop a large multimodal language model, which is designed to process visual content, compactly represent timelines and accurately interpret timeline editing instructions. Second, we introduce a novel method for automatically generating datasets for visual assembly tasks, enabling efficient training of our model without the need for human-labeled data. Third, we validate our approach by creating two novel datasets for image and video assembly, demonstrating that the Timeline Assembler substantially outperforms established baseline models, including the recent GPT-4o, in accurately executing complex assembly instructions across various real-world inspired scenarios.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Autoregressive Modeling of Film with Applications in Video Montage

    cs.CV 2026-07 conditional novelty 7.0 of 10

    An autoregressive transformer with an explicit cut token and footage-constrained decoding edits raw video into sequences that people rate as better than two prior automated editing methods.

  2. EditDuet: A Multi-Agent System for Video Non-Linear Editing

    cs.CV 2025-09 conditional novelty 6.0 of 10

    EditDuet builds B-roll timelines automatically through an Editor-Critic LLM agent loop, with a GPT-4o judge that approximates human preference for edited video.

Pith tools