Pith. sign in

REVIEW 3 cited by

Zelda: Video Analytics using Vision-Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.03785 v2 pith:DLSD7PYK submitted 2023-05-05 cs.DB

classification cs.DB
keywords zeldavideovlmsanalyticsqueryresultsaccuracydatasets
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Advances in ML have motivated the design of video analytics systems that allow for structured queries over video datasets. However, existing systems limit query expressivity, require users to specify an ML model per predicate, rely on complex optimizations that trade off accuracy for performance, and return large amounts of redundant and low-quality results. This paper focuses on the recently developed Vision-Language Models (VLMs) that allow users to query images using natural language like "cars during daytime at traffic intersections." Through an in-depth analysis, we show VLMs address three limitations of current video analytics systems: general expressivity, a single general purpose model to query many predicates, and are both simple and fast. However, VLMs still return large numbers of redundant and low-quality results that can overwhelm and burden users. In addition, VLMs often require manual prompt engineering to improve result relevance. We present Zelda: a video analytics system that uses VLMs to return both relevant and semantically diverse results for top-K queries on large video datasets. Zelda prompts the VLM with the user's query in natural language. Zelda then automatically adds discriminator and synonym terms to boost accuracy, and terms to identify low-quality frames. To improve result diversity, Zelda uses semantic-rich VLM embeddings in an algorithm that prunes similar frames while considering their relevance to the query and the number of top-K results requested. We evaluate Zelda across five datasets and 19 queries and quantitatively show it achieves higher mean average precision (up to 1.15x) and improves average pairwise similarity (up to 1.16x) compared to using VLMs out-of-the-box. We also compare Zelda to a state-of-the-art video analytics engine and show that Zelda retrieves results 7.5x (up to 10.4x) faster for the same accuracy and frame diversity.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LAVA: Language Driven Scalable and Versatile Traffic Video Analytics

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A natural-language video analytics system combining bandit-based sampling, open-vocabulary detection, and trajectory linking reports higher query accuracy than closed-world baselines on a new 18-predicate traffic benchmark.

  2. D\'ej\`a Vu: Efficient Video-Language Query Engine with Learning-based Inter-Frame Computation Reuse

    cs.DC 2025-06 conditional novelty 6.0 of 10

    A learned per-token reuse mechanism plus GPU-friendly memory and compute compaction accelerates ViT-based video embedding generation by up to 2.64x while keeping end-task accuracy within 2% of the original model.

  3. LOVO: Efficient Complex Object Query in Large-Scale Video Datasets

    cs.IR 2025-07 conditional novelty 5.0 of 10

    LOVO builds a one-time patch-embedding index over video keyframes and uses ANN search plus cross-modal rerank to answer complex object-text queries in large videos.

Pith tools