Pith. sign in

REVIEW 1 cited by

AutoAD-Zero: A Training-Free Framework for Zero-Shot Audio Description

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.15850 v2 pith:323L7T2P submitted 2024-07-22 cs.CV

classification cs.CV
keywords audiomodelsautoad-zerodescriptiongenerateinformationmoviesseries
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Our objective is to generate Audio Descriptions (ADs) for both movies and TV series in a training-free manner. We use the power of off-the-shelf Visual-Language Models (VLMs) and Large Language Models (LLMs), and develop visual and text prompting strategies for this task. Our contributions are three-fold: (i) We demonstrate that a VLM can successfully name and refer to characters if directly prompted with character information through visual indications without requiring any fine-tuning; (ii) A two-stage process is developed to generate ADs, with the first stage asking the VLM to comprehensively describe the video, followed by a second stage utilising a LLM to summarise dense textual information into one succinct AD sentence; (iii) A new dataset for TV audio description is formulated. Our approach, named AutoAD-Zero, demonstrates outstanding performance (even competitive with some models fine-tuned on ground truth ADs) in AD generation for both movies and TV series, achieving state-of-the-art CRITIC scores.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. HuMoCon: Concept Discovery for Human Motion Understanding

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A framework that combines explicit video-motion feature alignment with velocity-aware masked autoencoding to improve LLM-based human motion and video question answering.

Pith tools