Pith. sign in

REVIEW 3 cited by

AudioTime: A Temporally-aligned Audio-text Benchmark Dataset

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.02857 v1 pith:TY2VNF3S submitted 2024-07-03 cs.SD eess.AS

classification cs.SDeess.AS
keywords temporalaudiocontrolmodelsaudio-textaudiotimetemporally-alignedannotations
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent advancements in audio generation have enabled the creation of high-fidelity audio clips from free-form textual descriptions. However, temporal relationships, a critical feature for audio content, are currently underrepresented in mainstream models, resulting in an imprecise temporal controllability. Specifically, users cannot accurately control the timestamps of sound events using free-form text. We acknowledge that a significant factor is the absence of high-quality, temporally-aligned audio-text datasets, which are essential for training models with temporal control. The more temporally-aligned the annotations, the better the models can understand the precise relationship between audio outputs and temporal textual prompts. Therefore, we present a strongly aligned audio-text dataset, AudioTime. It provides text annotations rich in temporal information such as timestamps, duration, frequency, and ordering, covering almost all aspects of temporal control. Additionally, we offer a comprehensive test set and evaluation metric to assess the temporal control performance of various models. Examples are available on the https://zeyuxie29.github.io/AudioTime/

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Auto-AEG: Scalable Data Construction for Open-Vocabulary Audio Event Grounding

    cs.SD 2026-07 conditional novelty 6.0 of 10

    Automatically constructed synthetic exact-GT data plus multi-model pseudo-labels with interval-aware GRPO rewards improve LALM open-vocabulary audio event grounding on AEGBench and DESED.

  2. Auto-AEG: Scalable Data Construction for Open-Vocabulary Audio Event Grounding

    cs.SD 2026-07 conditional novelty 6.0 of 10

    Auto-AEG constructs audio event grounding supervision from synthetic and pseudo-labeled real audio and uses RL fine-tuning to improve open-vocabulary temporal localization by 73.9%/23.1% mIoU over zero-shot on the new...

  3. AI-Based Sound Effect Generation: A Narrative Review of Generative Models Across Input Modalities

    cs.SD 2026-08 conditional novelty 3.0 of 10

    A narrative review of 30 recent papers classifies AI sound-effect generators by input modality and summarizes reported progress and remaining limitations in temporal sync, evaluation, and controllability.

Pith tools