Pith. sign in

REVIEW 2 cited by

IteraTTA: An interface for exploring both text prompts and audio priors in generating music with text-to-audio models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2307.13005 v1 pith:QPARYHMN submitted 2023-07-24 eess.AS cs.AIcs.HCcs.SD

classification eess.AScs.AIcs.HCcs.SD
keywords usersaudiopromptstextmusicpriorstext-to-audioaudios
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent text-to-audio generation techniques have the potential to allow novice users to freely generate music audio. Even if they do not have musical knowledge, such as about chord progressions and instruments, users can try various text prompts to generate audio. However, compared to the image domain, gaining a clear understanding of the space of possible music audios is difficult because users cannot listen to the variations of the generated audios simultaneously. We therefore facilitate users in exploring not only text prompts but also audio priors that constrain the text-to-audio music generation process. This dual-sided exploration enables users to discern the impact of different text prompts and audio priors on the generation results through iterative comparison of them. Our developed interface, IteraTTA, is specifically designed to aid users in refining text prompts and selecting favorable audio priors from the generated audios. With this, users can progressively reach their loosely-specified goals while understanding and exploring the space of possible results. Our implementation and discussions highlight design considerations that are specifically required for text-to-audio models and how interaction techniques can contribute to their effectiveness.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. From Prompting to Describing: A Cross-Cultural Study of Language for AI-Generated Music

    cs.SD 2026-08 conditional novelty 6.0 of 10

    Prompts for AI music are dominated by genre and story terms, but genre words survive into perception while story-heavy prompts produce the largest semantic mismatch.

  2. Workflow-Based Evaluation of Music Generation Systems

    eess.AS 2025-06 conditional novelty 5.0 of 10

    A single-producer workflow evaluation of eight music AI tools finds they work as idea and sound generators but not as complete composers, and proposes a reusable framework.

Pith tools