Pith. sign in

REVIEW 1 cited by

Generating Sample-Based Musical Instruments Using Neural Audio Codec Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.15641 v1 pith:PZV5KWND submitted 2024-07-22 eess.AS cs.LGcs.SD

classification eess.AScs.LGcs.SD
keywords audioinstrumentsconsistencygeneratedmusicaltimbralapproachcodec
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this paper, we propose and investigate the use of neural audio codec language models for the automatic generation of sample-based musical instruments based on text or reference audio prompts. Our approach extends a generative audio framework to condition on pitch across an 88-key spectrum, velocity, and a combined text/audio embedding. We identify maintaining timbral consistency within the generated instruments as a major challenge. To tackle this issue, we introduce three distinct conditioning schemes. We analyze our methods through objective metrics and human listening tests, demonstrating that our approach can produce compelling musical instruments. Specifically, we introduce a new objective metric to evaluate the timbral consistency of the generated instruments and adapt the average Contrastive Language-Audio Pretraining (CLAP) score for the text-to-instrument case, noting that its naive application is unsuitable for assessing this task. Our findings reveal a complex interplay between timbral consistency, the quality of generated samples, and their correspondence to the input prompt.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. TokenSynth: A Token-based Neural Synthesizer for Instrument Cloning and Text-to-Instrument

    cs.SD 2025-02 conditional novelty 5.0 of 10

    TokenSynth uses a decoder-only transformer over audio tokens, conditioned on MIDI and CLAP timbre embeddings, to perform zero-shot instrument cloning, text-to-instrument synthesis, and text-guided timbre manipulation ...

Pith tools