Pith. sign in

REVIEW 1 cited by

SMITIN: Self-Monitored Inference-Time INtervention for Generative Music Transformers

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.02252 v2 pith:ECJ2Z42X submitted 2024-04-02 cs.SD eess.AS

classification cs.SDeess.AS
keywords generativeinterventionmusicaudiooutputsmitinapproachattention
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce Self-Monitored Inference-Time INtervention (SMITIN), an approach for controlling an autoregressive generative music transformer using classifier probes. These simple logistic regression probes are trained on the output of each attention head in the transformer using a small dataset of audio examples both exhibiting and missing a specific musical trait (e.g., the presence/absence of drums, or real/synthetic music). We then steer the attention heads in the probe direction, ensuring the generative model output captures the desired musical trait. Additionally, we monitor the probe output to avoid adding an excessive amount of intervention into the autoregressive generation, which could lead to temporally incoherent music. We validate our results objectively and subjectively for both audio continuation and text-to-music applications, demonstrating the ability to add controls to large generative models for which retraining or even fine-tuning is impractical for most musicians. Audio samples of the proposed intervention approach are available on our demo page http://tinyurl.com/smitin .

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. TokenSynth: A Token-based Neural Synthesizer for Instrument Cloning and Text-to-Instrument

    cs.SD 2025-02 conditional novelty 5.0 of 10

    TokenSynth uses a decoder-only transformer over audio tokens, conditioned on MIDI and CLAP timbre embeddings, to perform zero-shot instrument cloning, text-to-instrument synthesis, and text-guided timbre manipulation ...

Pith tools