Pith. sign in

REVIEW 1 cited by

A Language Model With Million Context Length For Raw Audio

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2206.08297 v3 pith:IC4J7MY2 submitted 2022-06-16 cs.SD cs.LGeess.AS

classification cs.SDcs.LGeess.AS
keywords contextdependenciesmodelingaudioscalestimecompareddataset
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Modeling long-term dependencies for audio signals is a particularly challenging problem, as even small-time scales yield on the order of a hundred thousand samples. With the recent advent of Transformers, neural architectures became good at modeling dependencies over longer time scales, but they suffered from quadratic constraints to scale them. We propose a generative auto-regressive architecture that can model audio waveforms over quite a large context, greater than 500,000 samples. Our work is adapted to learn time dependencies by learning a latent representation by a CNN front-end, and then learning dependencies over these representations using Transformer encoders, fully trained end-to-end: thereby allowing to learn representations as it deems fit for the next sample. Unlike previous works that compared different time scales to show improvement, we use a standard dataset, with the same number of parameters/context to show improvements. We achieve a state-of-the-art performance as compared to other approaches such as Wavenet, SaSHMI, and Sample-RNN on a standard dataset for modeling long-term structure. This work gives very exciting direction for the field, given improvements in context modeling that can be scaled with more data, as well as potentially better results by using billions/trillions of parameters.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music

    cs.SD 2024-12 conditional novelty 6.0 of 10

    A hybrid causal transformer that combines mel-spectrogram frames with EnCodec acoustic tokens matches or beats a 10-times larger token-only GPT on next-token likelihood for speech and music.

Pith tools