Pith. sign in

REVIEW 6 cited by

TS3-Codec: Transformer-Based Simple Streaming Single Codec

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.18803 v2 pith:GONIJGYN submitted 2024-11-27 eess.AS

classification eess.AS
keywords codects3-codecaudioconvolution-basedlayersstreamingtransformer-basedarchitecture
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Neural audio codecs (NACs) have garnered significant attention as key technologies for audio compression as well as audio representation for speech language models. While mainstream NAC models are predominantly convolution-based, the performance of NACs with a purely transformer-based, and convolution-free architecture remains unexplored. This paper introduces TS3-Codec, a Transformer-Based Simple Streaming Single Codec. TS3-Codec consists of only a stack of transformer layers with a few linear layers, offering greater simplicity and expressiveness by fully eliminating convolution layers that require careful hyperparameter tuning and large computations. Under the streaming setup, the proposed TS3-Codec achieves comparable or superior performance compared to the codec with state-of-the-art convolution-based architecture while requiring only 12% of the computation and 77% of bitrate. Furthermore, it significantly outperforms the convolution-based codec when using similar computational resources.

Discussion (0). Sign in to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Unlocking Temporal Flexibility: Neural Speech Codec with Variable Frame Rate

    eess.AS 2025-05 conditional novelty 7.0 of 10

    A neural speech codec that dynamically varies frame rate per segment using waveform entropy achieves competitive or better reconstruction quality at lower average frame rates than constant frame rate baselines.

  2. ReGen: Hierarchical Multi-Prompt Representation Generation for Efficient Waveform Diffusion Models

    cs.SD 2026-07 conditional novelty 6.5 of 10

    Hierarchical multi-prompt representation generation plus generalized flow matching yields high-quality single-stage waveform diffusion from 12.5 Hz latents and efficient LDM TTS.

  3. The Equalizer: Introducing Shape-Gain Decomposition in Neural Audio Codecs

    cs.SD 2026-02 conditional novelty 6.0 of 10

    Adding classical shape-gain decomposition to a neural audio codec makes it invariant to input gain and improves bitrate-distortion performance.

  4. Memory Efficient Audio Synthesis with Decoupled Temporal Depth Diffusion Transformers

    cs.SD 2026-07 conditional novelty 5.5 of 10

    A streaming encoder plus temporal and fully shared DiT-conditioned depth decoders converts semantic audio tokens to RVQ with constant memory and ~16× real-time on-device synthesis.

  5. NanoCodec: Towards High-Quality Ultra Fast Speech LLM Inference

    eess.AS 2025-08 conditional novelty 5.0 of 10

    NanoCodec achieves competitive speech quality at 12.5 frames per second and 0.6-1.78 kbps, with a causal decoder for low-latency speech LLM inference.

  6. MagiCodec: Simple Masked Gaussian-Injected Codec for High-Fidelity Reconstruction and Generation

    cs.SD 2025-05 conditional novelty 5.0 of 10

    A single-layer streaming Transformer codec with masked Gaussian noise injection during training reports state-of-the-art reconstruction and better downstream generation and understanding in 16 kHz English speech.

Pith tools