Efficient Compressed Wavelet Trees over Large Alphabets

Alberto Ord\'o\~nez; Francisco Claude; Gonzalo Navarro

arxiv: 1405.1220 · v1 · pith:54CPLQ6Cnew · submitted 2014-05-06 · 💻 cs.DS

Efficient Compressed Wavelet Trees over Large Alphabets

Francisco Claude , Gonzalo Navarro , Alberto Ord\'o\~nez This is my paper

classification 💻 cs.DS

keywords waveletcompressedmatrixspacetimetreealphabetslarge

0 comments

read the original abstract

The {\em wavelet tree} is a flexible data structure that permits representing sequences $S[1,n]$ of symbols over an alphabet of size $\sigma$, within compressed space and supporting a wide range of operations on $S$. When $\sigma$ is significant compared to $n$, current wavelet tree representations incur in noticeable space or time overheads. In this article we introduce the {\em wavelet matrix}, an alternative representation for large alphabets that retains all the properties of wavelet trees but is significantly faster. We also show how the wavelet matrix can be compressed up to the zero-order entropy of the sequence without sacrificing, and actually improving, its time performance. Our experimental results show that the wavelet matrix outperforms all the wavelet tree variants along the space/time tradeoff map.

This paper has not been read by Pith yet.

Efficient Compressed Wavelet Trees over Large Alphabets

discussion (0)