Efficient Text Classification Using Tree-structured Multi-linear Principal Component Analysis

C.-C. Jay Kuo; Yuanhang Su; Yuzhong Huang

arxiv: 1801.06607 · v2 · pith:4UIJJM4Nnew · submitted 2018-01-20 · 💻 cs.CL

Efficient Text Classification Using Tree-structured Multi-linear Principal Component Analysis

Yuanhang Su , Yuzhong Huang , C.-C. Jay Kuo This is my paper

classification 💻 cs.CL

keywords textcomponentdimensionprincipaltmpcaanalysisclassificationdata

0 comments

read the original abstract

A novel text data dimension reduction technique, called the tree-structured multi-linear principal component anal- ysis (TMPCA), is proposed in this work. Being different from traditional text dimension reduction methods that deal with the word-level representation, the TMPCA technique reduces the dimension of input sequences and sentences to simplify the following text classification tasks. It is shown mathematically and experimentally that the TMPCA tool demands much lower complexity (and, hence, less computing power) than the ordinary principal component analysis (PCA). Furthermore, it is demon- strated by experimental results that the support vector machine (SVM) method applied to the TMPCA-processed data achieves commensurable or better performance than the state-of-the-art recurrent neural network (RNN) approach.

This paper has not been read by Pith yet.

Efficient Text Classification Using Tree-structured Multi-linear Principal Component Analysis

discussion (0)