A two-stage CTC-trained acoustic model feeding a BLSTM classifier achieves 88.9 percent accuracy on ten Chinese dialects, beating a one-stage baseline by 10 percent.
Network structure The major network structure we use in the two-stage system can be divided to the CNN part and the RNN part, as described in Table 1
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2019 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Two-stage Training for Chinese Dialect Recognition
A two-stage CTC-trained acoustic model feeding a BLSTM classifier achieves 88.9 percent accuracy on ten Chinese dialects, beating a one-stage baseline by 10 percent.