Pith. sign in

REVIEW 1 cited by

PyKaldi2: Yet another speech toolkit based on Kaldi and PyTorch

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1907.05955 v3 pith:R63U3WGV submitted 2019-07-12 cs.CL eess.AS

PyKaldi2: Yet another speech toolkit based on Kaldi and PyTorch

classification cs.CL eess.AS
keywords pykaldi2trainingfeaturetoolkitfront-endimplementedkaldimodel
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

We introduce PyKaldi2 speech recognition toolkit implemented based on Kaldi and PyTorch. While similar toolkits are available built on top of the two, a key feature of PyKaldi2 is sequence training with criteria such as MMI, sMBR and MPE. In particular, we implemented the sequence training module with on-the-fly lattice generation during model training in order to simplify the training pipeline. To address the challenging acoustic environments in real applications, PyKaldi2 also supports on-the-fly noise and reverberation simulation to improve the model robustness. With this feature, it is possible to backpropogate the gradients from the sequence-level loss to the front-end feature extraction module, which, hopefully, can foster more research in the direction of joint front-end and backend learning. We performed benchmark experiments on Librispeech, and show that PyKaldi2 can achieve reasonable recognition accuracy. The toolkit is released under the MIT license.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Improving multichannel speech enhancement through accurate room-acoustic simulations

    eess.AS 2026-06 unverdicted novelty 5.0

    High-fidelity wave-based and hybrid room-acoustic simulations for training-data augmentation improve SpatialNet multichannel speech enhancement, delivering up to 38% relative median WER reduction on measured data vers...