Pith. sign in

REVIEW 1 cited by

The NPU-ASLP-LiAuto System Description for Visual Speech Recognition in CNVSRC 2023

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.06788 v2 pith:HC6TJNKE submitted 2024-01-07 eess.AS cs.AIcs.SD

classification eess.AScs.AIcs.SD
keywords taskvisualrecognitionspeechsystemcnvsrcdatafirst
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper delineates the visual speech recognition (VSR) system introduced by the NPU-ASLP-LiAuto (Team 237) in the first Chinese Continuous Visual Speech Recognition Challenge (CNVSRC) 2023, engaging in the fixed and open tracks of Single-Speaker VSR Task, and the open track of Multi-Speaker VSR Task. In terms of data processing, we leverage the lip motion extractor from the baseline1 to produce multi-scale video data. Besides, various augmentation techniques are applied during training, encompassing speed perturbation, random rotation, horizontal flipping, and color transformation. The VSR model adopts an end-to-end architecture with joint CTC/attention loss, comprising a ResNet3D visual frontend, an E-Branchformer encoder, and a Transformer decoder. Experiments show that our system achieves 34.76% CER for the Single-Speaker Task and 41.06% CER for the Multi-Speaker Task after multi-system fusion, ranking first place in all three tracks we participate.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CNVSRC 2024: The Second Chinese Continuous Visual Speech Recognition Challenge

    cs.CV 2025-05 conditional novelty 5.0 of 10

    CNVSRC 2024 lowers the baseline character error rate for Chinese visual speech recognition from 48.6% to 39.7% (single-speaker) and from 58.4% to 52.2% (multi-speaker), while adding a 200-hour dataset and documenting ...

Pith tools