Pith. sign in

REVIEW 1 cited by

A Weakly Supervised Data Labeling Framework for Machine Lexical Normalization in Vietnamese Social Media

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.20467 v1 pith:VJZW2XPM submitted 2024-09-30 cs.CL cs.AI

classification cs.CLcs.AI
keywords frameworkaccuracydatalabelinglanguagemedianormalizationsocial
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This study introduces an innovative automatic labeling framework to address the challenges of lexical normalization in social media texts for low-resource languages like Vietnamese. Social media data is rich and diverse, but the evolving and varied language used in these contexts makes manual labeling labor-intensive and expensive. To tackle these issues, we propose a framework that integrates semi-supervised learning with weak supervision techniques. This approach enhances the quality of training dataset and expands its size while minimizing manual labeling efforts. Our framework automatically labels raw data, converting non-standard vocabulary into standardized forms, thereby improving the accuracy and consistency of the training data. Experimental results demonstrate the effectiveness of our weak supervision framework in normalizing Vietnamese text, especially when utilizing Pre-trained Language Models. The proposed framework achieves an impressive F1-score of 82.72% and maintains vocabulary integrity with an accuracy of up to 99.22%. Additionally, it effectively handles undiacritized text under various conditions. This framework significantly enhances natural language normalization quality and improves the accuracy of various NLP tasks, leading to an average accuracy increase of 1-3%.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ViSoLex: An Open-Source Repository for Vietnamese Social Media Lexical Normalization

    cs.CL 2025-01 conditional novelty 4.0 of 10

    An open-source Vietnamese social media normalizer adds multitask NSW detection and dictionary lookup, reporting small F1 gains that are partly contradicted by its own table.

Pith tools