REVIEW 5 cited by
ZEN 2.0: Continue Training and Adaption for N-gram Enhanced Text Encoders
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Pre-trained text encoders have drawn sustaining attention in natural language processing (NLP) and shown their capability in obtaining promising results in different tasks. Recent studies illustrated that external self-supervised signals (or knowledge extracted by unsupervised learning, such as n-grams) are beneficial to provide useful semantic evidence for understanding languages such as Chinese, so as to improve the performance on various downstream tasks accordingly. To further enhance the encoders, in this paper, we propose to pre-train n-gram-enhanced encoders with a large volume of data and advanced techniques for training. Moreover, we try to extend the encoder to different languages as well as different domains, where it is confirmed that the same architecture is applicable to these varying circumstances and new state-of-the-art performance is observed from a long list of NLP tasks across languages and domains.
Forward citations
Cited by 5 Pith papers
-
Balanced Training Data Augmentation for Aspect-Based Sentiment Analysis
DPO-optimized LLM data augmentation with label balancing improves ABSA accuracy and F1 on most English benchmarks, but the balancing benefit is inconsistent.
-
ChiMed 2.0: Advancing Chinese Medical Dataset in Facilitating Large Language Modeling
ChiMed 2.0 is a 204.4M-character Chinese medical dataset spanning pretraining, SFT, and preference data that yields small gains on CMMLU and CEval medical subsets.
-
Computed Tomography Visual Question Answering with Cross-modal Feature Graphing
A cross-modal graph connecting CT slices and question tokens, aggregated by an attentive GCN, improves LLM-based CT visual question answering on M3D-VQA.
-
Detoxification of Large Language Models through Output-layer Fusion with a Calibration Model
A small calibration model trained on non-toxic text is aligned and fused into the final layer of LLaMA-2-based LLMs, modestly reducing toxicity on RealToxicityPrompts but with mixed perplexity results.
-
Representation Decomposition for Learning Similarity and Contrastness Across Modalities for Affective Computing
A method that decomposes CLIP image and text features into a shared low-rank component and modality-specific sparse components, then uses an attention-weighted soft prompt to guide an LLM for sentiment, emotion, and h...
Discussion (0). Sign in to comment.