REVIEW 13 cited by
AMMUS : A Survey of Transformer-based Pretrained Models in Natural Language Processing
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Transformer-based pretrained language models (T-PTLMs) have achieved great success in almost every NLP task. The evolution of these models started with GPT and BERT. These models are built on the top of transformers, self-supervised learning and transfer learning. Transformed-based PTLMs learn universal language representations from large volumes of text data using self-supervised learning and transfer this knowledge to downstream tasks. These models provide good background knowledge to downstream tasks which avoids training of downstream models from scratch. In this comprehensive survey paper, we initially give a brief overview of self-supervised learning. Next, we explain various core concepts like pretraining, pretraining methods, pretraining tasks, embeddings and downstream adaptation methods. Next, we present a new taxonomy of T-PTLMs and then give brief overview of various benchmarks including both intrinsic and extrinsic. We present a summary of various useful libraries to work with T-PTLMs. Finally, we highlight some of the future research directions which will further improve these models. We strongly believe that this comprehensive survey paper will serve as a good reference to learn the core concepts as well as to stay updated with the recent happenings in T-PTLMs.
Forward citations
Cited by 13 Pith papers
-
The Ripple Effect: On Unforeseen Complications of Backdoor Attacks
Backdoored language models distort predictions on unrelated fine-tuned tasks, collapsing triggered inputs to a single class, and a multi-task correction reduces this distortion without hurting attack success.
-
Semantic Homogenization in Italian Popular Music: A Diachronic Analysis
Sanremo lyrics exhibit rising semantic homogeneity over decades, consistently recovered by full-text, portion, topic and word-level embedding analyses.
-
SFi-Former: Sparse Flow Induced Attention for Graph Transformer
SFi-Former replaces dense graph transformer attention with sparse flows from an l1-regularized energy minimization, improving long-range graph benchmark accuracy and generalization.
-
A Hybrid Framework for Subject Analysis: Integrating Embedding-Based Regression Models with Large Language Models
Using an ML-predicted label count to constrain LLM generation and post-editing outputs to the LCSH vocabulary lifts subject-heading prediction F1 from 0.135 to 0.300 on a 2,100-book test set.
-
Few-shot Learning on AMS Circuits and Its Application to Parasitic Capacitance Prediction
A few-shot graph-pretraining pipeline, built from subgraph sampling and a hybrid graph transformer, predicts parasitic coupling capacitance on unseen AMS circuits with substantially lower error than prior graph baselines.
-
SCFormer: Structured Channel-wise Transformer with Cumulative Historical State for Multivariate Time Series Forecasting
SCFormer, a channel-wise Transformer with triangular or convolutional temporal constraints and a HiPPO cumulative history state, improves forecasting on several benchmarks but not on Traffic.
-
Brain Network Analysis Based on Fine-tuned Self-supervised Model for Brain Disease Diagnosis
An MLP adapter before a frozen pre-trained brain transformer improves AD vs NC and MCI vs NC classification on ADNI relative to BrainNetCNN and BrainGNN, but the gain over the transformer alone is untested.
-
Synthetic Time Series Forecasting with Transformer Architectures: Extensive Simulation Benchmarks
Autoformer and PatchTST outperform Informer across synthetic forecasting benchmarks, while a proposed Koopman-Transformer hybrid is illustrated on Van der Pol and Lorenz systems.
-
Non-Stationary Time Series Forecasting Based on Fourier Analysis and Cross Attention Mechanism
AEFIN combines frequency-domain decomposition, cross-attention, and a Fourier feature network to forecast non-stationary time series, with partial improvements over some baselines.
-
IntellectSeeker: A Personalized Literature Management System with the Probabilistic Model and Large Language Model
IntellectSeeker combines a fine-tuned GPT-3.5-turbo term translator and a probabilistic relevance filter for personalized academic search, reporting BLEU 0.93 and ROUGE-1 0.94 on a self-generated corpus.
-
A Lightweight Edge-CNN-Transformer Model for Detecting Coordinated Cyber and Digital Twin Attacks in Cooperative Smart Farming
A CNN-Transformer model with post-quantization compression is presented for detecting network attacks in cooperative smart farming, with claimed accuracies up to 97% on proprietary testbed data.
-
Explainability in Practice: A Survey of Explainable NLP Across Various Domains
A survey of explainable NLP across application domains, with evaluation metrics, but with several inaccurate paper-to-application mappings.
-
Survey of different Large Language Model Architectures: Trends, Benchmarks, and Challenges
A broad but error-prone survey of LLM and MLLM architectures, training methods, benchmarks, and challenges.
Discussion (0). Continue with ORCID to comment.