Pith. sign in

REVIEW 13 cited by

AMMUS : A Survey of Transformer-based Pretrained Models in Natural Language Processing

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2108.05542 v2 pith:S6UKXG3U submitted 2021-08-12 cs.CL

classification cs.CL
keywords modelsdownstreamlearningt-ptlmslanguagepretrainingself-supervisedsurvey
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Transformer-based pretrained language models (T-PTLMs) have achieved great success in almost every NLP task. The evolution of these models started with GPT and BERT. These models are built on the top of transformers, self-supervised learning and transfer learning. Transformed-based PTLMs learn universal language representations from large volumes of text data using self-supervised learning and transfer this knowledge to downstream tasks. These models provide good background knowledge to downstream tasks which avoids training of downstream models from scratch. In this comprehensive survey paper, we initially give a brief overview of self-supervised learning. Next, we explain various core concepts like pretraining, pretraining methods, pretraining tasks, embeddings and downstream adaptation methods. Next, we present a new taxonomy of T-PTLMs and then give brief overview of various benchmarks including both intrinsic and extrinsic. We present a summary of various useful libraries to work with T-PTLMs. Finally, we highlight some of the future research directions which will further improve these models. We strongly believe that this comprehensive survey paper will serve as a good reference to learn the core concepts as well as to stay updated with the recent happenings in T-PTLMs.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. The Ripple Effect: On Unforeseen Complications of Backdoor Attacks

    cs.CR 2025-05 conditional novelty 7.0 of 10

    Backdoored language models distort predictions on unrelated fine-tuned tasks, collapsing triggered inputs to a single class, and a multi-task correction reduces this distortion without hurting attack success.

  2. Semantic Homogenization in Italian Popular Music: A Diachronic Analysis

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Sanremo lyrics exhibit rising semantic homogeneity over decades, consistently recovered by full-text, portion, topic and word-level embedding analyses.

  3. SFi-Former: Sparse Flow Induced Attention for Graph Transformer

    cs.LG 2025-04 conditional novelty 6.0 of 10

    SFi-Former replaces dense graph transformer attention with sparse flows from an l1-regularized energy minimization, improving long-range graph benchmark accuracy and generalization.

  4. A Hybrid Framework for Subject Analysis: Integrating Embedding-Based Regression Models with Large Language Models

    cs.CL 2025-07 conditional novelty 5.0 of 10

    Using an ML-predicted label count to constrain LLM generation and post-editing outputs to the LCSH vocabulary lifts subject-heading prediction F1 from 0.135 to 0.300 on a 2,100-book test set.

  5. Few-shot Learning on AMS Circuits and Its Application to Parasitic Capacitance Prediction

    cs.LG 2025-07 conditional novelty 5.0 of 10

    A few-shot graph-pretraining pipeline, built from subgraph sampling and a hybrid graph transformer, predicts parasitic coupling capacitance on unseen AMS circuits with substantially lower error than prior graph baselines.

  6. SCFormer: Structured Channel-wise Transformer with Cumulative Historical State for Multivariate Time Series Forecasting

    cs.LG 2025-05 conditional novelty 5.0 of 10

    SCFormer, a channel-wise Transformer with triangular or convolutional temporal constraints and a HiPPO cumulative history state, improves forecasting on several benchmarks but not on Traffic.

  7. Brain Network Analysis Based on Fine-tuned Self-supervised Model for Brain Disease Diagnosis

    eess.IV 2025-06 conditional novelty 4.0 of 10

    An MLP adapter before a frozen pre-trained brain transformer improves AD vs NC and MCI vs NC classification on ADNI relative to BrainNetCNN and BrainGNN, but the gain over the transformer alone is untested.

  8. Synthetic Time Series Forecasting with Transformer Architectures: Extensive Simulation Benchmarks

    cs.LG 2025-05 reject novelty 4.0 of 10

    Autoformer and PatchTST outperform Informer across synthetic forecasting benchmarks, while a proposed Koopman-Transformer hybrid is illustrated on Van der Pol and Lorenz systems.

  9. Non-Stationary Time Series Forecasting Based on Fourier Analysis and Cross Attention Mechanism

    cs.LG 2025-05 reject novelty 4.0 of 10

    AEFIN combines frequency-domain decomposition, cross-attention, and a Fourier feature network to forecast non-stationary time series, with partial improvements over some baselines.

  10. IntellectSeeker: A Personalized Literature Management System with the Probabilistic Model and Large Language Model

    cs.IR 2024-12 reject novelty 4.0 of 10

    IntellectSeeker combines a fine-tuned GPT-3.5-turbo term translator and a probabilistic relevance filter for personalized academic search, reporting BLEU 0.93 and ROUGE-1 0.94 on a self-generated corpus.

  11. A Lightweight Edge-CNN-Transformer Model for Detecting Coordinated Cyber and Digital Twin Attacks in Cooperative Smart Farming

    cs.CR 2024-11 reject novelty 3.0 of 10

    A CNN-Transformer model with post-quantization compression is presented for detecting network attacks in cooperative smart farming, with claimed accuracies up to 97% on proprietary testbed data.

  12. Explainability in Practice: A Survey of Explainable NLP Across Various Domains

    cs.CL 2025-02 reject novelty 2.0 of 10

    A survey of explainable NLP across application domains, with evaluation metrics, but with several inaccurate paper-to-application mappings.

  13. Survey of different Large Language Model Architectures: Trends, Benchmarks, and Challenges

    cs.LG 2024-12 conditional

    A broad but error-prone survey of LLM and MLLM architectures, training methods, benchmarks, and challenges.

Pith tools