REVIEW 5 major objections 6 minor 299 references
Survey of different Large Language Model Architectures: Trends, Benchmarks, and Challenges
T0 review · 5 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read This survey claims that the evolution of large language models can be comprehensively mapped through three Transformer-based architectural families—encoder-only, decoder-only, and encoder-decoder—and provides a comparative analysis of…
desk verdict A useful-in-principle LLM survey, but the factual errors and duplicated copy make it unreliable in its current form. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The organizing machinery is the three-way architectural taxonomy built on the Transformer foundation: auto-encoding (encoder-only) models such as BERT, ERNIE, and ALBERT; auto-regressive (decoder-only) models such as GPT, LLaMA, PaLM, and KOSMOS-1; and sequence-to-sequence (encoder-decoder) models such as T5, BART, Pangu, and GLM. The Transformer—a neural architecture using multi-head self-attention and positional encoding—is treated as the shared substrate, and each family is defined by which part of the Transformer it keeps and what training objective it uses. This taxonomy carries the survey's comparative analysis: benchmark tables and lineage diagrams are organized by family, and fine-tuning and compression techniques are discussed as ways of adapting models within a family.
What would settle it
Spot-check at least ten model descriptions from the survey against the cited technical reports; for example, verify the release year, parameter count, and architectural details claimed for LLaMA, Megatron, and KOSMOS-1. Any material mismatch, such as dating LLaMA to 2022 when its cited source describes the 2023 LLaMA-2 model, would indicate that the survey's secondary summaries cannot be trusted without independent verification.
Extended reading notes
Core claim
On its own terms, the paper discovers nothing new about language models; its claim is that the recent history of LLMs is now mature enough to be summarized in a coherent narrative, and that the right narrative is an architectural one. Grouping models by encoder-only, decoder-only, and encoder-decoder designs, the survey traces evolution from 2018's GPT and BERT through the open-weights LLaMA family, the Pathways-based PaLM series, and the multimodal GPT-4, KOSMOS-1, and Gemini models. It further claims that standardized benchmarks—MMLU, SuperGLUE, HellaSwag, ARC, WinoGrande for language, NLVR2 and VQA for vision-language—plus a taxonomy of fine-tuning methods (LoRA and parameter-efficient techniques) and challenges (data quality, compression, distributed training, multimodality) provide a reliable basis for comparing models and guiding practice. The contribution is the synthesis and the comparative framing, not a new model or algorithm.
Load-bearing premise
The survey's value rests on the assumption that its selection of models, benchmarks, and methods is representative and that the descriptions of cited works faithfully reflect their primary sources; if either fails, the overview can mislead despite its breadth.
Editorial extensions
If this is right
- A reader can classify any new LLM release by its architectural family and predict its likely strengths and trade-offs, since the survey links family to task suitability.
- Practitioners can use the benchmark figures (MMLU, HellaSwag, ARC, WinoGrande, NLVR2, VQA) to compare models released in different years without re-running evaluations.
- The survey's catalog of LoRA and other parameter-efficient fine-tuning methods offers concrete, lower-cost routes for adapting large models to specialized tasks.
- The explicit list of challenges—data quality and bias, model compression, distributed computation, and multimodal alignment—marks where near-term research effort is most needed.
Reading between the lines
- Editorial inference: the survey's own comparison tables suggest that benchmark scores, not parameter counts, are becoming the field's common currency, but the paper does not argue for any particular evaluation protocol.
- Editorial inference: because the paper includes a few clear factual slips (for example, dating LLaMA to 2022 and citing the LLaMA-2 paper for LLaMA-1, and duplicating the Megatron description), a careful reader should treat specific numbers and dates as leads to check against primary sources rather than as authoritative.
- Editorial inference: the architectural taxonomy could be extended to organize future model families, but doing so would require explicit inclusion criteria, which the survey does not provide.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper is a survey of Large Language Model (LLM) and Multimodal LLM (MLLM) research, aiming to cover architectures (encoder-only, decoder-only, encoder-decoder), benchmark evaluations, fine-tuning and pre-training methods, applications, and challenges, with coverage of models up to mid-2024 and a reference list of 427 entries. It positions itself as a comprehensive, holistic map of the field, explicitly claiming in Section III to delve deeply into Architecture, Benchmark, and Challenges aspects.
Significance. If the survey were factually reliable, its breadth—spanning model evolution, PEFT taxonomies, benchmark comparisons, and challenge taxonomies—would make it a useful entry point for researchers and practitioners. The manuscript has genuine strengths: it compiles a large reference base, includes a comparative table of earlier surveys (Table 3), and offers a structured taxonomy of parameter-efficient fine-tuning methods (Table 5). However, the paper's value rests entirely on the accuracy and representativeness of its content, and the demonstrated factual errors, duplicated passages, and unverifiable benchmark figures currently undermine that value. The errors are correctable within the scope of a survey, so the appropriate path is a major revision rather than rejection, but the revision must be substantive rather than cosmetic.
major comments (5)
- [§VI-B-5] The section states that 'The LLaMA model was introduced by Meta in 2022' and cites reference [10], which is the LLaMA-2 paper, for claims about the original LLaMA. LLaMA-1 was released in February 2023 and has its own technical report; conflating the two releases corrupts the historical timeline that the survey claims to provide. This needs correction, including separate references for LLaMA-1 and LLaMA-2.
- [§VI-B-4] The Megatron paragraph is duplicated nearly verbatim: the text beginning 'Nvidia Megatron [220] is a framework proposed by Nvidia...' appears twice, with the second copy mislabeled as following from 'ChatGPT Nvidia's Megatron'. The two copies also give contradictory definitions: the first says inter-layer parallel is also known as tensor parallel, while the second says intra-layer parallelism is also known as tensor parallelism. This is a load-bearing editing failure in a technical survey, and the passage must be rewritten with a single, correct definition of tensor, pipeline, and data parallelism.
- [§VI-B-1, §VI-C-1, §VI-B-7] Multiple model descriptions contain factual errors that are not isolated typos. In §VI-B-1, GPT is described as having 'employed a 12-layer transformer encoder,' but GPT is decoder-only, contradicting the paper's own classification in §II-C. In §VI-C-1, BART is attributed to 'the Google research team,' whereas BART is from Facebook AI Research. In §VI-B-7, CogView is described as 'a 4T parameter Chinese multimodal LLM' (CogView is a 4B-parameter text-to-image model), and Mistral 7B is said to 'use mix-of-expert,' which is actually a property of Mixtral, not Mistral 7B. Together with the LLaMA error, these indicate a systemic reliability problem in the model descriptions, which is the core content of a survey.
- [§IV-A, Figures 5–11] The benchmark performance figures plot named models (MMLU, HellaSwag, ARC, WinoGrande, NLVR2, VQA) without any accompanying data tables, source citations, or evaluation-protocol details. The reader cannot verify the plotted values, determine whether they come from the original benchmark papers or secondary leaderboards, or assess comparability across models with different prompting and few-shot settings. Since the paper explicitly claims in §III to delve deeply into benchmarking, these unsupported figures are load-bearing for that claim and must be replaced or supplemented with tables that give exact scores, sources, and protocols.
- [§III, §IV] The paper does not state inclusion criteria for the models, benchmarks, or prior surveys it discusses. The claim of being 'comprehensive' and 'holistic' (abstract and §III) is therefore unsupported by any reproducible methodology. To make the survey's coverage verifiable, the authors should specify how models and benchmark results were selected (e.g., date ranges, release venues, availability of primary sources) and how the comparative table of prior surveys (Table 3) was compiled.
minor comments (6)
- [§VI-B-5, References] Reference [10] is cited for both LLaMA and LLaMA-2 claims; these should be separate references, with the original LLaMA report cited for LLaMA-1.
- [§II-A, Equations (1)–(2)] The positional-encoding formulas are printed as '100002i/dim' in the denominator; this should be typeset as 10000^(2i/dim) to be unambiguous.
- [§VIII-A-3, §VIII-B-1] Cross-references are incorrect: the text says 'figure 15 shows the common source of the datasets' but the relevant figure is Figure 28, and 'Table 7 shows transformer based pruning technology' but the pruning methods are in Table 9.
- [§VI-B-4] The phrase 'ChatGPT Nvidia's Megatron' appears in the running text as a leftover editing artifact and should be removed.
- [Throughout] There are numerous typos and terminological inconsistencies, including 'Megatron-Tuiring NLG' (§VI-B-4), 'the the datasets' (§VIII-A-3), and inconsistent spelling of 'auto-regressive' vs. 'autoregressive' and 'LLaMA' vs. 'Llama'.
- [Figures 12–14] The evolutionary tree and parameter-count figures lack explicit sources and some lack axis labels (e.g., Figure 13 uses a log scale without stating the base); these should be clarified for a reader to interpret the data.
Circularity Check
No circularity: the paper is a survey with no derived predictions, fitted parameters, or load-bearing self-citation chain.
full rationale
This manuscript is a literature survey of LLM and MLLM architectures, benchmarks, fine-tuning techniques, challenges, and applications. It does not propose a new model, derive a mathematical result, fit parameters to data, or make a prediction that is then validated against the same data. The central claims are descriptive: that the survey organizes prior work and that its coverage is comprehensive. Those claims rest on the accuracy and representativeness of the cited sources, not on any equation that defines a target quantity in terms of the paper's own outputs. There is no self-definitional step, no fitted input renamed as a prediction, and no invoked uniqueness theorem or ansatz imported from prior work by the same authors. The paper's known weaknesses—such as the LLaMA 1/LLaMA 2 dating and citation conflation in Section VI-B-5, the duplicated Megatron paragraph in Section VI-B-4, and unsourced benchmark figures—are factual reliability and correctness risks, not circular reasoning. A survey can be inaccurate without being circular. Accordingly, no circularity steps are identified, and the score is 0.
Assumptions & free parameters
assumptions (3)
- domain assumption The taxonomy of LLMs into encoder-only, decoder-only, and encoder-decoder architectures is complete enough to organize all relevant modern models.
- domain assumption The benchmarks selected in Section IV (MMLU, SuperGLUE, HellaSwag, ARC, WinoGrande, NLVR2, VQA) are representative of how LLM capability should be measured.
- domain assumption Figures 5 through 14 display correct scores sourced from credible leaderboards or papers, meaning the displayed data accurately represent the cited models.
Cite this review
Pith. "Pith review of Survey of different Large Language Model Architectures: Trends, Benchmarks, and Challenges." pith.science (2026). https://pith.science/paper/ST4DGU6P
@misc{pith2026241203220,
author = {Pith},
title = {Pith review of: Survey of different Large Language Model Architectures: Trends, Benchmarks, and Challenges},
year = {2026},
howpublished = {\url{https://pith.science/paper/ST4DGU6P}},
note = {Machine review of arXiv:2412.03220}
}
read the original abstract
Large Language Models (LLMs) represent a class of deep learning models adept at understanding natural language and generating coherent responses to various prompts or queries. These models far exceed the complexity of conventional neural networks, often encompassing dozens of neural network layers and containing billions to trillions of parameters. They are typically trained on vast datasets, utilizing architectures based on transformer blocks. Present-day LLMs are multi-functional, capable of performing a range of tasks from text generation and language translation to question answering, as well as code generation and analysis. An advanced subset of these models, known as Multimodal Large Language Models (MLLMs), extends LLM capabilities to process and interpret multiple data modalities, including images, audio, and video. This enhancement empowers MLLMs with capabilities like video editing, image comprehension, and captioning for visual content. This survey provides a comprehensive overview of the recent advancements in LLMs. We begin by tracing the evolution of LLMs and subsequently delve into the advent and nuances of MLLMs. We analyze emerging state-of-the-art MLLMs, exploring their technical features, strengths, and limitations. Additionally, we present a comparative analysis of these models and discuss their challenges, potential limitations, and prospects for future development.
Figures
Figures from the paper (20 more)
Reference graph
Works this paper leans on
-
[10]
H. Touvron, L. Martin, K. Stone, P . Albert, A. Almahairi, Y . Babaei, N. Bashlykov, S. Batra, P . Bhargava, S. Bhosale et al. , ‘‘Llama 2: Open foundation and fine-tuned chat models,’’ arXiv preprint arXiv:2307.09288, 2023
arXiv 2023
-
[220]
[Online]
NVIDIA, ‘‘Nvidia nemo_2023,’’ Oct 2023. [Online]. Available: https://docs.nvidia.com/deeplearning/nemo/user-guide/docs/en/main/ nlp/megatron.html
2023
-
[1]
V aswani, N
A. V aswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, ‘‘Attention is all you need,’’ Advances in neural information processing systems, vol. 30, 2017
2017
-
[2]
Radford, K
A. Radford, K. Narasimhan, T. Salimans, I. Sutskever et al., ‘‘Improving language understanding by generative pre-training,’’ 2018
2018
-
[3]
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, ‘‘Bert: Pre-training of deep bidirectional transformers for language understanding,’’ arXiv preprint arXiv:1810.04805, 2018
arXiv 2018
-
[4]
W. Zeng, X. Ren, T. Su, H. Wang, Y . Liao, Z. Wang, X. Jiang, Z. Y ang, K. Wang, X. Zhang et al., ‘‘Pangu-{\alpha}: Large-scale autoregressive pretrained chinese language models with auto-parallel computation,’’ arXiv preprint arXiv:2104.12369, 2021
arXiv 2021
-
[5]
X. Ren, P . Zhou, X. Meng, X. Huang, Y . Wang, W. Wang, P . Li, X. Zhang, A. Podolskiy, G. Arshinov et al. , ‘‘Pangu-{\Sigma}: Towards trillion parameter language model with sparse heterogeneous computing,’’ arXiv preprint arXiv:2303.10845, 2023
arXiv 2023
-
[6]
Graves and A
A. Graves and A. Graves, ‘‘Long short-term memory,’’ Supervised sequence labelling with recurrent neural networks , pp. 37–45, 2012
2012
Show all 299 references
-
[7]
L. R. Medsker and L. Jain, ‘‘Recurrent neural networks,’’ Design and Applications, vol. 5, no. 64-67, p. 2, 2001
2001
-
[8]
Zhang, X
Z. Zhang, X. Han, Z. Liu, X. Jiang, M. Sun, and Q. Liu, ‘‘Ernie: Enhanced language representation with informative entities,’’ arXiv preprint arXiv:1905.07129, 2019
1905 arXiv
-
[9]
Z. Lan, M. Chen, S. Goodman, K. Gimpel, P . Sharma, and R. Soricut, ‘‘Albert: A lite bert for self-supervised learning of language representations,’’ arXiv preprint arXiv:1909.11942, 2019
1909 arXiv
-
[11]
Raffel, N
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y . Zhou, W. Li, and P . J. Liu, ‘‘Exploring the limits of transfer learning with a unified text-to-text transformer,’’ The Journal of Machine Learning Research, vol. 21, no. 1, pp. 5485–5551, 2020
2020
-
[12]
Z. Du, Y . Qian, X. Liu, M. Ding, J. Qiu, Z. Y ang, and J. Tang, ‘‘Glm: General language model pretraining with autoregressive blank infilling,’’ arXiv preprint arXiv:2103.10360, 2021. 30 VOLUME 11, 2024 Minghao et al.: Survey of different Large Language Model Architectures: T...
2021 arXiv
-
[13]
D. P . Kingma, ‘‘Auto-encoding variational bayes,’’ arXiv preprint arXiv:1312.6114, 2013
2013 arXiv
-
[14]
J. Zhai, S. Zhang, J. Chen, and Q. He, ‘‘Autoencoder and its various variants,’’ in 2018 IEEE international conference on systems, man, and cybernetics (SMC). IEEE, 2018, pp. 415–419
2018
-
[15]
R. Wei, C. Garcia, A. El-Sayed, V . Peterson, and A. Mahmood, ‘‘V ariations in variational autoencoders-a comparative evaluation,’’ Ieee Access, vol. 8, pp. 153 651–153 670, 2020
2020
-
[16]
I. J. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y . Bengio, ‘‘Generative adversarial networks,’’ Science Robotics, vol. 3, pp. 2672–2680, 6 2014
2014
-
[17]
Radford, L
A. Radford, L. Metz, and S. Chintala, ‘‘Unsupervised representation learning with deep convolutional generative adversarial networks,’’ 4th International Conference on Learning Representations, ICLR 2016 - Conference Track Proceedings, 11 2015
2016
-
[18]
Mirza and S
M. Mirza and S. Osindero, ‘‘Conditional generative adversarial nets,’’ arXiv preprint: arXiv:1411.1784, 11 2014
2014 arXiv
-
[19]
Arjovsky, S
M. Arjovsky, S. Chintala, and L. Bottou, ‘‘Wasserstein gan,’’ arXiv preprint: arXiv:1701.07875, 1 2017
2017 arXiv
-
[20]
Gulrajani, F
I. Gulrajani, F. Ahmed, M. Arjovsky, V . Dumoulin, and A. Courville, ‘‘Improved training of wasserstein gans,’’ Advances in Neural Informa- tion Processing Systems, vol. 2017-December, pp. 5768–5778, 3 2017
2017
-
[21]
X. Mao, Q. Li, H. Xie, R. Y . Lau, Z. Wang, and S. P . Smolley, ‘‘Least squares generative adversarial networks,’’ Proceedings of the IEEE International Conference on Computer Vision , vol. 2017-October, pp. 2813–2821, 11 2016
2017
-
[22]
J. Y . Zhu, T. Park, P . Isola, and A. A. Efros, ‘‘Unpaired image-to-image translation using cycle-consistent adversarial networks,’’ Proceedings of the IEEE International Conference on Computer Vision , vol. 2017-October, pp. 2242–2251, 3 2017
2017
-
[23]
Karras, S
T. Karras, S. Laine, and T. Aila, ‘‘A style-based generator architecture for generative adversarial networks,’’IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 43, pp. 4217–4228, 12 2018
2018
-
[24]
Zhang, I
H. Zhang, I. Goodfellow, D. Metaxas, and A. Odena, ‘‘Self-attention gen- erative adversarial networks,’’36th International Conference on Machine Learning, ICML 2019, vol. 2019-June, pp. 12 744–12 753, 5 2018
2019
-
[25]
Brock, J
A. Brock, J. Donahue, and K. Simonyan, ‘‘Large scale gan training for high fidelity natural image synthesis,’’ 7th International Conference on Learning Representations, ICLR 2019 , 9 2018
2019
-
[26]
Karras, T
T. Karras, T. Aila, S. Laine, and J. Lehtinen, ‘‘Progressive growing of gans for improved quality, stability, and variation,’’ 6th International Conference on Learning Representations, ICLR 2018 - Conference Track Proceedings, 10 2017
2018
-
[27]
Y . Choi, M. Choi, M. Kim, J. W. Ha, S. Kim, and J. Choo, ‘‘Stargan: Uni- fied generative adversarial networks for multi-domain image-to-image translation,’’ Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition, pp. 8789–8797, 11 2017
2017
-
[28]
Z. Wang, Q. She, and T. E. Ward, ‘‘Generative adversarial networks in computer vision: A survey and taxonomy,’’ ACM Computing Surveys , vol. 54, 6 2019
2019
-
[29]
J. Gui, Z. Sun, Y . Wen, D. Tao, and J. Y e, ‘‘A review on generative adversarial networks: Algorithms, theory, and applications,’’ IEEE Transactions on Knowledge and Data Engineering , vol. 35, pp. 3313–3332, 4 2023
2023
-
[30]
Creswell, T
A. Creswell, T. White, V . Dumoulin, K. Arulkumaran, B. Sengupta, and A. A. Bharath, ‘‘Generative adversarial networks: An overview,’’ IEEE Signal Processing Magazine, vol. 35, pp. 53–65, 10 2017
2017
-
[31]
I. O. Gallegos, R. A. Rossi, J. Barrow, M. M. Tanjim, S. Kim, F. Der- noncourt, T. Y u, R. Zhang, and N. K. Ahmed, ‘‘Bias and fairness in large language models: A survey,’’ arXiv preprint arXiv:2309.00770, 2023
2023 arXiv
-
[32]
H. Shi, Z. Xu, H. Wang, W. Qin, W. Wang, Y . Wang, and H. Wang, ‘‘Continual learning of large language models: A comprehensive survey,’’ arXiv preprint arXiv:2404.16789, 2024
2024 arXiv
-
[33]
Z. Zhou, X. Ning, K. Hong, T. Fu, J. Xu, S. Li, Y . Lou, L. Wang, Z. Y uan, X. Li et al., ‘‘A survey on efficient inference for large language models,’’ arXiv preprint arXiv:2404.14294, 2024
2024 arXiv
-
[34]
X. Fang, W. Xu, F. A. Tan, J. Zhang, Z. Hu, Y . Qi, S. Nickleach, D. Socolinsky, S. Sengamedu, and C. Faloutsos, ‘‘Large language models (llms) on tabular data: Prediction, generation, and understanding–a survey,’’arXiv preprint arXiv:2402.17944, 2024
2024 arXiv
-
[35]
B. C. Das, M. H. Amini, and Y . Wu, ‘‘Security and privacy challenges of large language models: A survey,’’ arXiv preprint arXiv:2402.00888 , 2024
2024 arXiv
-
[36]
H. Jin, L. Hu, X. Li, P . Zhang, C. Chen, J. Zhuang, and H. Wang, ‘‘Jail- breakzoo: Survey, landscapes, and horizons in jailbreaking large language and vision-language models,’’ arXiv preprint arXiv:2407.01599, 2024
2024
-
[37]
L. Wu, Z. Zheng, Z. Qiu, H. Wang, H. Gu, T. Shen, C. Qin, C. Zhu, H. Zhu, Q. Liu et al., ‘‘A survey on large language models for recommendation,’’ arXiv preprint arXiv:2305.19860, 2023
2023 arXiv
-
[38]
Akyash and H
M. Akyash and H. M. Kamali, ‘‘Evolutionary large language models for hardware security: A comparative survey,’’ arXiv preprint arXiv:2404.16651, 2024
2024 arXiv
-
[39]
T. Bai, H. Liang, B. Wan, L. Y ang, B. Li, Y . Wang, B. Cui, C. He, B. Y uan, and W. Zhang, ‘‘A survey of multimodal large language model from a data-centric perspective,’’ arXiv preprint arXiv:2405.16640, 2024
2024 arXiv
-
[40]
H. Xiao, F. Zhou, X. Liu, T. Liu, Z. Li, X. Liu, and X. Huang, ‘‘A comprehensive survey of large language models and multimodal large language models in medicine,’’ arXiv preprint arXiv:2405.08603, 2024
2024 arXiv
-
[41]
H. Zhou, C. Hu, Y . Y uan, Y . Cui, Y . Jin, C. Chen, H. Wu, D. Y uan, L. Jiang, D. Wu et al. , ‘‘Large language model (llm) for telecommunications: A comprehensive survey on principles, key techniques, and opportunities,’’ arXiv preprint arXiv:2405.10825, 2024
2024 arXiv
-
[42]
Huang, K
Y . Huang, K. Tang, and M. Chen, ‘‘A comprehensive survey on evaluating large language model applications in the medical industry,’’ arXiv preprint arXiv:2404.15777, 2024
2024 arXiv
-
[43]
C. Qu, S. Dai, X. Wei, H. Cai, S. Wang, D. Yin, J. Xu, and J.-R. Wen, ‘‘Tool learning with large language models: A survey,’’ arXiv preprint arXiv:2405.17935, 2024
2024 arXiv
-
[44]
L. Qin, Q. Chen, X. Feng, Y . Wu, Y . Zhang, Y . Li, M. Li, W. Che, and P . S. Y u, ‘‘Large language models meet nlp: A survey,’’ arXiv preprint arXiv:2405.12819, 2024
2024 arXiv
-
[45]
Zhang, Y
Z. Zhang, Y . Sun, Z. Wang, Y . Nie, X. Ma, P . Sun, and R. Li, ‘‘Large language models for mobility in transportation systems: A survey on forecasting tasks,’’ arXiv preprint arXiv:2405.02357, 2024
2024 arXiv
-
[46]
Kukreja, T
S. Kukreja, T. Kumar, A. Purohit, A. Dasgupta, and D. Guha, ‘‘A literature survey on open source large language models,’’ in Proceedings of the 2024 7th International Conference on Computers in Management and Business, 2024, pp. 133–143
2024
-
[47]
L. Qin, Q. Chen, Y . Zhou, Z. Chen, Y . Li, L. Liao, M. Li, W. Che, and P . S. Y u, ‘‘Multilingual large language model: A survey of resources, taxonomy and frontiers,’’ arXiv preprint arXiv:2404.04925, 2024
2024 arXiv
-
[48]
S. Dai, C. Xu, S. Xu, L. Pang, Z. Dong, and J. Xu, ‘‘Unifying bias and un- fairness in information retrieval: A survey of challenges and opportunities with large language models,’’ arXiv preprint arXiv:2404.11457, 2024
2024 arXiv
-
[49]
S. Yin, C. Fu, S. Zhao, K. Li, X. Sun, T. Xu, and E. Chen, ‘‘A survey on multimodal large language models,’’ arXiv preprint arXiv:2306.13549 , 2023
2023 arXiv
-
[50]
Zhang, X
Z. Zhang, X. Bo, C. Ma, R. Li, X. Chen, Q. Dai, J. Zhu, Z. Dong, and J.-R. Wen, ‘‘A survey on the memory mechanism of large language model based agents,’’ arXiv preprint arXiv:2404.13501, 2024
2024 arXiv
-
[51]
S. Hu, T. Huang, F. Ilhan, S. Tekin, G. Liu, R. Kompella, and L. Liu, ‘‘A survey on large language model-based game agents,’’ arXiv preprint arXiv:2404.02039, 2024
2024 arXiv
-
[52]
Z. Bai, P . Wang, T. Xiao, T. He, Z. Han, Z. Zhang, and M. Z. Shou, ‘‘Hallucination of multimodal large language models: A survey,’’ arXiv preprint arXiv:2404.18930, 2024
2024 arXiv
-
[53]
Huang and J
Y . Huang and J. Huang, ‘‘A survey on retrieval-augmented text generation for large language models,’’ arXiv preprint arXiv:2404.10981, 2024
2024 arXiv
-
[54]
J. Li, T. Tang, W. X. Zhao, J.-Y . Nie, and J.-R. Wen, ‘‘Pre-trained language models for text generation: A survey,’’ ACM Computing Surveys, vol. 56, no. 9, pp. 1–39, 2024
2024
-
[55]
Y . Liu, Y . Y ao, J.-F. Ton, X. Zhang, R. G. H. Cheng, Y . Klochkov, M. F. Taufiq, and H. Li, ‘‘Trustworthy llms: a survey and guideline for evaluating large language models’ alignment,’’ arXiv preprint arXiv:2308.05374, 2023
2023 arXiv
-
[56]
Y . Y ao, J. Duan, K. Xu, Y . Cai, Z. Sun, and Y . Zhang, ‘‘A survey on large language model (llm) security and privacy: The good, the bad, and the ugly,’’High-Confidence Computing, p. 100211, 2024
2024
-
[57]
Chang, X
Y . Chang, X. Wang, J. Wang, Y . Wu, L. Y ang, K. Zhu, H. Chen, X. Yi, C. Wang, Y . Wang et al. , ‘‘A survey on evaluation of large language models,’’ ACM Transactions on Intelligent Systems and Technology , vol. 15, no. 3, pp. 1–45, 2024
2024
-
[58]
Y . Gao, Y . Xiong, X. Gao, K. Jia, J. Pan, Y . Bi, Y . Dai, J. Sun, and H. Wang, ‘‘Retrieval-augmented generation for large language models: A survey,’’arXiv preprint arXiv:2312.10997, 2023. VOLUME 11, 2024 31 Minghao et al.: Survey of different Large Language Model Architect...
2023 arXiv
-
[59]
Zhang, L
S. Zhang, L. Dong, X. Li, S. Zhang, X. Sun, S. Wang, J. Li, R. Hu, T. Zhang, F. Wu et al., ‘‘Instruction tuning for large language models: A survey,’’arXiv preprint arXiv:2308.10792, 2023
2023
-
[60]
R. Hong, X. Pang, and C. Zhang, ‘‘Advances in reasoning by prompting large language models: A survey,’’ Cybernetics and Intelligence, 2024
2024
-
[61]
B. Y an, K. Li, M. Xu, Y . Dong, Y . Zhang, Z. Ren, and X. Cheng, ‘‘On protecting the data privacy of large language models (llms): A survey,’’ arXiv preprint arXiv:2403.05156, 2024
2024 arXiv
-
[62]
Y . Cao, H. Zhao, Y . Cheng, T. Shu, G. Liu, G. Liang, J. Zhao, and Y . Li, ‘‘Survey on large language model-enhanced reinforcement learning: Con- cept, taxonomy, and methods,’’ arXiv preprint arXiv:2404.00282, 2024
2024 arXiv
-
[63]
X. Liu, P . Xu, J. Wu, J. Y uan, Y . Y ang, Y . Zhou, F. Liu, T. Guan, H. Wang, T. Y uet al., ‘‘Large language models and causal inference in collabora- tion: A comprehensive survey,’’ arXiv preprint arXiv:2403.09606, 2024
2024 arXiv
-
[64]
Esmradi, D
A. Esmradi, D. W. Yip, and C. F. Chan, ‘‘A comprehensive survey of attack techniques, implementation, and mitigation strategies in large language models,’’ in International Conference on Ubiquitous Security . Springer, 2023, pp. 76–95
2023
-
[65]
A. G. Chowdhury, M. M. Islam, V . Kumar, F. H. Shezan, V . Jain, and A. Chadha, ‘‘Breaking down the defenses: A comparative survey of at- tacks on large language models,’’arXiv preprint arXiv:2403.04786, 2024
2024
-
[66]
Sun, ‘‘A short survey of viewing large language models in legal aspect,’’ arXiv preprint arXiv:2303.09136, 2023
Z. Sun, ‘‘A short survey of viewing large language models in legal aspect,’’ arXiv preprint arXiv:2303.09136, 2023
2023 arXiv
-
[67]
H. Zhao, H. Chen, F. Y ang, N. Liu, H. Deng, H. Cai, S. Wang, D. Yin, and M. Du, ‘‘Explainability for large language models: A survey,’’ ACM Transactions on Intelligent Systems and Technology , vol. 15, no. 2, pp. 1–38, 2024
2024
-
[68]
Y . Zhu, H. Y uan, S. Wang, J. Liu, W. Liu, C. Deng, Z. Dou, and J.-R. Wen, ‘‘Large language models for information retrieval: A survey,’’arXiv preprint arXiv:2308.07107, 2023
2023
-
[69]
J. Li, Y . Liu, C. Liu, L. Shi, X. Ren, Y . Zheng, Y . Liu, and Y . Xue, ‘‘A cross-language investigation into jailbreak attacks in large language models,’’ arXiv preprint arXiv:2401.16765, 2024
2024 arXiv
-
[70]
Z. Xi, W. Chen, X. Guo, W. He, Y . Ding, B. Hong, M. Zhang, J. Wang, S. Jin, E. Zhou et al. , ‘‘The rise and potential of large language model based agents: A survey,’’ arXiv preprint arXiv:2309.07864, 2023
2023 arXiv
-
[71]
Huang, W
L. Huang, W. Y u, W. Ma, W. Zhong, Z. Feng, H. Wang, Q. Chen, W. Peng, X. Feng, B. Qin et al. , ‘‘A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions,’’ arXiv preprint arXiv:2311.05232, 2023
2023 arXiv
-
[72]
Shayegani, M
E. Shayegani, M. A. A. Mamun, Y . Fu, P . Zaree, Y . Dong, and N. Abu-Ghazaleh, ‘‘Survey of vulnerabilities in large language models revealed by adversarial attacks,’’arXiv preprint arXiv:2310.10844, 2023
2023 arXiv
-
[73]
Zhang, H
H. Zhang, H. Song, S. Li, M. Zhou, and D. Song, ‘‘A survey of controllable text generation using transformer-based pre-trained language models,’’ ACM Computing Surveys, vol. 56, no. 3, pp. 1–37, 2023
2023
-
[74]
B. Min, H. Ross, E. Sulem, A. P . B. V eyseh, T. H. Nguyen, O. Sainz, E. Agirre, I. Heintz, and D. Roth, ‘‘Recent advances in natural language processing via large pre-trained language models: A survey,’’ ACM Computing Surveys, vol. 56, no. 2, pp. 1–40, 2023
2023
-
[75]
Zhang, Y
Y . Zhang, Y . Li, L. Cui, D. Cai, L. Liu, T. Fu, X. Huang, E. Zhao, Y . Zhang, Y . Chenet al., ‘‘Siren’s song in the ai ocean: a survey on hallucination in large language models,’’ arXiv preprint arXiv:2309.01219, 2023
2023 arXiv
-
[76]
X. Zhu, J. Li, Y . Liu, C. Ma, and W. Wang, ‘‘A survey on model compres- sion for large language models,’’ arXiv preprint arXiv:2308.07633, 2023
2023 arXiv
-
[77]
L. Hu, Z. Liu, Z. Zhao, L. Hou, L. Nie, and J. Li, ‘‘A survey of knowledge enhanced pre-trained language models,’’ IEEE Transactions on Knowledge and Data Engineering, vol. 36, no. 4, pp. 1413–1430, 2024
2024
-
[78]
B. Wang, Q. Xie, J. Pei, Z. Chen, P . Tiwari, Z. Li, and J. Fu, ‘‘Pre-trained language models in biomedical domain: A systematic survey,’’ ACM Computing Surveys, vol. 56, no. 3, pp. 1–52, 2023
2023
-
[79]
Huang and K
J. Huang and K. C.-C. Chang, ‘‘Towards reasoning in large language models: A survey,’’ arXiv preprint arXiv:2212.10403, 2022
2022 arXiv
-
[80]
Kasneci, K
E. Kasneci, K. Seßler, S. Küchemann, M. Bannert, D. Dementieva, F. Fischer, U. Gasser, G. Groh, S. Günnemann, E. Hüllermeier et al. , ‘‘Chatgpt for good? on opportunities and challenges of large language models for education,’’ Learning and individual differences , vol. 103, p...
2023
-
[81]
W. X. Zhao, K. Zhou, J. Li, T. Tang, X. Wang, Y . Hou, Y . Min, B. Zhang, J. Zhang, Z. Dong et al. , ‘‘A survey of large language models,’’ arXiv preprint arXiv:2303.18223, 2023
2023 arXiv
-
[82]
Mialon, R
G. Mialon, R. Dessì, M. Lomeli, C. Nalmpantis, R. Pasunuru, R. Raileanu, B. Rozière, T. Schick, J. Dwivedi-Y u, A. Celikyilmazet al., ‘‘Augmented language models: a survey,’’ arXiv preprint arXiv:2302.07842, 2023
2023 arXiv
-
[83]
K. S. Kalyan, A. Rajasekharan, and S. Sangeetha, ‘‘Ammu: a survey of transformer-based biomedical pretrained language models,’’ Journal of biomedical informatics, vol. 126, p. 103982, 2022
2022
-
[84]
——, ‘‘Ammus: A survey of transformer-based pretrained models in natural language processing,’’ arXiv preprint arXiv:2108.05542, 2021
2021 arXiv
-
[85]
M. Zaib, Q. Z. Sheng, and W. Emma Zhang, ‘‘A short survey of pre-trained language models for conversational ai-a new age in nlp,’’ in Proceedings of the Australasian computer science week multiconference , 2020, pp. 1–4
2020
-
[86]
Houlsby, A
N. Houlsby, A. Giurgiu, S. Jastrzebski, B. Morrone, Q. De Laroussilhe, A. Gesmundo, M. Attariyan, and S. Gelly, ‘‘Parameter-efficient transfer learning for nlp,’’ in International Conference on Machine Learning . PMLR, 2019, pp. 2790–2799
2019
-
[87]
J. He, C. Zhou, X. Ma, T. Berg-Kirkpatrick, and G. Neubig, ‘‘Towards a unified view of parameter-efficient transfer learning,’’ arXiv preprint arXiv:2110.04366, 2021
2021 arXiv
-
[88]
Y . Zhu, J. Feng, C. Zhao, M. Wang, and L. Li, ‘‘Counter-interference adapter for multilingual machine translation,’’ arXiv preprint arXiv:2104.08154, 2021
2021 arXiv
-
[89]
T. Lei, J. Bai, S. Brahma, J. Ainslie, K. Lee, Y . Zhou, N. Du, V . Y . Zhao, Y . Wu, B. Li et al. , ‘‘Conditional adapters: Parameter-efficient transfer learning with fast inference,’’ arXiv preprint arXiv:2304.04947, 2023
2023 arXiv
-
[90]
Pfeiffer, A
J. Pfeiffer, A. Kamath, A. Rücklé, K. Cho, and I. Gurevych, ‘‘Adapterfusion: Non-destructive task composition for transfer learning,’’ arXiv preprint arXiv:2005.00247, 2020
2005 arXiv
-
[91]
Y . Wang, S. Mukherjee, X. Liu, J. Gao, A. H. Awadallah, and J. Gao, ‘‘Adamix: Mixture-of-adapter for parameter-efficient tuning of large lan- guage models,’’arXiv preprint arXiv:2205.12410, vol. 1, no. 2, p. 4, 2022
2022 arXiv
-
[92]
H. Zhao, J. Fu, and Z. He, ‘‘Prototype-based hyperadapter for sample- efficient multi-task tuning,’’ arXiv preprint arXiv:2310.11670, 2023
2023 arXiv
-
[93]
Chronopoulou, M
A. Chronopoulou, M. E. Peters, A. Fraser, and J. Dodge, ‘‘Adaptersoup: Weight averaging to improve generalization of pretrained language models,’’ arXiv preprint arXiv:2302.07027, 2023
2023 arXiv
-
[94]
He, R.-Z
S. He, R.-Z. Fan, L. Ding, L. Shen, T. Zhou, and D. Tao, ‘‘Mera: Merging pretrained adapters for few-shot learning,’’ arXiv preprint arXiv:2308.15982, 2023
2023 arXiv
-
[95]
R. K. Mahabadi, S. Ruder, M. Dehghani, and J. Henderson, ‘‘Parameter-efficient multi-task fine-tuning for transformers via shared hypernetworks,’’arXiv preprint arXiv:2106.04489, 2021
2021 arXiv
-
[96]
X. L. Li and P . Liang, ‘‘Prefix-tuning: Optimizing continuous prompts for generation,’’ arXiv preprint arXiv:2101.00190, 2021
2021 arXiv
-
[97]
J. Li, W. Aitken, R. Bhambhoria, and X. Zhu, ‘‘Prefix propagation: Parameter-efficient tuning for long sequences,’’ arXiv preprint arXiv:2305.12086, 2023
2023 arXiv
-
[98]
X. Liu, K. Ji, Y . Fu, W. L. Tam, Z. Du, Z. Y ang, and J. Tang, ‘‘P-tuning v2: Prompt tuning can be comparable to fine-tuning universally across scales and tasks,’’ arXiv preprint arXiv:2110.07602, 2021
2021 arXiv
-
[99]
Zhang, C
Z.-R. Zhang, C. Tan, H. Xu, C. Wang, J. Huang, and S. Huang, ‘‘Towards adaptive prefix tuning for parameter-efficient language model fine-tuning,’’ arXiv preprint arXiv:2305.15212, 2023
2023 arXiv
-
[100]
X. Liu, Y . Zheng, Z. Du, M. Ding, Y . Qian, Z. Y ang, and J. Tang, ‘‘Gpt understands, too,’’ arXiv preprint arXiv:2103.10385, 2021
2021 arXiv
-
[101]
Lester, R
B. Lester, R. Al-Rfou, and N. Constant, ‘‘The power of scale for parameter-efficient prompt tuning,’’ arXiv preprint arXiv:2104.08691 , 2021
2021 arXiv
-
[102]
F. Ma, C. Zhang, L. Ren, J. Wang, Q. Wang, W. Wu, X. Quan, and D. Song, ‘‘Xprompt: Exploring the extreme of prompt tuning,’’ arXiv preprint arXiv:2210.04457, 2022
2022 arXiv
-
[103]
Z. Wu, S. Wang, J. Gu, R. Hou, Y . Dong, V . Vydiswaran, and H. Ma, ‘‘Idpg: An instance-dependent prompt generation method,’’ arXiv preprint arXiv:2204.04497, 2022
2022 arXiv
-
[104]
X. Liu, T. Sun, X. Huang, and X. Qiu, ‘‘Late prompt tuning: A late prompt could be better than many prompts,’’ arXiv preprint arXiv:2210.11292 , 2022
2022 arXiv
-
[105]
Zhu and M
W. Zhu and M. Tan, ‘‘Spt: Learning to selectively insert prompts for better prompt tuning,’’ in Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , 2023, pp. 11 862–11 878
2023
-
[106]
Q. Wang, Y . Mao, J. Wang, H. Y u, S. Nie, S. Wang, F. Feng, L. Huang, X. Quan, Z. Xu et al. , ‘‘Aprompt: Attention prompt tuning for efficient adaptation of pre-trained language models,’’ in Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processin...
2023
-
[107]
T. Vu, B. Lester, N. Constant, R. Al-Rfou, and D. Cer, ‘‘Spot: Better frozen model adaptation through soft prompt transfer,’’ arXiv preprint arXiv:2110.07904, 2021
2021 arXiv
-
[108]
Y . Su, X. Wang, Y . Qin, C.-M. Chan, Y . Lin, H. Wang, K. Wen, Z. Liu, P . Li, J. Li et al. , ‘‘On transferability of prompt tuning for natural language processing,’’ arXiv preprint arXiv:2111.06719, 2021
2021 arXiv
-
[109]
J. Wu, T. Y u, R. Wang, Z. Song, R. Zhang, H. Zhao, C. Lu, S. Li, and R. Henao, ‘‘Infoprompt: Information-theoretic soft prompt tuning for nat- ural language understanding,’’ arXiv preprint arXiv:2306.04933, 2023
2023 arXiv
-
[110]
L. Chen, H. Huang, and M. Cheng, ‘‘Ptp: Boosting stability and performance of prompt tuning with perturbation-based regularizer,’’ arXiv preprint arXiv:2305.02423, 2023
2023 arXiv
-
[111]
Y . Qin, X. Wang, Y . Su, Y . Lin, N. Ding, J. Yi, W. Chen, Z. Liu, J. Li, L. Hou et al. , ‘‘Exploring universal intrinsic task subspace via prompt tuning,’’ arXiv preprint arXiv:2110.07867, 2021
2021 arXiv
-
[112]
J.-Y . Choi, J. Kim, J.-H. Park, W.-L. Mok, and S. Lee, ‘‘Smop: Towards efficient and effective prompt tuning with sparse mixture-of-prompts,’’ in Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023, pp. 14 306–14 316
2023
-
[113]
Shi and A
Z. Shi and A. Lipani, ‘‘Dept: Decomposed prompt tuning for parameter-efficient fine-tuning,’’arXiv preprint arXiv:2309.05173, 2023
2023 arXiv
-
[114]
H. Liu, D. Tam, M. Muqeeth, J. Mohta, T. Huang, M. Bansal, and C. A. Raffel, ‘‘Few-shot parameter-efficient fine-tuning is better and cheaper than in-context learning,’’ Advances in Neural Information Processing Systems, vol. 35, pp. 1950–1965, 2022
1950
-
[115]
Zadouri, A
T. Zadouri, A. Üstün, A. Ahmadian, B. Ermiş, A. Locatelli, and S. Hooker, ‘‘Pushing mixture of experts to the limit: Extremely parameter efficient moe for instruction tuning,’’ arXiv preprint arXiv:2309.05444, 2023
2023 arXiv
-
[116]
D. Lian, D. Zhou, J. Feng, and X. Wang, ‘‘Scaling & shifting your features: A new baseline for efficient model tuning,’’ Advances in Neural Information Processing Systems, vol. 35, pp. 109–123, 2022
2022
-
[117]
X. Lu, F. Brahman, P . West, J. Jang, K. Chandu, A. Ravichander, L. Qin, P . Ammanabrolu, L. Jiang, S. Ramnath et al. , ‘‘Inference-time policy adapters (ipa): Tailoring extreme-scale lms without fine-tuning,’’ arXiv preprint arXiv:2305.15065, 2023
2023 arXiv
-
[118]
D. Guo, A. M. Rush, and Y . Kim, ‘‘Parameter-efficient transfer learning with diff pruning,’’ arXiv preprint arXiv:2012.07463, 2020
2012 arXiv
-
[119]
Lawton, A
N. Lawton, A. Kumar, G. Thattai, A. Galstyan, and G. V . Steeg, ‘‘Neural architecture search for parameter-efficient fine-tuning of large pre-trained language models,’’ arXiv preprint arXiv:2305.16597, 2023
2023 arXiv
-
[120]
B. Liao, Y . Meng, and C. Monz, ‘‘Parameter-efficient fine-tuning without introducing new latency,’’ arXiv preprint arXiv:2305.16742, 2023
2023 arXiv
-
[121]
Y .-L. Sung, V . Nair, and C. A. Raffel, ‘‘Training neural networks with fixed sparse masks,’’ Advances in Neural Information Processing Systems, vol. 34, pp. 24 193–24 205, 2021
2021
-
[122]
S. S. S. Das, R. H. Zhang, P . Shi, W. Yin, and R. Zhang, ‘‘Unified low-resource sequence labeling by sample-aware dynamic sparse finetuning,’’ arXiv preprint arXiv:2311.03748, 2023
2023 arXiv
-
[123]
Ansell, E
A. Ansell, E. M. Ponti, A. Korhonen, and I. Vulić, ‘‘Composable sparse fine-tuning for cross-lingual transfer,’’ arXiv preprint arXiv:2110.07560, 2021
2021 arXiv
-
[124]
Z. Fu, H. Y ang, A. M.-C. So, W. Lam, L. Bing, and N. Collier, ‘‘On the effectiveness of parameter-efficient fine-tuning,’’ in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 37, 2023, pp. 12 799–12 807
2023
-
[125]
R. Xu, F. Luo, Z. Zhang, C. Tan, B. Chang, S. Huang, and F. Huang, ‘‘Raise a child in large language model: Towards effective and generalizable fine-tuning,’’ arXiv preprint arXiv:2109.05687, 2021
2021 arXiv
-
[126]
Vucetic, M
D. Vucetic, M. Tayaranian, M. Ziaeefard, J. J. Clark, B. H. Meyer, and W. J. Gross, ‘‘Efficient fine-tuning of bert models on the edge,’’ in 2022 IEEE International Symposium on Circuits and Systems (ISCAS). IEEE, 2022, pp. 1838–1842
2022
-
[127]
E. B. Zaken, S. Ravfogel, and Y . Goldberg, ‘‘Bitfit: Simple parameter-efficient fine-tuning for transformer-based masked language-models,’’ arXiv preprint arXiv:2106.10199, 2021
2021
-
[128]
Gheini, X
M. Gheini, X. Ren, and J. May, ‘‘Cross-attention is all you need: Adapting pretrained transformers for machine translation,’’ arXiv preprint arXiv:2104.08771, 2021
2021 arXiv
-
[129]
Aghajanyan, L
A. Aghajanyan, L. Zettlemoyer, and S. Gupta, ‘‘Intrinsic dimensionality explains the effectiveness of language model fine-tuning,’’ arXiv preprint arXiv:2012.13255, 2020
2012 arXiv
-
[131]
Karimi Mahabadi, J
R. Karimi Mahabadi, J. Henderson, and S. Ruder, ‘‘Compacter: Efficient low-rank hypercomplex adapter layers,’’Advances in Neural Information Processing Systems, vol. 34, pp. 1022–1035, 2021
2021
-
[132]
Edalati, M
A. Edalati, M. Tahaei, I. Kobyzev, V . P . Nia, J. J. Clark, and M. Rezagholizadeh, ‘‘Krona: Parameter efficient tuning with kronecker adapter,’’arXiv preprint arXiv:2212.10650, 2022
2022 arXiv
-
[133]
X. He, C. Li, P . Zhang, J. Y ang, and X. E. Wang, ‘‘Parameter-efficient model adaptation for vision transformers,’’ in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 37, 2023, pp. 817–825
2023
-
[134]
D. J. Kopiczko, T. Blankevoort, and Y . M. Asano, ‘‘V era: V ector-based random matrix adaptation,’’ arXiv preprint arXiv:2310.11454, 2023
2023 arXiv
-
[135]
Liu, C.-Y
S.-Y . Liu, C.-Y . Wang, H. Yin, P . Molchanov, Y .-C. F. Wang, K.-T. Cheng, and M.-H. Chen, ‘‘Dora: Weight-decomposed low-rank adaptation,’’ arXiv preprint arXiv:2402.09353, 2024
2024 arXiv
-
[136]
V alipour, M
M. V alipour, M. Rezagholizadeh, I. Kobyzev, and A. Ghodsi, ‘‘Dylora: Parameter efficient tuning of pre-trained models using dynamic search- free low-rank adaptation,’’ arXiv preprint arXiv:2210.07558, 2022
2022 arXiv
-
[137]
Zhang, M
Q. Zhang, M. Chen, A. Bukharin, P . He, Y . Cheng, W. Chen, and T. Zhao, ‘‘Adaptive budget allocation for parameter-efficient fine-tuning,’’ arXiv preprint arXiv:2303.10512, 2023
2023 arXiv
-
[138]
N. Ding, X. Lv, Q. Wang, Y . Chen, B. Zhou, Z. Liu, and M. Sun, ‘‘Sparse low-rank adaptation of pre-trained language models,’’ arXiv preprint arXiv:2311.11696, 2023
2023 arXiv
-
[139]
Haobo, H
S. Haobo, H. Zhao, S. Majumder, and T. Lin, ‘‘Increasing model capacity for free: A simple strategy for parameter efficient fine-tuning,’’ in The Twelfth International Conference on Learning Representations , 2023
2023
-
[140]
Zhang, R
R. Zhang, R. Qiang, S. A. Somayajula, and P . Xie, ‘‘Autolora: Automatically tuning matrix ranks in low-rank adaptation based on meta learning,’’ arXiv preprint arXiv:2403.09113, 2024
2024 arXiv
-
[141]
A. X. Y ang, M. Robeyns, X. Wang, and L. Aitchison, ‘‘Bayesian low-rank adaptation for large language models,’’arXiv preprint arXiv:2308.13111, 2023
2023 arXiv
-
[142]
Y . Lin, X. Ma, X. Chu, Y . Jin, Z. Y ang, Y . Wang, and H. Mei, ‘‘Lora dropout as a sparsity regularizer for overfitting control,’’ arXiv preprint arXiv:2404.09610, 2024
2024 arXiv
-
[143]
X. Meng, D. Dai, W. Luo, Z. Y ang, S. Wu, X. Wang, P . Wang, Q. Dong, L. Chen, and Z. Sui, ‘‘Periodiclora: Breaking the low-rank bottleneck in lora optimization,’’ arXiv preprint arXiv:2402.16141, 2024
2024 arXiv
-
[144]
Hayou, N
S. Hayou, N. Ghosh, and B. Y u, ‘‘Lora+: Efficient low rank adaptation of large models,’’ arXiv preprint arXiv:2402.12354, 2024
2024 arXiv
-
[145]
Y . Chen, S. Qian, H. Tang, X. Lai, Z. Liu, S. Han, and J. Jia, ‘‘Longlora: Efficient fine-tuning of long-context large language models,’’ 2024. [Online]. Available: https://arxiv.org/abs/2309.12307
2024 arXiv
-
[146]
Huang, Q
C. Huang, Q. Liu, B. Y . Lin, T. Pang, C. Du, and M. Lin, ‘‘Lorahub: Efficient cross-task generalization via dynamic lora composition,’’ arXiv preprint arXiv:2307.13269, 2023
2023 arXiv
-
[147]
Q. Liu, X. Wu, X. Zhao, Y . Zhu, D. Xu, F. Tian, and Y . Zheng, ‘‘Moelora: An moe-based parameter efficient fine-tuning method for multi-task medical applications,’’ arXiv preprint arXiv:2310.18339, 2023
2023 arXiv
-
[148]
W. Feng, C. Hao, Y . Zhang, Y . Han, and H. Wang, ‘‘Mixture-of-loras: An efficient multitask tuning for large language models,’’ arXiv preprint arXiv:2403.03432, 2024
2024 arXiv
-
[149]
X. Wu, S. Huang, and F. Wei, ‘‘Mixture of lora experts,’’ arXiv preprint arXiv:2404.13628, 2024
2024 arXiv
-
[150]
D. Li, Y . Ma, N. Wang, Z. Cheng, L. Duan, J. Zuo, C. Y ang, and M. Tang, ‘‘Mixlora: Enhancing large language models fine-tuning with lora based mixture of experts,’’ arXiv preprint arXiv:2404.15159, 2024
2024 arXiv
-
[151]
Y . Mao, L. Mathias, R. Hou, A. Almahairi, H. Ma, J. Han, W.-t. Yih, and M. Khabsa, ‘‘Unipelt: A unified framework for parameter-efficient language model tuning,’’ arXiv preprint arXiv:2110.07577, 2021
2021 arXiv
-
[152]
J. Chen, A. Zhang, X. Shi, M. Li, A. Smola, and D. Y ang, ‘‘Parameter-efficient fine-tuning design spaces,’’ arXiv preprint arXiv:2301.01821, 2023
2023 arXiv
-
[153]
Zhang, K
Y . Zhang, K. Zhou, and Z. Liu, ‘‘Neural prompt search,’’ 2022
2022
-
[154]
Zhong, J
S. Zhong, J. Mo, and Z. Liu, ‘‘Autopet challenge 2022: Automatic segmentation of whole-body tumor lesion based on deep learning and fdg pet/ct,’’ arXiv preprint arXiv:2209.01212, 2022
2022 arXiv
-
[155]
Z. Hu, Y . Lan, L. Wang, W. Xu, E.-P . Lim, R. K.-W. Lee, L. Bing, and S. Poria, ‘‘Llm-adapters: An adapter family for parameter-efficient fine-tuning of large language models,’’ arXiv preprint arXiv:2304.01933, 2023
2023 arXiv
-
[156]
S. Hu, Z. Zhang, N. Ding, Y . Wang, Y . Wang, Z. Liu, and M. Sun, ‘‘Sparse structure search for parameter-efficient tuning,’’ arXiv preprint arXiv:2206.07382, 2022. VOLUME 11, 2024 33 Minghao et al.: Survey of different Large Language Model Architectures: Trends, Benchmarks, a...
2022 arXiv
-
[157]
E. J. Hu, Y . Shen, P . Wallis, Z. Allen-Zhu, Y . Li, S. Wang, L. Wang, and W. Chen, ‘‘Lora: Low-rank adaptation of large language models,’’ 2021. [Online]. Available: https://arxiv.org/abs/2106.09685
2021 arXiv
-
[158]
H. Liu, C. Li, Q. Wu, and Y . J. Lee, ‘‘Visual instruction tuning,’’ arXiv preprint arXiv:2304.08485, 2023
2023 arXiv
-
[159]
P . Lu, S. Mishra, T. Xia, L. Qiu, K.-W. Chang, S.-C. Zhu, O. Tafjord, P . Clark, and A. Kalyan, ‘‘Learn to explain: Multimodal reasoning via thought chains for science question answering,’’ 2022. [Online]. Available: https://arxiv.org/abs/2209.09513
2022 arXiv
-
[160]
Y .-L. Sung, J. Cho, and M. Bansal, ‘‘Vl-adapter: Parameter-efficient transfer learning for vision-and-language tasks,’’ 2022. [Online]. Available: https://arxiv.org/abs/2112.06825
2022 arXiv
-
[161]
Y . Cui, W. Che, T. Liu, B. Qin, and Z. Y ang, ‘‘Pre-training with whole word masking for chinese bert,’’ IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 29, pp. 3504–3514, 2021
2021
-
[162]
Y . Cui, W. Che, T. Liu, B. Qin, Z. Y ang, S. Wang, and G. Hu, ‘‘Pre-training with whole word masking for chinese bert,’’ arXiv preprint arXiv:1906.08101, 2019
1906 arXiv
-
[163]
Joshi, D
M. Joshi, D. Chen, Y . Liu, D. S. Weld, L. Zettlemoyer, and O. Levy, ‘‘Spanbert: Improving pre-training by representing and predicting spans,’’ Transactions of the association for computational linguistics , vol. 8, pp. 64–77, 2020
2020
-
[164]
V . Sanh, L. Debut, J. Chaumond, and T. Wolf, ‘‘Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter,’’ arXiv preprint arXiv:1910.01108, 2019
1910 arXiv
-
[165]
Y . Liu, M. Ott, N. Goyal, J. Du, M. Joshi, D. Chen, O. Levy, M. Lewis, L. Zettlemoyer, and V . Stoyanov, ‘‘Roberta: A robustly optimized bert pretraining approach,’’ arXiv preprint arXiv:1907.11692, 2019
1907 arXiv
-
[166]
X. Jiao, Y . Yin, L. Shang, X. Jiang, X. Chen, L. Li, F. Wang, and Q. Liu, ‘‘Tinybert: Distilling bert for natural language understanding,’’ arXiv preprint arXiv:1909.10351, 2019
1909 arXiv
-
[167]
L. H. Li, M. Y atskar, D. Yin, C.-J. Hsieh, and K.-W. Chang, ‘‘Visualbert: A simple and performant baseline for vision and language,’’ arXiv preprint arXiv:1908.03557, 2019
1908 arXiv
-
[168]
Y . Cui, W. Che, T. Liu, B. Qin, S. Wang, and G. Hu, ‘‘Revisiting pre-trained models for chinese natural language processing,’’ arXiv preprint arXiv:2004.13922, 2020
2004 arXiv
-
[169]
H. Bao, L. Dong, S. Piao, and F. Wei, ‘‘Beit: Bert pre-training of image transformers,’’ arXiv preprint arXiv:2106.08254, 2021
2021 arXiv
-
[170]
Z. Peng, L. Dong, H. Bao, Q. Y e, and F. Wei, ‘‘Beit v2: Masked image modeling with vector-quantized visual tokenizers,’’ arXiv preprint arXiv:2208.06366, 2022
2022 arXiv
-
[171]
W. Wang, H. Bao, L. Dong, J. Bjorck, Z. Peng, Q. Liu, K. Aggarwal, O. K. Mohammed, S. Singhal, S. Som et al. , ‘‘Image as a foreign language: Beit pretraining for all vision and vision-language tasks,’’ arXiv preprint arXiv:2208.10442, 2022
2022 arXiv
-
[172]
V ahdat, E
A. V ahdat, E. Andriyash, and W. Macready, ‘‘Dvae#: Discrete variational autoencoders with relaxed boltzmann priors,’’ Advances in Neural Information Processing Systems, vol. 31, 2018
2018
-
[173]
Xu, ‘‘Roberta-wwm-ext fine-tuning for chinese text classification,’’ arXiv preprint arXiv:2103.00492, 2021
Z. Xu, ‘‘Roberta-wwm-ext fine-tuning for chinese text classification,’’ arXiv preprint arXiv:2103.00492, 2021
2021 arXiv
-
[174]
Y . Sun, S. Wang, Y . Li, S. Feng, H. Tian, H. Wu, and H. Wang, ‘‘Ernie 2.0: A continual pre-training framework for language understanding,’’ in Proceedings of the AAAI conference on artificial intelligence , vol. 34, 2020, pp. 8968–8975
2020
-
[175]
Y . Sun, S. Wang, S. Feng, S. Ding, C. Pang, J. Shang, J. Liu, X. Chen, Y . Zhao, Y . Lu et al. , ‘‘Ernie 3.0: Large-scale knowledge enhanced pre-training for language understanding and generation,’’ arXiv preprint arXiv:2107.02137, 2021
2021 arXiv
-
[176]
F. Y u, J. Tang, W. Yin, Y . Sun, H. Tian, H. Wu, and H. Wang, ‘‘Ernie-vil: Knowledge enhanced vision-language representations through scene graphs,’’ inProceedings of the AAAI Conference on Artificial Intelligence, vol. 35, 2021, pp. 3208–3216
2021
-
[177]
B. Shan, W. Yin, Y . Sun, H. Tian, H. Wu, and H. Wang, ‘‘Ernie-vil 2.0: Multi-view contrastive learning for image-text pre-training,’’ arXiv preprint arXiv:2209.15270, 2022
2022 arXiv
-
[178]
Clark, M.-T
K. Clark, M.-T. Luong, Q. V . Le, and C. D. Manning, ‘‘Electra: Pre-training text encoders as discriminators rather than generators,’’ arXiv preprint arXiv:2003.10555, 2020
2003 arXiv
-
[179]
Goodfellow, J
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y . Bengio, ‘‘Generative adversarial networks,’’ Communications of the ACM , vol. 63, no. 11, pp. 139–144, 2020
2020
-
[180]
P . He, X. Liu, J. Gao, and W. Chen, ‘‘Deberta: Decoding-enhanced bert with disentangled attention,’’ arXiv preprint arXiv:2006.03654, 2020
2006 arXiv
-
[181]
Z. Dai, Z. Y ang, Y . Y ang, J. Carbonell, Q. V . Le, and R. Salakhutdinov, ‘‘Transformer-xl: Attentive language models beyond a fixed-length context,’’arXiv preprint arXiv:1901.02860, 2019
1901 arXiv
-
[182]
Y ang, Z
Z. Y ang, Z. Dai, Y . Y ang, J. Carbonell, R. R. Salakhutdinov, and Q. V . Le, ‘‘Xlnet: Generalized autoregressive pretraining for language understand- ing,’’ Advances in neural information processing systems , vol. 32, 2019
2019
-
[183]
Y . Zhu, R. Kiros, R. Zemel, R. Salakhutdinov, R. Urtasun, A. Torralba, and S. Fidler, ‘‘Aligning books and movies: Towards story-like visual explanations by watching movies and reading books,’’ in The IEEE International Conference on Computer Vision (ICCV) , December 2015
2015
-
[184]
Radford, J
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever et al. , ‘‘Language models are unsupervised multitask learners,’’ OpenAI blog , vol. 1, no. 8, p. 9, 2019
2019
-
[185]
McCann, N
B. McCann, N. S. Keskar, C. Xiong, and R. Socher, ‘‘The natural language decathlon: Multitask learning as question answering,’’ arXiv preprint arXiv:1806.08730, 2018
2018 arXiv
-
[186]
Brown, B
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P . Dhariwal, A. Neelakantan, P . Shyam, G. Sastry, A. Askellet al., ‘‘Language models are few-shot learners,’’ Advances in neural information processing systems, vol. 33, pp. 1877–1901, 2020
1901
-
[187]
C. Finn, P . Abbeel, and S. Levine, ‘‘Model-agnostic meta-learning for fast adaptation of deep networks,’’ in International conference on machine learning. PMLR, 2017, pp. 1126–1135
2017
-
[188]
[Online]
‘‘commoncrawl,’’ 2022. [Online]. Available: https://commoncrawl.org/
2022
-
[189]
Ouyang, J
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P . Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray et al. , ‘‘Training language models to follow instructions with human feedback,’’ Advances in Neural Information Processing Systems, vol. 35, pp. 27 730–27 744, 2022
2022
-
[190]
[Online]
OpenAI, ‘‘Chatgpt,’’ 2022. [Online]. Available: https://openai.com/chatgpt
2022
-
[191]
——, ‘‘Gpt-4 technical report,’’ 2023
2023
-
[192]
Black, L
S. Black, L. Gao, P . Wang, C. Leahy, and S. R. Biderman, ‘‘Gpt-neo: Large scale autoregressive language modeling with mesh-tensorflow,’’ 2021. [Online]. Available: https://api.semanticscholar.org/CorpusID:245758737
2021
-
[193]
Z. Hu, Y . Dong, K. Wang, K.-W. Chang, and Y . Sun, ‘‘Gpt-gnn: Generative pre-training of graph neural networks,’’ in Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, 2020, pp. 1857–1867
2020
-
[194]
Komatsuzaki, ‘‘Gpt-j-6b: 6b jax-based transformer,’’ Jun 2021
A. Komatsuzaki, ‘‘Gpt-j-6b: 6b jax-based transformer,’’ Jun 2021. [Online]. Available: https://arankomatsuzaki.wordpress.com/2021/06/ 04/gpt-j/
2021
-
[195]
Huebner, ‘‘Mesh transformer jax,’’ Feb 2023
C. Huebner, ‘‘Mesh transformer jax,’’ Feb 2023. [Online]. Available: https://www.eleuther.ai/artifacts/mtj
2023
-
[196]
Black, S
S. Black, S. Biderman, E. Hallahan, Q. Anthony, L. Gao, L. Golding, H. He, C. Leahy, K. McDonell, J. Phang et al. , ‘‘Gpt-neox-20b: An open-source autoregressive language model,’’ arXiv preprint arXiv:2204.06745, 2022
2022 arXiv
-
[197]
Microsoft, ‘‘Deepspeed is a deep learning optimization library that makes distributed training and inference easy, efficient, and effective.’’
-
[198]
Zhang, S
Y . Zhang, S. Sun, M. Galley, Y .-C. Chen, C. Brockett, X. Gao, J. Gao, J. Liu, and B. Dolan, ‘‘Dialogpt: Large-scale generative pre-training for conversational response generation,’’ arXiv preprint arXiv:1911.00536 , 2019
1911 arXiv
-
[199]
Google, ‘‘Introducing pathways: A next-generation ai architecture,’’
-
[200]
——, ‘‘Pathways language model (palm): Scaling to 540 billion parameters for breakthrough performance,’’
-
[201]
Available: https://blog.google/technology/ai/ introducing-pathways-next-generation-ai-architecture/
[Online]. Available: https://blog.google/technology/ai/ introducing-pathways-next-generation-ai-architecture/
-
[202]
J. Su, Y . Lu, S. Pan, A. Murtadha, B. Wen, and Y . Liu, ‘‘Roformer: Enhanced transformer with rotary position embedding,’’ arXiv preprint arXiv:2104.09864, 2021
2021 arXiv
-
[203]
Available: https://blog.research.google/2022/04/ pathways-language-model-palm-scaling-to.html
[Online]. Available: https://blog.research.google/2022/04/ pathways-language-model-palm-scaling-to.html
2022
-
[204]
Shazeer, ‘‘Glu variants improve transformer,’’ arXiv preprint arXiv:2002.05202, 2020
N. Shazeer, ‘‘Glu variants improve transformer,’’ arXiv preprint arXiv:2002.05202, 2020
2002 arXiv
-
[205]
T.-Y . Lin, M. Maire, S. Belongie, J. Hays, P . Perona, D. Ramanan, P . Dollár, and C. L. Zitnick, ‘‘Microsoft coco: Common objects in context,’’ in Computer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part V 13 . Springe...
2014
-
[206]
Driess, F
D. Driess, F. Xia, M. S. Sajjadi, C. Lynch, A. Chowdhery, B. Ichter, A. Wahid, J. Tompson, Q. Vuong, T. Y u et al. , ‘‘Palm-e: An embodied multimodal language model,’’ arXiv preprint arXiv:2303.03378, 2023
2023 arXiv
-
[207]
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, ‘‘Imagenet: A large-scale hierarchical image database,’’ in 2009 IEEE Conference on Computer Vision and Pattern Recognition, 2009, pp. 248–255. 34 VOLUME 11, 2024 Minghao et al.: Survey of different Large Language M...
2009
-
[208]
Mehta, A
H. Mehta, A. Thakurta, A. Kurakin, and A. Cutkosky, ‘‘Large scale transfer learning for differentially private image classification,’’ arXiv preprint arXiv:2205.02973, 2022
2022 arXiv
-
[209]
Krishna, Y
R. Krishna, Y . Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y . Kalantidis, L.-J. Li, D. A. Shammaet al., ‘‘Visual genome: Connecting language and vision using crowdsourced dense image annotations,’’ International journal of computer vision , vol. 123, pp. 32–73, 2017
2017
-
[210]
Dehghani, J
M. Dehghani, J. Djolonga, B. Mustafa, P . Padlewski, J. Heek, J. Gilmer, A. P . Steiner, M. Caron, R. Geirhos, I. Alabdulmohsin et al. , ‘‘Scaling vision transformers to 22 billion parameters,’’ inInternational Conference on Machine Learning. PMLR, 2023, pp. 7480–7512
2023
-
[211]
X. Chen, X. Wang, S. Changpinyo, A. Piergiovanni, P . Padlewski, D. Salz, S. Goodman, A. Grycner, B. Mustafa, L. Beyer et al. , ‘‘Pali: A jointly-scaled multilingual language-image model,’’ arXiv preprint arXiv:2209.06794, 2022
2022 arXiv
-
[212]
Huang, L
S. Huang, L. Dong, W. Wang, Y . Hao, S. Singhal, S. Ma, T. Lv, L. Cui, O. K. Mohammed, Q. Liu et al., ‘‘Language is not all you need: Aligning perception with language models,’’ arXiv preprint arXiv:2302.14045 , 2023
2023 arXiv
-
[213]
Lewkowycz, A
A. Lewkowycz, A. Andreassen, D. Dohan, E. Dyer, H. Michalewski, V . Ramasesh, A. Slone, C. Anil, I. Schlag, T. Gutman-Soloet al., ‘‘Solving quantitative reasoning problems with language models,’’ Advances in Neural Information Processing Systems, vol. 35, pp. 3843–3857, 2022
2022
-
[214]
Dosovitskiy, L
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, and T. Unterthiner, ‘‘Transformers for image recognition at scale,’’ arXiv preprint arXiv:2010.11929, 2020
2010 arXiv
-
[215]
L. Xue, N. Constant, A. Roberts, M. Kale, R. Al-Rfou, A. Siddhant, A. Barua, and C. Raffel, ‘‘mt5: A massively multilingual pre-trained text-to-text transformer,’’arXiv preprint arXiv:2010.11934, 2020
2010 arXiv
-
[216]
X. Zhai, A. Kolesnikov, N. Houlsby, and L. Beyer, ‘‘Scaling vision transformers,’’ inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 12 104–12 113
2022
-
[217]
D. Jin, E. Pan, N. Oufattole, W.-H. Weng, H. Fang, and P . Szolovits, ‘‘What disease does this patient have? a large-scale open domain question answering dataset from medical exams,’’ Applied Sciences , vol. 11, no. 14, p. 6421, 2021
2021
-
[218]
R. Anil, A. M. Dai, O. Firat, M. Johnson, D. Lepikhin, A. Passos, S. Shakeri, E. Taropa, P . Bailey, Z. Chenet al., ‘‘Palm 2 technical report,’’ arXiv preprint arXiv:2305.10403, 2023
2023 arXiv
-
[219]
Singhal, T
K. Singhal, T. Tu, J. Gottweis, R. Sayres, E. Wulczyn, L. Hou, K. Clark, S. Pfohl, H. Cole-Lewis, D. Neal et al. , ‘‘Towards expert-level medical question answering with large language models,’’ arXiv preprint arXiv:2305.09617, 2023
2023 arXiv
-
[221]
Singhal, S
K. Singhal, S. Azizi, T. Tu, S. S. Mahdavi, J. Wei, H. W. Chung, N. Scales, A. Tanwani, H. Cole-Lewis, S. Pfohl et al. , ‘‘Large language models encode clinical knowledge,’’ arXiv preprint arXiv:2212.13138, 2022
2022 arXiv
-
[222]
H. Wang, S. Ma, S. Huang, L. Dong, W. Wang, Z. Peng, Y . Wu, P . Bajaj, S. Singhal, A. Benhaim et al., ‘‘Foundation transformers,’’arXiv preprint arXiv:2210.06423, 2022
2022 arXiv
-
[223]
V . A. Korthikanti, J. Casper, S. Lym, L. McAfee, M. Andersch, M. Shoeybi, and B. Catanzaro, ‘‘Reducing activation recomputation in large transformer models,’’ Proceedings of Machine Learning and Systems, vol. 5, 2023
2023
-
[224]
Shoeybi, M
M. Shoeybi, M. Patwary, R. Puri, P . LeGresley, J. Casper, and B. Catan- zaro, ‘‘Megatron-lm: Training multi-billion parameter language models using model parallelism,’’ arXiv preprint arXiv:1909.08053, 2019
1909 arXiv
-
[225]
Narayanan, M
D. Narayanan, M. Shoeybi, J. Casper, P . LeGresley, M. Patwary, V . Korthikanti, D. V ainbrand, P . Kashinkunti, J. Bernauer, B. Catanzaro et al. , ‘‘Efficient large-scale language model training on gpu clusters using megatron-lm,’’ in Proceedings of the International Conferen...
2021
-
[226]
Smith, M
S. Smith, M. Patwary, B. Norick, P . LeGresley, S. Rajbhandari, J. Casper, Z. Liu, S. Prabhumoye, G. Zerveas, V . Korthikanti et al. , ‘‘Using deepspeed and megatron to train megatron-turing nlg 530b, a large-scale generative language model,’’ arXiv preprint arXiv:2201.11990, 2022
2022 arXiv
-
[227]
‘‘Turing-nlg: A 17-billion-parameter language model by microsoft,’’ Feb
-
[228]
P . Xu, M. Patwary, M. Shoeybi, R. Puri, P . Fung, A. Anandkumar, and B. Catanzaro, ‘‘Megatron-cntrl: Controllable story generation with external knowledge using large-scale language models,’’ arXiv preprint arXiv:2010.00840, 2020
2010 arXiv
-
[229]
Rajbhandari, J
S. Rajbhandari, J. Rasley, O. Ruwase, and Y . He, ‘‘Zero: Memory optimizations toward training trillion parameter models,’’ in SC20: International Conference for High Performance Computing, Networking, Storage and Analysis. IEEE, 2020, pp. 1–16
2020
-
[230]
Taori, I
R. Taori, I. Gulrajani, T. Zhang, Y . Dubois, X. Li, C. Guestrin, P . Liang, and T. B. Hashimoto, ‘‘Alpaca: A strong, replicable instruction-following model,’’ 2019
2019
-
[231]
H.-C. Shin, Y . Zhang, E. Bakhturina, R. Puri, M. Patwary, M. Shoeybi, and R. Mani, ‘‘Biomegatron: Larger biomedical domain language model,’’ arXiv preprint arXiv:2010.06060, 2020
2010 arXiv
-
[232]
[Online]
‘‘Guanaco - generative universal assistant for natural-language adaptive context-aware omnilingual outputs,’’ Feb 2020. [Online]. Available: https://guanaco-model.github.io/
2020
-
[233]
Zhang and R
B. Zhang and R. Sennrich, ‘‘Root mean square layer normalization,’’ Advances in Neural Information Processing Systems , vol. 32, 2019
2019
-
[234]
Chiang, Z
W.-L. Chiang, Z. Li, Z. Lin, Y . Sheng, Z. Wu, H. Zhang, L. Zheng, S. Zhuang, Y . Zhuang, J. E. Gonzalez, I. Stoica, and E. P . Xing, ‘‘Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality,’’ March 2023. [Online]. Available: https://lmsys.org/blog/2023-03-...
2023
-
[235]
[Online]
OpenAI, ‘‘Models - openai,’’ Feb 2020. [Online]. Available: https://platform.openai.com/docs/models
2020
-
[236]
[Online]
Databricks, ‘‘dolly,’’ Feb 2020. [Online]. Available: https://github.com/databrickslabs/dolly
2020
-
[237]
[Online]
‘‘Alpaca-lora,’’ Feb 2020. [Online]. Available: https: //github.com/tloen/alpaca-lora
2020
-
[238]
Biderman, H
S. Biderman, H. Schoelkopf, Q. G. Anthony, H. Bradley, K. O’Brien, E. Hallahan, M. A. Khan, S. Purohit, U. S. Prashanth, E. Raff et al. , ‘‘Pythia: A suite for analyzing large language models across training and scaling,’’ in International Conference on Machine Learning . PMLR...
2023
-
[239]
[Online]
ShareGPT, ‘‘Sharegpt,’’ Feb 2020. [Online]. Available: https://sharegpt.com/
2020
-
[240]
Zheng, W.-L
L. Zheng, W.-L. Chiang, Y . Sheng, S. Zhuang, Z. Wu, Y . Zhuang, Z. Lin, Z. Li, D. Li, E. P . Xing, H. Zhang, J. E. Gonzalez, and I. Stoica, ‘‘Judging llm-as-a-judge with mt-bench and chatbot arena,’’ 2023
2023
-
[241]
Conover, M
M. Conover, M. Hayes, A. Mathur, J. Xie, J. Wan, S. Shah, A. Ghodsi, P . Wendell, M. Zaharia, and R. Xin, ‘‘Free dolly: Introducing the world’s first truly open instruction-tuned llm,’’
-
[242]
L. L. Ziang Leng, Qiyuan Chen, ‘‘Luotuo: Chinese-alpaca-lora,’’ March 2023. [Online]. Available: https://github.com/LC1332/Chinese-alpaca-lora
2023
-
[243]
Y . Cui, Z. Y ang, and X. Y ao, ‘‘Efficient and effective text encoding for chinese llama and alpaca,’’ arXiv preprint arXiv:2304.08177 , 2023. [Online]. Available: https://arxiv.org/abs/2304.08177
2023 arXiv
-
[244]
X. Geng, A. Gudibande, H. Liu, E. Wallace, P . Abbeel, S. Levine, and D. Song, ‘‘Koala: A dialogue model for academic research,’’ Apr 2023. [Online]. Available: https://bair.berkeley.edu/blog/2023/04/03/koala/
2023
-
[245]
P . Gao, J. Han, R. Zhang, Z. Lin, S. Geng, A. Zhou, W. Zhang, P . Lu, C. He, X. Y ue et al. , ‘‘Llama-adapter v2: Parameter-efficient visual instruction model,’’ arXiv preprint arXiv:2304.15010, 2023
2023 arXiv
-
[246]
C. Xu, D. Guo, N. Duan, and J. McAuley, ‘‘Baize: An open-source chat model with parameter-efficient tuning on self-chat data,’’ arXiv preprint arXiv:2304.01196, 2023
2023 arXiv
-
[247]
LLaV A, ‘‘Chinese llava,’’ July 2023
C. LLaV A, ‘‘Chinese llava,’’ July 2023. [Online]. Available: https://github.com/LinkSoul-AI/Chinese-LLaV A VOLUME 11, 2024 35 Minghao et al.: Survey of different Large Language Model Architectures: Trends, Benchmarks, and Challenges
2023
-
[248]
Y . Shu, S. Dong, G. Chen, W. Huang, R. Zhang, D. Shi, Q. Xiang, and Y . Shi, ‘‘Llasm: Large language and speech model,’’ arXiv preprint arXiv:2308.15930, 2023
2023 arXiv
-
[249]
Zhang, J
R. Zhang, J. Han, A. Zhou, X. Hu, S. Y an, P . Lu, H. Li, P . Gao, and Y . Qiao, ‘‘Llama-adapter: Efficient fine-tuning of language models with zero-init attention,’’ arXiv preprint arXiv:2303.16199, 2023
2023 arXiv
-
[250]
Zhang, X
H. Zhang, X. Li, and L. Bing, ‘‘Video-llama: An instruction-tuned audio-visual language model for video understanding,’’ arXiv preprint arXiv:2306.02858, 2023
2023 arXiv
-
[251]
D. Zhu, J. Chen, X. Shen, X. Li, and M. Elhoseiny, ‘‘Minigpt-4: Enhancing vision-language understanding with advanced large language models,’’ arXiv preprint arXiv:2304.10592, 2023
2023 arXiv
-
[252]
K. Li, Y . He, Y . Wang, Y . Li, W. Wang, P . Luo, Y . Wang, L. Wang, and Y . Qiao, ‘‘Videochat: Chat-centric video understanding,’’ arXiv preprint arXiv:2305.06355, 2023
2023 arXiv
-
[253]
Girdhar, A
R. Girdhar, A. El-Nouby, Z. Liu, M. Singh, K. V . Alwala, A. Joulin, and I. Misra, ‘‘Imagebind: One embedding space to bind them all,’’ in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 15 180–15 190
2023
-
[254]
[Online]
Visual-LLaMA, ‘‘Visual-llama,’’ 2023. [Online]. Available: https://github.com/feizc/Visual-LLaMA
2023
-
[255]
Kudo and J
T. Kudo and J. Richardson, ‘‘Sentencepiece: A simple and language independent subword tokenizer and detokenizer for neural text processing,’’ arXiv preprint arXiv:1808.06226, 2018
2018 arXiv
-
[256]
Zhang, J
Q. Zhang, J. Zhang, Y . Xu, and D. Tao, ‘‘Vision transformer with quadrangle attention,’’ arXiv preprint arXiv:2303.15105, 2023
2023 arXiv
-
[257]
Hennigan, T
T. Hennigan, T. Cai, T. Norman, L. Martens, and I. Babuschkin, ‘‘Haiku: Sonnet for JAX,’’ 2020. [Online]. Available: http://github.com/deepmind/dm-haiku
2020
-
[258]
Hoffmann, S
J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. d. L. Casas, L. A. Hendricks, J. Welbl, A. Clark et al., ‘‘Training compute-optimal large language models,’’ arXiv preprint arXiv:2203.15556, 2022
2022 arXiv
-
[259]
J. W. Rae, S. Borgeaud, T. Cai, K. Millican, J. Hoffmann, F. Song, J. Aslanides, S. Henderson, R. Ring, S. Y oung et al., ‘‘Scaling language models: Methods, analysis & insights from training gopher,’’ arXiv preprint arXiv:2112.11446, 2021
2021 arXiv
-
[260]
Brock, S
A. Brock, S. De, S. L. Smith, and K. Simonyan, ‘‘High-performance large-scale image recognition without normalization,’’ in International Conference on Machine Learning. PMLR, 2021, pp. 1059–1071
2021
-
[261]
Bradbury, R
J. Bradbury, R. Frostig, P . Hawkins, M. J. Johnson, C. Leary, D. Maclaurin, G. Necula, A. Paszke, J. V anderPlas, S. Wanderman-Milne, and Q. Zhang, ‘‘JAX: composable transformations of Python+NumPy programs,’’ 2018. [Online]. Available: http://github.com/google/jax
2018
-
[262]
Lieber, O
O. Lieber, O. Sharir, B. Lenz, and Y . Shoham, ‘‘Jurassic-1: Technical details and evaluation,’’ White Paper . AI21 Labs, vol. 1, 2021
2021
-
[263]
Karpas, O
E. Karpas, O. Abend, Y . Belinkov, B. Lenz, O. Lieber, N. Ratner, Y . Shoham, H. Bata, Y . Levine, K. Leyton-Brownet al., ‘‘Mrkl systems: A modular, neuro-symbolic architecture that combines large language models, external knowledge sources and discrete reasoning,’’ arXiv prep...
2022 arXiv
-
[264]
Alayrac, J
J.-B. Alayrac, J. Donahue, P . Luc, A. Miech, I. Barr, Y . Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds et al. , ‘‘Flamingo: a visual language model for few-shot learning,’’ Advances in Neural Information Processing Systems, vol. 35, pp. 23 716–23 736, 2022
2022
-
[265]
[Online]
Anthropic, ‘‘Introducing claude,’’ Mar 2023. [Online]. Available: https://www.anthropic.com/index/introducing-claude
2023
-
[266]
Lab, ‘‘Ai21 studio logo,’’ August 2021
A. Lab, ‘‘Ai21 studio logo,’’ August 2021. [Online]. Available: https://www.ai21.com/studio
2021
-
[267]
L. Mei, J. Mao, Z. Wang, C. Gan, and J. B. Tenenbaum, ‘‘Falcon: fast visual concept learning by integrating images, linguistic descriptions, and conceptual relations,’’ arXiv preprint arXiv:2203.16639, 2022
2022 arXiv
-
[268]
Penedo, Q
G. Penedo, Q. Malartic, D. Hesslow, R. Cojocaru, A. Cappelli, H. Alobeidli, B. Pannier, E. Almazrouei, and J. Launay, ‘‘The refinedweb dataset for falcon llm: outperforming curated corpora with web data, and web data only,’’ arXiv preprint arXiv:2306.01116, 2023
2023 arXiv
-
[269]
Lab, ‘‘Announcing jurassic-2 and task-specific apis,’’ March 2022
A. Lab, ‘‘Announcing jurassic-2 and task-specific apis,’’ March 2022. [Online]. Available: https://www.ai21.com/blog/introducing-j2
2022
-
[270]
Radford, J
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P . Mishkin, J. Clark et al. , ‘‘Learning transferable visual models from natural language supervision,’’ in International conference on machine learning. PMLR, 2021, pp. 8748–8763
2021
-
[271]
[Online]
——, ‘‘Claude 2,’’ Jul 2023. [Online]. Available: https://www.anthropic.com/index/claude-2
2023
-
[272]
[Online]
OpenAI, ‘‘Dall ·e 3,’’ September 2023. [Online]. Available: https://openai.com/dall-e-3
2023
-
[273]
Radford, J
A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever, ‘‘Robust speech recognition via large-scale weak supervision,’’ in International Conference on Machine Learning . PMLR, 2023, pp. 28 492–28 518
2023
-
[274]
Ramesh, M
A. Ramesh, M. Pavlov, G. Goh, S. Gray, C. V oss, A. Radford, M. Chen, and I. Sutskever, ‘‘Zero-shot text-to-image generation,’’ in International Conference on Machine Learning. PMLR, 2021, pp. 8821–8831
2021
-
[275]
R. Li, L. B. Allal, Y . Zi, N. Muennighoff, D. Kocetkov, C. Mou, M. Marone, C. Akiki, J. Li, J. Chim et al. , ‘‘Starcoder: may the source be with you!’’ arXiv preprint arXiv:2305.06161, 2023
2023 arXiv
-
[276]
Ramesh, P
A. Ramesh, P . Dhariwal, A. Nichol, C. Chu, and M. Chen, ‘‘Hierarchical text-conditional image generation with clip latents,’’ arXiv preprint arXiv:2204.06125, vol. 1, no. 2, p. 3, 2022
2022 arXiv
-
[277]
Thoppilan, D
R. Thoppilan, D. De Freitas, J. Hall, N. Shazeer, A. Kulshreshtha, H.-T. Cheng, A. Jin, T. Bos, L. Baker, Y . Du et al., ‘‘Lamda: Language models for dialog applications,’’ arXiv preprint arXiv:2201.08239, 2022
2022 arXiv
-
[278]
G. H. Cohen, ‘‘Align: a program to superimpose protein coordinates, accounting for insertions and deletions,’’ Journal of applied crystallography, vol. 30, no. 6, pp. 1160–1161, 1997
1997
-
[279]
M. Chen, J. Tworek, H. Jun, Q. Y uan, H. P . d. O. Pinto, J. Kaplan, H. Ed- wards, Y . Burda, N. Joseph, G. Brockman et al. , ‘‘Evaluating large lan- guage models trained on code,’’ arXiv preprint arXiv:2107.03374, 2021
2021 arXiv
-
[280]
N. Du, Y . Huang, A. M. Dai, S. Tong, D. Lepikhin, Y . Xu, M. Krikun, Y . Zhou, A. W. Y u, O. Firatet al., ‘‘Glam: Efficient scaling of language models with mixture-of-experts,’’ in International Conference on Machine Learning. PMLR, 2022, pp. 5547–5569
2022
-
[281]
D. So, Q. Le, and C. Liang, ‘‘The evolved transformer,’’ in International conference on machine learning. PMLR, 2019, pp. 5877–5886
2019
-
[282]
[Online]
‘‘Phi-1.5,’’ 2023. [Online]. Available: https://huggingface.co/microsoft/ phi-1_5
2023
-
[283]
[Online]
‘‘Phi-2: The surprising power of small language models,’’ 2023. [Online]. Available: https://www.microsoft.com/en-us/research/blog/ phi-2-the-surprising-power-of-small-language-models/
2023
-
[284]
Koonce and B
B. Koonce and B. Koonce, ‘‘Efficientnet,’’ Convolutional Neural Networks with Swift for Tensorflow: Image Recognition and Dataset Categorization, pp. 109–123, 2021
2021
-
[285]
H. Xu, Q. Y e, M. Y an, Y . Shi, J. Y e, Y . Xu, C. Li, B. Bi, Q. Qian, W. Wang et al. , ‘‘mplug-2: A modularized multi-modal foundation model across text, image and video,’’ arXiv preprint arXiv:2302.00402, 2023
2023 arXiv
-
[286]
G. Team, R. Anil, S. Borgeaud, Y . Wu, J.-B. Alayrac, J. Y u, R. Soricut, J. Schalkwyk, A. M. Dai, A. Hauth et al. , ‘‘Gemini: a family of highly capable multimodal models,’’ arXiv preprint arXiv:2312.11805, 2023
2023 arXiv
-
[287]
Y . Shen, K. Song, X. Tan, D. Li, W. Lu, and Y . Zhuang, ‘‘Hugginggpt: Solving ai tasks with chatgpt and its friends in huggingface,’’ arXiv preprint arXiv:2303.17580, 2023
2023 arXiv
-
[288]
H. Xu, Q. Y e, X. Wu, M. Y an, Y . Miao, J. Y e, G. Xu, A. Hu, Y . Shi, G. Xu et al. , ‘‘Y ouku-mplug: A 10 million large-scale chinese video-language dataset for pre-training and benchmarks,’’ arXiv preprint arXiv:2306.04362, 2023
2023 arXiv
-
[289]
C. Li, H. Xu, J. Tian, W. Wang, M. Y an, B. Bi, J. Y e, H. Chen, G. Xu, Z. Cao et al., ‘‘mplug: Effective and efficient vision-language learning by cross-modal skip-connections,’’ arXiv preprint arXiv:2205.12005, 2022
2022 arXiv
-
[290]
Sanders, D
K. Sanders, D. Etter, R. Kriz, and B. V an Durme, ‘‘Multivent: Multilingual videos of events with aligned natural text,’’ arXiv preprint arXiv:2307.03153, 2023
2023 arXiv
-
[291]
Y ang, L
Z. Y ang, L. Li, J. Wang, K. Lin, E. Azarnasab, F. Ahmed, Z. Liu, C. Liu, M. Zeng, and L. Wang, ‘‘Mm-react: Prompting chatgpt for multimodal reasoning and action,’’ arXiv preprint arXiv:2303.11381, 2023
2023 arXiv
-
[292]
S. Bao, H. He, F. Wang, H. Wu, and H. Wang, ‘‘Plato: Pre-trained dialogue generation model with discrete latent variable,’’ arXiv preprint arXiv:1910.07931, 2019
1910 arXiv
-
[293]
S. Bao, H. He, F. Wang, H. Wu, H. Wang, W. Wu, Z. Guo, Z. Liu, and X. Xu, ‘‘Plato-2: Towards building an open-domain chatbot via curriculum learning,’’ arXiv preprint arXiv:2006.16779, 2020
2006 arXiv
-
[294]
J. Y e, A. Hu, H. Xu, Q. Y e, M. Y an, Y . Dan, C. Zhao, G. Xu, C. Li, J. Tian et al. , ‘‘mplug-docowl: Modularized multimodal large language model for document understanding,’’ arXiv preprint arXiv:2307.02499, 2023
2023 arXiv
-
[295]
Y . Huo, M. Zhang, G. Liu, H. Lu, Y . Gao, G. Y ang, J. Wen, H. Zhang, B. Xu, W. Zheng et al., ‘‘Wenlan: Bridging vision and language by large- scale multi-modal pre-training,’’ arXiv preprint arXiv:2103.06561, 2021
2021 arXiv
-
[296]
Soltan, S
S. Soltan, S. Ananthakrishnan, J. FitzGerald, R. Gupta, W. Hamza, H. Khan, C. Peris, S. Rawls, A. Rosenbaum, A. Rumshisky et al. , ‘‘Alexatm 20b: Few-shot learning using a large-scale multilingual seq2seq model,’’ arXiv preprint arXiv:2208.01448, 2022
2022 arXiv
-
[299]
[Online]
BAAI, ‘‘Baai 23,’’ 2023. [Online]. Available: https://2023.baai.ac.cn/about
2023
-
[2020]
Available: https://www.microsoft.com/en-us/research/ blog/turing-nlg-a-17-billion-parameter-language-model-by-microsoft/
[Online]. Available: https://www.microsoft.com/en-us/research/ blog/turing-nlg-a-17-billion-parameter-language-model-by-microsoft/
-
[2022]
Available: https://github.com/microsoft/DeepSpeed
[Online]. Available: https://github.com/microsoft/DeepSpeed
-
[2023]
Available: https://www.databricks.com/blog/2023/04/12/ dolly-first-open-commercially-viable-instruction-tuned-llm
[Online]. Available: https://www.databricks.com/blog/2023/04/12/ dolly-first-open-commercially-viable-instruction-tuned-llm
2023
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.