...

A Systematic Survey of Large Language Models: Architectures, Training Paradigms, Capabilities, and Limitations

Abstract

LLMs represent an unprecedented breakthrough in the development of artificial intelligence. They exhibit extraordinary capabilities in both understanding and generating human language as well as reasoning using knowledge from diverse areas. By employing transformer architecture, which is scalable from a few million to hundreds of billions of parameters through self-supervised training on web corpus data, they have achieved new levels of performance across virtually all kinds of NLP tasks. This paper provides an organized and thorough review of LLMs in terms of their structure (encoder-only, decoder-only, encoder-decoder, and mixture-of-experts) and training methods (masked vs causal language modelling; instruction-tuning; reinforcement learning via human feedback; parameter-efficient fine-tuning); a review of the empirical scaling laws that have been established for these models, as well as exploring the notion of emergent abilities such as in-context learning, chain-of-thought reasoning, code generation, and retrieval-based generation and persistent limitations such as hallucinations, societal bias, costs associated with computation and alignment problems. Other items reviewed include evaluation standards, benchmark measures, and real-world use cases. Various open research questions and future directions for research are discussed to support future investigations into this area.

Authors

Files

Link of Paper

Seraphinite AcceleratorOptimized by Seraphinite Accelerator
Turns on site high speed to be attractive for people and search engines.