Few technologies have climbed as fast as large language models, the systems behind ChatGPT, Claude, Gemini and their rivals. Yet for all the headlines, the field has lacked something basic: a shared map of what these machines can actually do. A new systematic review published in Neural Computing and Applications by Mohamed Nazih Omri of the University of Sousse and Borhen Louhichi of Imam Mohammad Ibn Saud Islamic University attempts to draw that map, sorting the sprawling LLM landscape into a rigorous taxonomy and backing it with comparative analysis of the leading proprietary and open-source systems, from GPT-4 and Gemini to Claude, LLaMA and DeepSeek.
The paper’s central organizational contribution is a six-part taxonomy of language model capabilities: Generation, Extraction, Classification, Transformation, Problem-solving and Comprehension. The categories are deliberately functional rather than technical. Generation covers the production of new text, code and media; Extraction pulls structured information out of unstructured documents; Classification assigns labels to inputs; Transformation rewrites content between formats, styles or languages; Problem-solving encompasses reasoning, planning and mathematical work; and Comprehension describes the models’ ability to interpret meaning across long and complex contexts. The authors argue that this framework lets researchers compare models across vendors and scales on equal footing, instead of relying on marketing benchmarks that shift from release to release.
Beneath the taxonomy sits a more ambitious theoretical claim. The review introduces two formal constructs, an Emergence Function and an Architectural Efficiency Coefficient, intended to describe how capabilities appear as models scale and how efficiently a given architecture converts parameters and energy into performance. The Architectural Efficiency Coefficient is defined as performance divided by the product of parameters and energy, multiplied by the logarithm of throughput. In an illustrative calculation, the authors estimate a coefficient of roughly 2.30 for mixture-of-experts architectures, using the fact that Mixtral 8x7B reaches about 92 percent of GPT-4’s MMLU score while activating only 13 billion of its 47 billion parameters per token. Crucially, the authors are candid that these formulations are presented as testable hypotheses, not validated laws, and an appendix spells out exactly what would be needed to confirm them: controlled training runs, identical hardware, direct energy measurements and published replication code.
The survey traces the technical lineage that produced today’s giants. The transformer architecture, introduced in 2017 with the paper ‘Attention Is All You Need’, replaced the recurrent networks of the previous decade with self-attention mechanisms that process entire sequences in parallel. Scaling laws published in 2020 showed that model performance improves predictably with parameters, data and compute, setting off the race toward ever-larger systems. The review follows that arc through BERT and its bidirectional pretraining, the GPT series’ few-shot learning, instruction tuning with human feedback, and the recent shift toward inference-time scaling, in which models such as OpenAI’s o1 and o3 spend more computation at answer time to chain together reasoning steps.
Training methodology receives its own systematic treatment. The authors describe the now-standard multi-stage pipeline: massive self-supervised pretraining on trillion-token corpora, followed by supervised fine-tuning on curated instruction data, followed by reinforcement learning from human feedback to align outputs with human preferences. They highlight curriculum learning, in which training data is ordered from simple to complex, and efficient architectural paradigms such as mixture-of-experts sparsity, which lets models grow in total size without a proportional rise in per-token cost. Techniques like low-rank adaptation and quantized fine-tuning, which allow large models to be customized on modest hardware, feature prominently as the field’s answer to the crushing economics of full retraining.
The empirical comparisons reveal a striking convergence. Proprietary flagships from OpenAI, Google DeepMind and Anthropic still lead many benchmarks, but open-weight models have closed much of the gap. DeepSeek’s R1 model, which uses reinforcement learning to incentivize reasoning, demonstrated that open systems can match frontier reasoning performance at a fraction of the training cost. Meta’s Llama family, Alibaba’s Qwen series, Mistral’s compact 7-billion-parameter model and the Falcon models from the Technology Innovation Institute show that the open ecosystem now spans everything from edge-deployable assistants to near-frontier generalists with 128,000-token context windows. The review frames this as a structural shift: capability is no longer the sole property of closed labs.
Applications receive equally broad coverage. In healthcare, models such as BioGPT and Med-PaLM have shown the ability to encode clinical knowledge, while synthetic medical text generation has improved diagnostic code classification accuracy by up to 17.8 percent in reported experiments. In education, AI tutors built on these models promise personalized instruction, though the authors note open questions about learning outcomes. In law, GPT-4-based systems are already used for document analysis. Creative domains, software development through Code Llama-style models, and autonomous agents that combine language models with tools, memory and planning round out the survey of real-world deployment.
The review is notably unsparing about the field’s problems. Hallucination, the confident generation of false statements, remains a fundamental weakness tied to the models’ statistical nature. Bias embedded in training data propagates into outputs, and mitigation techniques are still immature. Computational cost and the carbon footprint of training large models raise sustainability concerns, even as analysts project that emissions may plateau and shrink with efficiency gains. Privacy risks persist, with research demonstrating that training data can be extracted from deployed models. The authors also cite the ‘stochastic parrots’ critique, which questions whether scale alone can deliver genuine understanding, and they catalog jailbreaking and adversarial attacks that undermine safety guardrails. Regulatory frameworks, including the EU AI Act, are presented as an emerging constraint on deployment.
Looking forward, the survey identifies several trends it expects to shape the next phase. Small language models, distilled and compressed versions of their giant cousins, are positioned as the pragmatic future for many applications, offering most of the utility at a fraction of the cost. Multimodal integration, exemplified by Gemini’s native handling of text, images and audio and by embodied models such as PaLM-E, is dissolving the boundary between language and perception. Retrieval-augmented generation, which grounds model outputs in external documents at query time, offers a partial remedy for hallucination and stale knowledge. The authors also point to interpretability research, from mechanistic analyses of transformer circuits to attribution methods, as essential for building systems whose decisions can be trusted and audited.
What makes the paper unusual among surveys is its insistence on intellectual honesty about its own contributions. The mathematical framework for emergence and efficiency is offered as a hypothesis-generating scaffold, with the authors explicitly listing what they did not do: no model was trained from scratch, no energy consumption was measured directly, and no statistical significance testing was performed. That transparency, combined with the six-category taxonomy and the synthesis of architectural trade-offs, positions the review as both a reference work for practitioners navigating a crowded market and a starting point for the empirical studies that will determine whether the field’s scaling obsession gives way to something smarter: architectures that squeeze more capability out of every parameter and every joule.
Subject of Research: A systematic taxonomy and empirical analysis of large language model architectures, training methods and applications
Article Title: Decoding the Giants: a systematic taxonomy and empirical analysis of large language model architectures, training, and applications
Article References: Omri, M. N., & Louhichi, B. (2026). Decoding the Giants: a systematic taxonomy and empirical analysis of large language model architectures, training, and applications. Neural Computing and Applications, 38(19), Article 759. https://doi.org/10.1007/s00521-026-12365-9
Image Credits: AI Generated
DOI: 10.1007/s00521-026-12365-9
Keywords: large language models, transformers, taxonomy, emergence, scaling laws, mixture of experts, reinforcement learning from human feedback, small language models, multimodal AI, retrieval-augmented generation, hallucination, AI ethics
Cite Scienmag News
Denise Maddox. (September 30, 2026). New Map of AI Giants Sorts the World’s Language Models Into Six Powers. Scienmag. https://scienmag.com/new-map-of-ai-giants-sorts-the-worlds-language-models-into-six-powers/
Denise Maddox. "New Map of AI Giants Sorts the World’s Language Models Into Six Powers." Scienmag, 30 September 2026, https://scienmag.com/new-map-of-ai-giants-sorts-the-worlds-language-models-into-six-powers/. Accessed 30 September 2026.
Denise Maddox. "New Map of AI Giants Sorts the World’s Language Models Into Six Powers." Scienmag. September 30, 2026. https://scienmag.com/new-map-of-ai-giants-sorts-the-worlds-language-models-into-six-powers/

