THE ROLE OF CLUSTERING TECHNIQUES AND MIXTURE OF EXPERTS ARCHITECTURE BEHIND THE SUCCESS OF DEEPSEEK

Main Article Content

suriya pumchalerm
Nopakhun Nanthasenee
Utsanee Yeesoonkaew

Abstract

This article is an academic review and synthesis of concepts, with the objective of presenting guidelines for the development of artificial intelligence (AI), which is currently confronted with significant challenges, particularly constraints related to computational complexity, intensive resource consumption, and system scalability. Consequently, AI system designers have increasingly focused on developing conceptual frameworks and architectural paradigms that can maintain high computational performance while reducing resource costs. Among the approaches that have attracted considerable attention are clustering techniques, one form of machine learning, and the Mixture of Experts (MoE) architecture, which functions to analyze structural and computational interrelationships within models. Clustering techniques aim to discover latent patterns and intrinsic structures within data by leveraging deep representations to group instances with similar characteristics. This mechanism helps reduce the complexity of input data and enhances processing efficiency. In parallel, the MoE architecture embodies a conditional computation paradigm in which a model is divided into multiple specialized sub-networks, or “experts,” and a routing mechanism selectively activates only those components most relevant to a given input. Such an approach enables more efficient resource utilization and improves the scalability of large-scale systems. The integration of clustering techniques with the MoE architecture therefore yields a computational framework capable of reducing computational burden, increasing flexibility, and sustaining high model performance. This conceptual integration reflects a key direction in the development of the AI system “DeepSeek,” which emphasizes efficiency, scalability, and sustainable deployment across diverse application contexts.

Article Details

How to Cite
[1]
suriya pumchalerm, N. Nanthasenee, and U. Yeesoonkaew, “THE ROLE OF CLUSTERING TECHNIQUES AND MIXTURE OF EXPERTS ARCHITECTURE BEHIND THE SUCCESS OF DEEPSEEK”, JSCI-SBU, vol. 6, no. 1, p. e263912, Jun. 2026.
Section
Academic Article

References

A. Vaswani et al., “Attention Is All You Need,” in Advances in Neural Information Processing Systems, vol. 30, 2017, pp. 5998–6008.

T. B. Brown et al., “Language Models are Few-Shot Learners,” in Advances in Neural Information Processing Systems, vol. 33, 2020, pp. 1877–1901.

N. Shazeer et al., “Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer,” in Proc. Int. Conf. Learning Representations (ICLR), 2017.

A. Paszke et al., “PyTorch: An Imperative Style, High-Performance Deep Learning Library,” in Advances in Neural Information Processing Systems, vol. 32, 2019, pp. 8024–8035.

R. Bommasani et al., “On the Opportunities and Risks of Foundation Models,” arXiv preprint arXiv:2108.07258, 2021. [Online]. Available: https://arxiv.org/abs/2108.07258. [Accessed: Feb. 2, 2026].

I. Goodfellow, Y. Bengio, and A. Courville, Deep Learning. Cambridge, MA, USA: MIT Press, 2016.

C. M. Bishop, Pattern Recognition and Machine Learning. New York, NY, USA: Springer, 2006.

A. K. Jain, “Data clustering: 50 years beyond k-means,” Pattern Recognition Letters, vol. 31, no. 8, pp. 651–666, Jun. 2010, doi: 10.1016/j.patrec.2009.09.011.

T. Zhang, H. Guo, W. Lu, T. Dai, S.-T. Xia, and J. Wang, “SPARSEEVAL: Efficient Evaluation of Large Language Models by Sparse Optimization,” in Proc. Int. Conf. Learning Representations (ICLR), 2026. [Online]. Available: https://openreview.net/forum?id=CZAzAedGSV. [Accessed: Feb. 10, 2026].

M. Mitchell, Complexity: A Guided Tour. Oxford, U.K.: Oxford University Press, 2009.

S. Russell and P. Norvig, Artificial Intelligence: A Modern Approach, 4th ed. Hoboken, NJ, USA: Pearson, 2020.

W. Fedus, B. Zoph, and N. Shazeer, “Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity,” Journal of Machine Learning Research, vol. 23, no. 120, pp. 1–39, 2022.

S. Xu, “The Complete Guide to DeepSeek Models: V3, R1, V4 and Beyond,” BentoML Blog, Apr. 24, 2026. [Online]. Available: https://www.bentoml.com/blog/the-complete-guide-to-deepseek-models-from-v3-to-r1-and-beyond. [Accessed: Feb. 10, 2026].

J. Dean and S. Ghemawat, “MapReduce: Simplified Data Processing on Large Clusters,” in Proc. 6th Symp. Operating Systems Design and Implementation (OSDI), San Francisco, CA, USA, Dec. 2004, pp. 137–150.

J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding,” in Proc. 2019 Conf. North American Chapter Assoc. Comput. Linguistics: Human Language Technologies (NAACL-HLT), Minneapolis, MN, USA, 2019, pp. 4171–4186, doi: 10.18653/v1/N19-1423.

DeepSeek-AI, “DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model,” arXiv preprint arXiv:2405.04434, 2024. [Online]. Available: https://arxiv.org/abs/2405.04434. [Accessed: Feb. 2, 2026].

D. Dai et al., “DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models,” in Proc. 62nd Annu. Meeting Assoc. Comput. Linguistics (ACL), Bangkok, Thailand, Aug. 2024, pp. 1280–1297, doi: 10.18653/v1/2024.acl-long.70.