Musk bets on ultra-large-scale AI computing power, xAI plans to deploy more than 1.2 million NVIDIA GPUs

📅 2026-09-25

Abstract:

XAI, Elon Musk’s artificial intelligence company, is further expanding its data center scale and plans to deploy more than 1.2 million Nvidia AI GPUs in the next few years. As the Colossus data center continues to expand, xAI hopes to train and run the next generation Grok model through large-scale computing infrastructure, and launch more intense AI infrastructure competition with competitors such as OpenAI.

xAI’s current largest AI infrastructure project is the Colossus data center located in Memphis, Tennessee, USA. Construction of the project begins in 2024, and the first phase, Colossus 1, is already operational. Musk has previously called Colossus the "computing power super factory" of xAI, whose core goal is to centrally deploy a large number of GPUs as quickly as possible.

Colossus initially used NVIDIA H100 GPUs. The previously announced scale of the first phase is about 200,000 H100s. With subsequent expansion, the data center has begun to add more advanced Blackwell series GPUs. The latest information shows that in addition to the original H100, Colossus 1 also deploys about 30,000 Nvidia GB200s, so the current number of GPUs in Colossus 1 has reached about 230,000.

xAI’s second phase, Colossus 2, is even larger. The project is also located in Memphis and plans to deploy more than 500,000 NVIDIA GPUs. According to current plans, Colossus 2 will become one of the world's largest AI data center projects, and its computing power will far exceed that of traditional supercomputers.

If all Colossus 1, Colossus 2 and subsequent construction projects are included, the number of NVIDIA GPUs xAI plans to use in the future will reach approximately 1.2 million or more. This number does not mean that all these GPUs have been installed, but rather corresponds to the overall scale of xAI’s future data center expansion plans.

xAI’s radical expansion of the number of GPUs is directly related to the rapid growth in demand for computing power for AI model training. The parameter scale, training data, and reinforcement learning process of large-scale language models require more and more computing resources, and the computing power required for model training and inference is evolving from the scale of traditional data centers to the scale of supercomputing clusters.

In this case, having a large number of GPUs has become an important foundation for AI companies to expand their model capabilities. Musk has repeatedly emphasized that xAI needs to expand its computing capabilities as quickly as possible and reduce its dependence on third-party cloud computing platforms by building its own supercomputing clusters.

The speed of construction of Colossus is also a major feature of the xAI infrastructure strategy. The first-stage data center was renovated on the basis of an abandoned industrial facility, and xAI expanded it into a large-scale AI training facility with hundreds of thousands of GPUs in a short period of time through a large number of prefabricated equipment, parallel construction, and rapid deployment of GPUs.

xAI is also continuing to expand Colossus’ data center building and power supply capabilities. Due to the extremely high power consumption of modern AI GPUs, the expansion of data centers is not just about simply increasing the number of servers. They must also build a large amount of power, cooling, network and storage infrastructure at the same time.

xAI has already faced power supply problems before. The power consumption of Colossus 1 reaches hundreds of megawatts, and as the number of GPUs further increases, its energy requirements will also increase significantly. In order to support subsequent expansion, xAI has been looking for new sources of power and building corresponding power generation and distribution infrastructure.

As NVIDIA Blackwell and subsequent Rubin series GPUs gradually enter the market, xAI's data center will continue to undergo hardware upgrades. Compared with H100, the new generation of GPU can provide higher AI computing performance, but the power consumption of a single GPU is also increasing. Therefore, future large AI data centers must solve the problems of computing density and energy supply at the same time.

Musk also previously stated that xAI plans to significantly increase the overall power capacity of the data center around 2027 and expand the computing power to several times its previous scale. Relevant plans show that the total power scale of xAI's future data centers may reach several gigawatts or even 10 gigawatts.

Such a huge infrastructure means that xAI is trying to build an AI "super factory" similar to that of large cloud computing companies. GPU servers, network equipment, storage systems, power facilities and liquid cooling systems will be combined into a huge computing cluster for training and running AI models such as Grok.

This strategy also makes the competition between xAI and OpenAI increasingly direct. OpenAI has also been building AI data centers on a large scale in recent years and has established a large-scale infrastructure partnership with NVIDIA. NVIDIA has announced plans to provide OpenAI with at least 10GW of AI systems, corresponding to millions of GPUs, and plans to gradually invest up to US$100 billion in OpenAI with infrastructure deployment.

OpenAI’s plan means that the AI ​​industry is entering a new stage with “gigawatt-level data centers” as the basic competitive unit. In the past, competition among AI companies focused more on model algorithms, talent, and data, but now the computing infrastructure itself has become an important factor in determining the speed and scale of model training.

xAI chooses to build its own large-scale Colossus data center, which allows the company to more directly control GPU deployment, network structure, data center software and training infrastructure. For companies that need to continuously train large-scale models, this approach can reduce dependence on external cloud platforms while allowing engineering teams to optimize hardware and software for their own models.

However, the scale of 1.2 million GPUs is still a future goal, not the number that xAI has currently deployed. The speed of data center construction, GPU supply, power supply, cooling facilities, network equipment and capital investment will all affect the final deployment speed. Therefore, the actual number of GPUs running and the completion time may still vary.

Regardless of whether the final figures can reach the scale currently planned, the Colossus being built by xAI has shown that the AI ​​infrastructure competition is entering a new stage. As OpenAI, Google, Meta, Microsoft and other AI companies build gigawatt-level data centers, the competition among AI models in the future will largely evolve into a competition between data center scale, power supply and advanced GPU acquisition capabilities.

Related tags

Related articles

Comments

0/500
Captcha (click to refresh)
No comments yet