Abstract:
Some leading artificial intelligence companies are promoting the mobile phone industry to develop devices that can continuously run 100 billion parameter large language models, with the target time pointing to 2028. However, this idea faces two practical challenges: First, running such a large model requires a large amount of memory, and second, artificial intelligence data centers are competing for global memory production capacity, resulting in tight supply of consumer-grade memory and continued price increases.

Qualcomm CEO Amon said in an interview with Fireside Alpha that some companies at the forefront of the artificial intelligence industry are exploring mobile phones that can continuously run hundreds of billions of parameter models. He did not disclose the specific names of these companies, but relevant statements show that artificial intelligence companies' expectations for mobile devices are changing: in the future, mobile phones may no longer be just terminals that occasionally call cloud AI services, but personal computing devices that can continuously run large models locally and actively handle tasks.
However, Amon’s statement was more about describing the goals proposed by industry customers rather than Qualcomm’s officially announced product roadmap. At present, Qualcomm has not yet provided a complete hardware solution that can achieve this goal in 2028, nor has it promised that it will launch a commercial mobile phone with corresponding capabilities by then.
From a technical perspective, the 100 billion parameter model has much higher requirements for mobile phone hardware than currently common on-device AI models. The parameters of a large language model determine the large number of values that the model needs to save and call, and these data usually need to be stored in memory so that the processor can quickly access them during inference.
Take a model with 100 billion parameters as an example. If 4-bit quantization is used, each parameter theoretically only requires 4 bits of storage, so only the model parameters themselves require about 50GB of memory. In actual operation, space must also be reserved for the operating system, other applications, context cache, and temporary data during inference. Therefore, 50GB is not the full memory capacity required by a mobile phone to run the model smoothly.
If the model is further quantized to 2 bits, each parameter only requires 2 bits, and the theoretical storage requirement for model parameters can be reduced to approximately 25GB. However, this solution still requires mobile phones to have a memory capacity far exceeding that of current mainstream products, and extremely low-precision quantization may significantly damage the output quality of the model, especially in complex inference tasks.
It should be noted that the number of parameters does not directly determine the full capabilities of the model. The training data, architecture, training methods and inference techniques of different models will all affect the final performance. It is entirely possible for a small, optimized model to outperform a larger model on a specific task. Therefore, even if the mobile phone cannot directly run the 100 billion parameter model for the time being, it does not mean that it cannot provide an actual use experience close to that of a large model.
Currently, there is still a clear gap between the memory configuration of smartphones and the capacity required by hundreds of billions of parameter models. In the past, some Android flagship phones once provided 24GB LPDDR5X memory. However, with the tight supply of memory and rising prices, mobile phone manufacturers began to re-evaluate the commercial value of high-memory versions. For ordinary consumers, increasing memory capacity means higher hardware costs, and it is difficult for mobile phone manufacturers to determine whether users are willing to pay extra for running larger local AI models.
This creates a contradiction: Artificial intelligence companies hope that mobile phones will have more and more powerful local computing capabilities, but the memory required to achieve this goal is becoming more expensive and difficult to obtain.
The global memory supply is tight, which is closely related to the rapid expansion of artificial intelligence data centers. Large technology companies continue to invest in AI servers, which require large amounts of DRAM, high-bandwidth memory, and high-speed storage devices. Memory manufacturers are also paying more and more attention to artificial intelligence-related products with strong demand and higher profits when arranging production. Traditional memory required for consumer electronic equipment is therefore under supply pressure.
Although the low-power DRAM used in smartphones and the high-bandwidth memory used in AI accelerators are different in product structure and manufacturing requirements, they are still affected by changes in production capacity allocation and market supply and demand in the entire semiconductor industry. The stronger the demand for memory from AI infrastructure, the harder it will be for consumer electronics manufacturers to maintain previous conditions in terms of price and supply.
This pressure has been reflected in mobile phone manufacturing costs. Previously published analysis by market research firm Counterpoint Research showed that in the second quarter of 2026, smartphone memory prices increased by more than 80% from the previous quarter. Among mobile phones at different price points, rising memory costs have a significant impact on the overall material costs.
Counterpoint pointed out that the bill of materials cost of mid-range mobile phones increased by approximately 52% year-on-year, and memory accounted for approximately 40% of the overall bill of materials cost. High-end mobile phones have also been affected, and DRAM has even surpassed system-level chips to become the most expensive single component in some flagship mobile phones.
This means that mobile phone manufacturers not only need to consider whether they have the ability to configure 50GB or more memory for 100 billion parameter models, but also must answer a more realistic question: Are consumers willing to pay for these memories?
If just to run large models, the memory of the mobile phone needs to be increased from the currently common 12GB or 16GB to 64GB or even higher, and the cost of the whole machine may increase significantly. Coupled with more powerful processors, advanced packaging and thermal design, the price of the final product may be further away from the range acceptable to ordinary consumers.
Qualcomm and Apple are trying to alleviate this problem in terms of chip architecture and packaging technology. Related reports mentioned that Apple A20 Pro and Qualcomm's new generation Snapdragon flagship platform are exploring chip designs and memory organization methods that are more conducive to the operation of large models, hoping to improve data access efficiency and reduce data transmission overhead between the processor and memory.
However, improving encapsulation and data access efficiency does not mean that physical memory requirements will automatically disappear. Even if the processor can utilize memory more efficiently, the 100 billion parameter model itself still needs to save a lot of parameter data. To make models of this scale truly adaptable to mobile phones, the industry may need to make further breakthroughs in model compression, inference architecture, and storage access.
One possible solution is more radical quantification technology. By reducing the number of bits used per parameter, the model can run in a smaller memory space. However, quantification does not come without costs. For tasks that require complex logical reasoning, too low numerical accuracy may lead to significant degradation in model performance. Therefore, how to strike a balance between memory usage and model quality will become an important research direction for end-side AI.
Another solution is the hybrid expert model, which is the MoE architecture. Unlike traditional intensive architectures that require calling the entire model for each inference, the MoE model selects part of the expert network to participate in calculations based on specific tasks. With proper design, only a small set of parameters in the model need to be activated each time a token is processed.
If expert offloading technology is further adopted, the system can retain the experts currently needed in DRAM and store the experts that are not needed temporarily in the flash memory. When other experts are needed for the inference process, the relevant data is read from flash memory to memory.
This approach can reduce the amount of data that the model must reside in DRAM at the same time, but at the expense of increased storage access and data transfer overhead. The access speed and latency characteristics of UFS flash memory used in mobile phones are significantly different from DRAM. If the model frequently reads parameters from flash memory, inference speed may be affected.
Apple has also previously studied the solution of using the device's built-in storage to run large language models. This type of technology is expected to reduce dependence on DRAM capacity, but whether it can achieve sufficiently low latency, high throughput and acceptable energy consumption on mobile phones still requires further verification.
In addition, continuous running of large models will also cause heat dissipation and battery life problems. The internal space of mobile phones is limited and they cannot be equipped with large-scale cooling systems like desktop computers or servers. If the processor performs high-intensity AI inference for a long time, the chip temperature may continue to rise, subsequently triggering the frequency reduction mechanism, resulting in performance degradation.
This problem is particularly acute for agents that need to operate around the clock. Traditional chatbots usually start processing tasks after the user makes a request, but the new generation of AI agents may need to continuously monitor information, analyze context, and proactively perform operations at the appropriate time. If the phone must run large models around the clock, not only does the memory need to be available at all times, but the processor's power consumption and battery consumption must also be tightly controlled.
Therefore, even if a future mobile phone can successfully load a 100 billion parameter model, it does not mean that it can continue to run at the ideal speed. Whether the model can truly reside in memory, how many words it can process per second, how much heat it generates during operation, and how much impact it has on battery life are all important indicators for measuring the actual use value.
Compared to directly plugging a 100 billion parameter model into a mobile phone, a more realistic solution may be to run a distilled and optimized medium-sized model.
Model distillation usually takes advantage of the capabilities of a large model to train a smaller model, so that the latter can retain the performance of the large model in a specific task as much as possible while the parameter scale is significantly reduced. Wccftech believes that a properly distilled 20 billion to 30 billion parameter model may approach the performance of a 100 billion parameter model on certain tasks, and is therefore more suitable to become the main force of local AI in future smartphones.
The advantage of this solution is that it does not have to rely on extremely large memory capacity and has the opportunity to provide quite powerful reasoning capabilities. For mobile phone manufacturers, instead of simply pursuing how many parameters can be accommodated, it is better to achieve a better actual experience with limited hardware resources through model architecture, quantification, distillation and inference optimization.
Of course, the performance of different models varies greatly, and a 20 to 30 billion parameter model cannot replace the 100 billion parameter model on all tasks. Capabilities such as complex reasoning, expertise, and long-context processing still depend on the training and architecture design of specific models. But from a productization perspective, being able to complete most daily tasks at reasonable power consumption and cost may be more valuable than pursuing pure parameter scale.
Qualcomm CEO’s statement also reflects that the smartphone industry is looking for new growth directions. As the marginal benefits of traditional hardware upgrades gradually decrease, mobile phone manufacturers hope to use AI agents to redefine how devices are used.
Future smartphones may no longer just be devices waiting for users to open applications and enter commands, but rather personal assistants that can continuously understand needs, arrange tasks, and call different services within the scope of user authorization. Such systems require stronger local reasoning capabilities, as well as more efficient memory management, low-power computing, and security mechanisms.
However, there is still a distance between industry goals and the final product. Amon did not announce the name of the specific partner company, nor did it provide a complete hardware roadmap or mass production schedule. 2028 is more suitable to be understood as the target time that some AI companies hope to achieve this capability, rather than the product release date that Qualcomm has confirmed.
Judging from the current technical conditions, it is not completely impossible for mobile phones to run hundreds of billions of parameter models in 2028, but the implementation method may be different from what people imagine by "loading all the complete models into the mobile phone memory." More aggressive quantization, MoE architecture, expert offloading, dedicated memory design, and model distillation may all be part of achieving the goal.
At the same time, global memory supply and price trends will directly affect the commercial viability of this direction. If the supply of consumer-grade DRAM is squeezed by AI data centers for a long time, even if it is technically possible to manufacture mobile phones with 64GB or more memory, manufacturers may not be willing to launch such products on a large scale.
Therefore, Qualcomm’s vision of a 100-billion-parameter mobile phone really needs to be solved not just in terms of chip performance, but also in the comprehensive balance between model size, memory capacity, manufacturing cost, heat dissipation, battery life and consumer demand. For ordinary users, the most noteworthy thing in the future may not be whether the mobile phone can run a model with 100 billion parameters, but whether it can provide reliable, fast and truly useful local AI services without significantly increasing the price or seriously sacrificing battery life.
Comments