Developer turns iPhone 17 Pro Max into MacBook Pro’s second GPU

📅 2026-10-03

Abstract:

An experiment completed by a developer shows that connecting the iPhone 17 Pro Max to a MacBook Pro equipped with an M4 Pro chip through customized software can significantly improve the operating efficiency of the local large language model. When running the Qwen3.8-27B model, the system's prefill performance is improved by up to 44%.

The background of this experiment is that large language models have extremely high requirements on video memory and system memory capacity. Although the MacBook Pro equipped with the M4 Pro chip has 24GB of unified memory, it is still limited by memory capacity and context window size when running a large model with 27 billion parameter levels. Developers therefore tried to use iPhone 17 Pro Max as an additional computing node to run models in conjunction with MacBook Pro through the USB-C interface.

According to the plan announced by the developer, the MacBook Pro is responsible for processing the computing tasks from layer 1 to layer 40 in each 256 Token batch, and transmits the intermediate activation data to the iPhone in real time. Subsequently, the iPhone 17 Pro Max uses the GPU in the A19 Pro chip to continue performing layer 41 to layer 64 operations, while the Mac starts processing the next batch of data simultaneously. Through this pipelined division of labor, two devices can participate in the model inference process at the same time.

Experimental results show that compared to running the model only on MacBook Pro, the pre-filling speed is significantly improved after introducing iPhone. Under the 8K context window, the processing speed increases from 132 Tokens per second to 177 Tokens, an increase of 35%; under the 16K context window, the performance increases from 109 Tokens per second to 157 Tokens, an increase of 44%; in the 32K context window scenario, the processing speed increases from 101 Tokens per second to 130 Tokens, an increase of approximately 29%.

In addition to GPU collaborative computing, the neural network engine in the A19 Pro chip also participates in some tasks. The developers stated that every 16K historical context will be converted into a dedicated neural network model for processing. At the 140,000 Token context scale, the writing time of a single Token is shortened from 279 milliseconds when relying only on the GPU to 176 milliseconds, further improving the operating efficiency in long context scenarios.

However, this solution is not without limitations. Developers pointed out that the iPhone mainly helps improve the performance of the pre-filling stage, while in text generation tasks below 64K, most of the actual reasoning work is still mainly done by the Mac. In addition, this solution requires specially developed software support and is currently not suitable for direct deployment by ordinary users.

Nonetheless, this experiment demonstrates the potential of consumer-grade Apple devices for local AI computing. As the performance of mobile chips continues to improve, a more flexible collaborative computing system may be formed between smartphones, tablets, and personal computers in the future, thereby reducing the reliance on a single device hardware configuration to run large artificial intelligence models.

Related tags

Related articles

Comments

0/500
Captcha (click to refresh)
No comments yet