Alibaba releases Qwen3.8-Omni-Flash, with multiple performances approaching Gemini 3.8 Flash

📅 2026-09-18

Abstract:

Alibaba's Qwen team recently released a new generation of native full-modal model Qwen3.8-Omni-Flash, which further combines audio and video understanding with AI agent capabilities.

The model supports four input forms: text, image, audio and video, and provides a context window of 1 million Tokens. Compared with the previous generation Qwen3.5-Omni-Plus, Qwen officials claim that its average score in 29 benchmark tests has increased by more than 25%, while significantly reducing audio and audio and video processing costs.

1789705181_banner-latest.en.webp

One of the biggest changes in Qwen3.8-Omni-Flash is that it does not simply splice together speech recognition, image recognition and other capabilities, but instead natively processes different types of information from the model level. Alibaba hopes to use this to enable the model to not only "understand" or "understand" audio and video content, but to further plan tasks, call tools, and complete actual work after understanding this information.

In terms of audio capabilities, Qwen3.8-Omni-Flash’s performance is particularly outstanding. According to data released by Qwen, this model has surpassed Gemini 3.8 Flash in some audio benchmark tests. For example, in the WildClawBench-MM multi-modal tool call test, Qwen3.8-Omni-Flash scored 71.0 points, while Gemini 3.8 Flash scored 58.9 points; in the DailyOmni test, Qwen scored 85.1 points and Gemini scored 84.0 points; in SpotSoundBench, the two scored 67.2 points and 39.7 points respectively; in the MMAU test, they scored 81.8 points and 76.9 points.

Multi-speaker conference speech recognition is also the focus of this upgrade. The AliMeeting test data released by Qwen shows that the speaker error rate DER of Qwen3.8-Omni-Flash is 3.35, and the word error rate cpWER with speaker attribution is 17.18, while the previous generation Qwen3.5-Omni-Plus is 88.11 and 89.61 respectively. This means the new model changes significantly in this test.

However, Qwen3.8-Omni-Flash is not ahead of Gemini 3.8 Flash in all projects. For example, in the AgenticVBench test, Gemini 3.8 Flash scored 45.0 points and Qwen scored 36.8 points; in OmniGAIA, they scored 78.6 and 74.0 points respectively. In some video understanding tests, Gemini also maintained its advantage. Therefore, Qwen's emphasis on "approaching Gemini" is more reflected in the overall audio and audio and video capabilities, and does not mean an overall lead in all multi-modal projects.

Another important change in Qwen3.8-Omni-Flash is the price. Qwen says its estimated cost per hour for audio input has dropped by more than 98% and the cost per hour for audio plus video input has dropped by more than 93% compared to the previous model. According to the official calculation method, the so-called "hourly" cost is calculated by multiplying the processing price of two minutes of material by 30, in which the video adopts the input condition of 720p and 1 frame per second.

The current international regional price announced by Alibaba Cloud is US$0.15 per million Token input, the input Token hitting the cache is US$0.016, and the output Token per million is US$0.47. Prices in mainland China and some other regions will vary. Since audio and video can be used directly as model inputs, this pricing structure is particularly suitable for scenarios such as meeting recording, long audio analysis, video review, subtitle generation, and processing of large amounts of media content.

The 1 million Token context window further amplifies this advantage. The maximum input length of Qwen3.8-Omni-Flash is close to 992,000 Tokens, the maximum input length in thinking mode is approximately 984,000 Tokens, and the maximum output length reaches 131,000 Tokens. This means that developers can directly hand over very large-scale audio or video content to the model for processing, without having to frequently slice and extract frames as in the past, and then use multiple models to complete speech recognition, visual understanding and text reasoning respectively.

Alibaba Cloud's official information shows that Qwen3.8-Omni-Flash can handle audio input in 113 languages ​​and dialects and supports spatial audio analysis. Through relevant parameters, the model can handle two-channel stereo and four-channel spatial audio. At the same time, the model supports function calling, network search, context caching and other capabilities, and can be directly used in more complex agent workflows.

From the perspective of application direction, the Qwen team focuses on "intelligent delivery" this time. For example, the model can analyze long videos, locate key content, and further call tools to complete tasks such as video editing, content organization, meeting summary, and subtitle generation. Qwen also announced Qwen-MM-Plugins for multi-modal agent development, and plans to launch Qwen-Live-Harness to help developers further build intelligent agent applications based on audio and video input.

Compared with the previous generation Qwen3.5-Omni, the focus of this upgrade is not just to improve the recognition accuracy in the traditional sense, but to try to change the way audio and video AI is used. In the past, multi-modal applications usually required extracting video frames and transcribing audio, and then handing the different results to the language model for processing; Qwen3.8-Omni-Flash tries to unify these links into a native full-modal model as much as possible, and uses 1 million Token contexts to directly process longer original content.

Qwen3.8-Omni-Flash currently provides services through Alibaba Cloud Bailian, supporting regions such as Beijing, Singapore, Hong Kong, China, Tokyo, Japan, Frankfurt, Germany, and Virginia, the United States. The model itself does not directly open the weights like some Qwen previous models. Alibaba mainly opens the supporting multi-modal plug-ins and related development tools this time.

Overall, the core of this Qwen3.8-Omni-Flash update is not just to improve model running scores, but to simultaneously reduce bass video processing costs, expand context scale, and combine audio and video understanding with tool invocation. In particular, the audio input cost has been reduced by more than 98%. If the actual use cost can reach the official calculation level, it will make large-scale automatic analysis of long-term meeting recordings, podcasts, live broadcasts and video materials easier, and further intensify the competition between Qwen and multi-modal models such as Google Gemini.

Related tags

Related articles

Comments

0/500
Captcha (click to refresh)
No comments yet