Qwen’s New Omni-Modal Agent Turns Audio and Video Into Action

Alibaba’s Qwen team has released Qwen3.8-Omni-Flash, a model that brings audio-video understanding, reasoning, and tool use into one system. The model accepts text, images, audio, and video, then returns text while moving through a workflow built to understand content, plan a task, execute with tools, and deliver the result.
“Meet Qwen3.8-Omni-Flash, Qwen’s first omni-modal model built around agentic capabilities!” the Qwen team wrote on September 18, 2026. That description points to the model’s central shift: it is designed not only to interpret rich media, but also to gather evidence, make plans, and act across several rounds.
A Million-Token Window for Long, Rich Tasks
Qwen3.8-Omni-Flash is built on the Qwen3.8-Flash-Next architecture and carries a 1 million token context window. QwenCloud lists 991,000 maximum input tokens and 131,000 maximum output tokens, with a maximum reasoning length of 262,000 tokens.
Those figures give the model room to work across long audio and video inputs while keeping the task, evidence, reasoning, and final response inside one workflow. The agent starts from the question, gathers evidence over several rounds, and uses tools as it works toward the result.
The model also connects media understanding with action instead of treating them as separate steps. Audio-video understanding, reasoning, and tool use are integrated into one model, creating a single path from incoming content to an answer.
Stronger Agent Results With Less Token Use
The evaluation results show gains across both media understanding and agent tasks. On OmniVideoBench, accuracy rises from 63.4 to 67.8, while token use drops from 145,736 to 79,117, a reduction of about 45.7%.
Across 29 evaluations, the average score improves more than 25% over Qwen3.5-Omni-Plus. The individual results show where the model adds force:
- WildClawBench-MM improves by 36.5 points.
- AgenticVBench improves by 22.3 points.
- UniClawBench reaches 69.6.
- LongAudioSpan gains 8.3 points.
- OmniVideoBench gains 9.6 points.
- OmniCap-IF CSR improves by 8.5 points.
- OmniCap-IF ISR improves by 14.1 points.
The agent gains reach an average of 19.5 points across WildClawBench-MM and UniClawBench. Audio-visual performance is close to Gemini 3.8 Flash, while overall audio performance is above Gemini 3.8 Flash.
These results matter because the model is handling more than recognition. It must understand the question, inspect content, gather evidence, plan the task, use tools, and produce the result. The combination of higher scores and lower token use gives the agent workflow a sharper edge.
Hosted Access, Broad Inputs, and Developer Tools
Qwen3.8-Omni-Flash is deployable as a hosted API today and is live on QwenCloud, Alibaba Cloud Model Studio, and Qwen Studio. The API follows both the DashScope and OpenAI protocols, with support for Chat Completions and the Responses API.
Its supported capabilities include function calling, web search, structured outputs, context caching, and batch calls. Calling the API takes a few lines with the OpenAI SDK, giving the model a direct route into existing API workflows.
QwenCloud pricing is $0.15 per 1 million input tokens and $0.47 per 1 million output tokens. Implicit cache hits cost $0.016 per 1 million tokens. Audio input costs over 98% less per hour compared to previous models, while audio-visual input costs over 93% less per hour.
The input limits are built for long recordings and video files. Video files up to 2 hours and 2 GB by URL are supported, and audio files up to 3 hours are supported. Audio input is available in 113 languages and dialects, while video sampled at up to 15 fps produces stable results.
Availability spans 6 regions: Beijing, Singapore, Hong Kong, Tokyo, Frankfurt, and Virginia. That regional reach pairs with a hosted deployment model, letting the same system handle text, images, audio, video, reasoning, and tools through one service.
Open-Source Pieces Around a Closed Launch
No open weights were announced at launch for Qwen3.8-Omni-Flash. The base model shipped with open weights in August 2026, but the new omni-modal agent itself arrives through hosted access.
The Qwen team is open-sourcing Qwen-MM-Plugins and Qwen-Live Harness, adding open components around the model’s audio, video, and tool-use direction. That gives the release a two-part shape: a hosted model for immediate API deployment, plus open tools that expand the surrounding workflow.
Qwen3.8-Omni-Flash now puts the focus on what an omni-modal model can do after it understands content. With a 1 million token context window, long audio and video support, tool use, web search, structured outputs, and evidence gathering across several rounds, Qwen is pushing its first agentic omni-modal model toward tasks that move from perception to action.




