Key Notes
- Alibaba released Qwen3.8-Omni-Flash, its first omni-modal model built around agentic capabilities, combining native audio and video understanding with reasoning and tool use in a 1 million token context window.
- The company said audio input costs dropped more than 98% and audio-visual costs more than 93% compared with its predecessor, with benchmark scores it says approach or exceed Google's Gemini 3.8 Flash on several audio-focused evaluations.
- The model is available only through Alibaba's hosted platforms, with no open weights announced, and all performance and pricing figures come from Alibaba's own launch materials pending independent review.
Alibaba has released Qwen3.8-Omni-Flash, describing it as the Qwen team’s first omni-modal model built around agentic capabilities, combining native audio and video understanding with reasoning and tool use in a single system that can watch or listen to content, plan a task, and act on it.
The model accepts text, images, audio and video and returns text, with a 1 million token context window. It is available now through Qwen Chat, Alibaba Cloud’s Model Studio API and QwenCloud, on an OpenAI-compatible endpoint, but is not open-weight, meaning it can only be run through Alibaba’s hosted infrastructure rather than downloaded and self-hosted.
A Sharp Cut in Audio and Video Costs
Alibaba’s headline claim centers on price rather than raw capability. The company said audio input now costs more than 98% less per hour than its predecessor, Qwen3.5-Omni-Plus, while combined audio-visual input costs have dropped more than 93%. Alibaba’s own account put the reduction in video input costs at roughly 89%. Model Studio lists the API at $0.15 per million input tokens and $0.47 per million output tokens, with cached tokens priced at $0.016 per million, matching the list price of Qwen3.8-Flash, the text-and-vision model launched alongside it.
On benchmarks, Alibaba reported average scores rising more than 25% across 29 evaluations compared with Qwen3.5-Omni-Plus, including a 19.5 point average gain on agent-focused tests. On OmniVideoBench specifically, the company said the model improved from 63.4 to 67.8 while cutting token consumption nearly in half, from 145,736 tokens to 79,117. All of these figures come from Alibaba’s own launch materials, and no independent evaluation had been published as of this writing.
How It Compares to Gemini
Alibaba positioned the new model against Google’s Gemini 3.8 Flash, saying its audio-video capability approaches Gemini’s while its overall audio performance surpasses it. According to figures Alibaba published, Qwen3.8-Omni-Flash scored 71.0 against 58.9 on WildClawBench-MM and 67.2 against 39.7 on SpotSoundBench, two audio-centric evaluations, while trailing on at least one pure video-reasoning benchmark, AgenticVBench. Independent commentators have noted that Alibaba’s comparisons are selective and vendor-reported, and that third-party verification is still pending.
Built for Longer, Tool-Using Workflows
Alibaba said the model is designed for agentic evidence gathering rather than simple description, citing use cases such as auto-editing vlogs, translating short videos, summarizing meetings and turning full-length films into structured recap reports. The company also released Qwen-MM-Plugins as open-source integration infrastructure, letting compatible agent harnesses such as Claude Code, Codex and Gemini CLI expose multimodal operations as callable tools through an optional Model Context Protocol server.
The release adds to a broader pattern among AI labs of cutting prices on multimodal and reasoning-heavy workloads to expand usage, a trend also visible in Anthropic’s recent Claude Fable 5.1 release, which similarly emphasized lower costs for long, iterative sessions. Alibaba has not disclosed regional availability, rate limits, or data retention policies for the new model beyond what is documented in its current API reference, and has not said when or whether open weights for Qwen3.8-Omni-Flash will follow, as they did roughly a month after its underlying Qwen3.8-Flash-Next base model shipped in August.
QwenCloud’s documentation lists a maximum of 991,000 input tokens and 131,000 output tokens within the model’s overall 1 million token window, with reasoning capped at 262,000 tokens, and the service accepts video files up to two hours and 2 gigabytes by URL and audio files up to three hours.
Because the model has no generated-speech output, Alibaba’s documentation points developers who need synthesized speech to the separate Qwen3.5-Omni model instead. The release drew a large discussion thread on Hacker News, where commenters focused on how closely the pricing and stated benchmarks track Google’s Gemini 3.8 Flash, a comparison Alibaba highlighted directly in its own launch materials.
Alibaba has released the Qwen family under a mix of open and closed terms over the past year, with some Qwen3 variants shipped as open weights that developers can download and run themselves, while flagship or specialized releases like Qwen3.8-Omni-Flash have increasingly launched as hosted-only API products first. That pattern mirrors a broader industry shift in which labs keep their most capable or most expensive-to-serve multimodal models behind an API rather than releasing the weights immediately, even as they continue to open-source smaller or earlier-generation models.
Disclaimer: AIstify is an independent media brand owned and operated by NuvexMedia LLC, publishing news, research, and insights on artificial intelligence, emerging technologies, automation, and related industries. NuvexMedia LLC invests in and collaborates with companies across the AI, technology, software, and digital innovation sectors. These relationships do not influence AIstify’s editorial coverage, and the publication maintains full editorial independence to provide accurate, timely, and objective information. © 2026 NuvexMedia LLC. All rights reserved. This content is for informational purposes only and should not be considered legal, tax, investment, financial, or other professional advice.