ModelsAlibaba
Alibaba's omnimodal flash model takes audio and video into a million tokens of context
1Mtoken context across text, image, audio and video
Alibaba's Qwen team released Qwen3.8-Omni-Flash, a native omnimodal model that takes text, images, audio and video in one pass and returns text, with a 1M token context window and tool use built in. Alibaba reports an average improvement of more than 26% across 30 evaluations against Qwen3.5-Omni-Plus, and results close to Gemini 3.8 Flash on audio and video reasoning. It is API only, priced on Alibaba Cloud Model Studio at $0.15 per million input tokens and $0.47 per million output on the international endpoint, and no weights have been published, which is a change of habit for a team whose last three releases were downloadable.