
ModelQwen3.8 Omni Flash by Qwen
Text+image+audio+video->text
Free credits included. No credit card required.
Qwen3.8 Omni Flash is Alibaba's next-generation native omnimodal model, the first Qwen model built around agentic capabilities with native audio-video understanding. It processes text, image, audio, and video in a single model with a 1M-token context window, supporting long-horizon workflows such as music video creation, film commentary, and meeting transcription with speaker diarization. Built on the Qwen3.8-Flash-Next MoE architecture, it delivers audio performance that exceeds Gemini 3.8 Flash and audio-visual performance approaching it, while keeping text performance comparable to a text-only model of the same size. The model supports 74 languages and 39 Chinese dialects for speech recognition, and includes agentic perception for efficient long-video understanding using significantly fewer tokens than static approaches.