Key Takeaways
- Qwen3.8-Omni-Flash, a next-generation AI model, has been launched with a 1-million-token context window.
- The model improves performance across 29 evaluations by more than 25% compared to its predecessor.
- Qwen3.8-Omni-Flash supports real-time meetings, video editing, and content creation.
Qwen, a leading AI platform, has unveiled Qwen3.8-Omni-Flash, a cutting-edge model designed to handle text, images, audio, and video. This new version boasts a 1-million-token context window, significantly enhancing its ability to process complex tasks.
According to Qwen, the new model has achieved a 25% improvement in average scores across 29 evaluations, surpassing its predecessor, Qwen3.5-Omni-Plus. The company claims that audio input costs have dropped by over 98%, and audio-visual input costs have reduced by more than 93%.
Qwen3.8-Omni-Flash excels in handling long videos, where its agent can selectively process frames based on user questions. This approach increases accuracy and reduces token usage. For instance, on the OmniVideoBench, accuracy improved from 63.4 to 67.8 while token usage was reduced by 45.7%.
The model is also adept at analyzing specific elements such as characters, camera shots, lighting, and sound, based on user queries. This feature makes it particularly useful for content creators and editors who need detailed insights into their projects.
In the realm of meetings, Qwen3.8-Omni-Flash can transcribe discussions, generate minutes, and extract action items. When connected to tools, it can automate tasks such as sending emails, organizing tasks, or initiating coding based on meeting requirements.
For video editing and content creation, Qwen3.8-Omni-Flash offers a comprehensive solution. It can analyze music before planning and creating music videos, and its video translation workflow handles transcription, translation, voice cloning, dubbing, audio mixing, and final review.
The model also supports real-time processing, making it suitable for live audio and video inputs. It can combine visual information with spatial sound to determine the origin of sounds, aiding in localization and navigation. Qwen3.8-Omni-Flash supports speech recognition in 74 languages, including Urdu and Punjabi, and speech generation in 29 languages.
To facilitate multimodal agent workflows, Qwen has expanded Qwen-MM-Plugins and open-sourced Qwen-Live Harness, providing tools for long-term memory, task delegation, and real-time interaction.
Qwen3.8-Omni-Flash is now available through the Qianwen AI Platform, targeting workflows including video editing, music video creation, film commentary, audio-visual summarization, and real-time conversations.





