Key Takeaways
- Xiaomi has released Xiaomi-CocktailASR-1, an open-source speech recognition model.
- The model can identify and transcribe one person’s voice in recordings with multiple speakers.
- Developers can access the model through GitHub and Hugging Face.
Xiaomi has unveiled a new open-source artificial intelligence (AI) model designed to isolate one person’s voice in recordings where multiple speakers are present. The model, named Xiaomi-CocktailASR-1, addresses the 'cocktail party problem'—a common challenge for automatic speech recognition systems where overlapping voices can confuse the technology.
Developed to enhance the accuracy of speech recognition in various scenarios, Xiaomi-CocktailASR-1 works by first taking a short audio sample of the target speaker. It then uses this sample as a reference to process a recording containing multiple speakers, focusing on identifying and transcribing only the selected person’s speech while ignoring others.
According to Xiaomi, the model has achieved state-of-the-art results across several multi-speaker speech recognition benchmarks and outperforms existing systems designed for similar tasks. It maintains competitive performance even in recordings with only one speaker, ensuring that the multi-speaker capabilities do not significantly reduce standard transcription quality.
The model is equipped with a reasoning mode that provides additional information on how it arrived at a transcription, offering transparency and reliability. If the selected speaker is not present in a recording, Xiaomi-CocktailASR-1 returns an empty result, avoiding the transcription of the wrong person.
Xiaomi has a history of releasing open-source AI models, with previous releases including MiMo-V2-Flash and Xiaomi Robotics-0. Xiaomi-CocktailASR-1 is now available for developers to access and further develop through GitHub and Hugging Face.
This development could have significant implications for various applications, including meetings, interviews, and group conversations, where multiple people speak over one another. The technology could also be beneficial in scenarios requiring accurate transcription of specific speakers, such as legal proceedings or customer service calls.





