LIVE Watch Now Punjabi News Channel from Pakistan
Breaking
PM Directs Facilitation Desks to Assist Citizens with Fuel Relief SchemeTaiwan’s NSTC Pushes for Quantum Breakthroughs in Chip ManufacturingJazzCash Boosts Savings Among BISP BeneficiariesPunjab Forest Department Recovers Rs 9 Billion Worth of LandSBP Prohibits Bank Employees from Applying for Affordable Housing SchemeAudit Reveals Failures in Double-Shift School ProgramGmail Settings to Combat AI Scams and SpamFinance Division Defends Performance of State-Owned EnterprisesGold prices surge by Rs7,900 per tola in PakistanFinance Division refutes SOE financial decline claimsPML-N Nominates Dr Najeeb Naqi for AJK PresidencyCantonment Board Workers Tackle Maintenance Tasks in KarachiJI warns talks with government may end over petroleum levyPakistan Extends Airspace Ban on Indian Aircraft Until October 24Lahore High Court Imposes Rs. 100,000 Fine on Woman for Contempt PetitionLahore to Launch Vehicular Emission Scanning SystemLahore Traffic Police to Equip 2.3 Million Motorcyclists with Safety RodsMediaTek Launches Dimensity 9600M: Minor Refresh of 9500FBR Admits IRIS Portal Issues Ahead of Tax Filing DeadlineRawalpindi Reports New Dengue Case Amid Intensified Control EffortsPM Directs Facilitation Desks to Assist Citizens with Fuel Relief SchemeTaiwan’s NSTC Pushes for Quantum Breakthroughs in Chip ManufacturingJazzCash Boosts Savings Among BISP BeneficiariesPunjab Forest Department Recovers Rs 9 Billion Worth of LandSBP Prohibits Bank Employees from Applying for Affordable Housing SchemeAudit Reveals Failures in Double-Shift School ProgramGmail Settings to Combat AI Scams and SpamFinance Division Defends Performance of State-Owned EnterprisesGold prices surge by Rs7,900 per tola in PakistanFinance Division refutes SOE financial decline claimsPML-N Nominates Dr Najeeb Naqi for AJK PresidencyCantonment Board Workers Tackle Maintenance Tasks in KarachiJI warns talks with government may end over petroleum levyPakistan Extends Airspace Ban on Indian Aircraft Until October 24Lahore High Court Imposes Rs. 100,000 Fine on Woman for Contempt PetitionLahore to Launch Vehicular Emission Scanning SystemLahore Traffic Police to Equip 2.3 Million Motorcyclists with Safety RodsMediaTek Launches Dimensity 9600M: Minor Refresh of 9500FBR Admits IRIS Portal Issues Ahead of Tax Filing DeadlineRawalpindi Reports New Dengue Case Amid Intensified Control Efforts
◕ SundialUpdated 10 hours ago
Trending
Technology

Xiaomi Releases Open-Source AI to Isolate Voices in Overlapping Speech

Xiaomi has released an open-source AI model, Xiaomi-CocktailASR-1, to accurately transcribe one person’s voice in overlapping speech recordings.

Add Sun News on Google News
Xiaomi Releases Open-Source AI to Isolate Voices in Overlapping Speech
A developer working with an AI model on a computer screen.

Key Takeaways

  • Xiaomi has released Xiaomi-CocktailASR-1, an open-source speech recognition model.
  • The model can identify and transcribe one person’s voice in recordings with multiple speakers.
  • Developers can access the model through GitHub and Hugging Face.

Xiaomi has unveiled a new open-source artificial intelligence (AI) model designed to isolate one person’s voice in recordings where multiple speakers are present. The model, named Xiaomi-CocktailASR-1, addresses the 'cocktail party problem'—a common challenge for automatic speech recognition systems where overlapping voices can confuse the technology.

Developed to enhance the accuracy of speech recognition in various scenarios, Xiaomi-CocktailASR-1 works by first taking a short audio sample of the target speaker. It then uses this sample as a reference to process a recording containing multiple speakers, focusing on identifying and transcribing only the selected person’s speech while ignoring others.

According to Xiaomi, the model has achieved state-of-the-art results across several multi-speaker speech recognition benchmarks and outperforms existing systems designed for similar tasks. It maintains competitive performance even in recordings with only one speaker, ensuring that the multi-speaker capabilities do not significantly reduce standard transcription quality.

AdvertisementFollow Sun NewsWatch Live

The model is equipped with a reasoning mode that provides additional information on how it arrived at a transcription, offering transparency and reliability. If the selected speaker is not present in a recording, Xiaomi-CocktailASR-1 returns an empty result, avoiding the transcription of the wrong person.

Xiaomi has a history of releasing open-source AI models, with previous releases including MiMo-V2-Flash and Xiaomi Robotics-0. Xiaomi-CocktailASR-1 is now available for developers to access and further develop through GitHub and Hugging Face.

This development could have significant implications for various applications, including meetings, interviews, and group conversations, where multiple people speak over one another. The technology could also be beneficial in scenarios requiring accurate transcription of specific speakers, such as legal proceedings or customer service calls.