Xiaomi has launched a new version of its AI models, MiMo V2.5, which allow converting text to voice and vice versa
Xiaomi launches new AI voice models
The company Xiaomi has introduced two versions of its voice AI models that work with text and audio:
Model Feature Key Features MiMo‑V2.5‑TTS Text → speech 3 variants, free for a limited period on MiMo Studio; speed, tone, and emotion adjustment; ability to create new voices from a short phrase (VoiceDesign) and clone a voice from a small set of samples (VoiceClone). MiMo‑V2.5‑ASR Speech → text Speech recognition in challenging conditions, support for Chinese dialects + English; bilingual dialogues and song lyrics (vocals over music); operation under heavy noise; automatic punctuation based on intonation.
How to use MiMo‑V2.5‑TTS
1. Basic mode – selects one of the preset voices and allows changing speed, tone, and emotional shade.
2. VoiceDesign – enters a short phrase, after which the system generates a new voice timbre.
3. VoiceClone – uploads several samples of the chosen voice; the model reproduces it while preserving style and instructions.
To achieve the desired sound, the user can:
- Add special tags to the text;
- Describe the voice in plain natural language (in Chinese or English);
- Create scenarios for virtual productions where multiple voices interact simultaneously.
MiMo‑V2.5‑ASR Features
- Multilingualism – support for several Chinese dialects and English.
- Song transcription – the model isolates vocals even with background music.
- Noise robustness – accurate recognition in conditions of strong external noise.
- Intonation-based punctuation – automatically inserts punctuation marks, reducing the need for manual editing.
Thus, Xiaomi expands its capabilities in voice technology, offering flexible tools for both synthesized speech creation and precise spoken information recognition.
Comments (0)
Share your thoughts — please be polite and stay on topic.
Log in to comment