Microsoft unveils real-time speech-to-text model with 2.5% word error rate
Microsoft has launched its first real-time streaming speech-to-text model, MAI-Transcribe-2-Streaming, capable of transcribing 60 languages with automatic language detection.
The model, available at a discounted rate of $0.54 per hour (approximately 3.6 RMB), offers a 0.13-second delay and a 2.5% word error rate, outperforming 28 other streaming models.
In tests, it achieved a final transcription delay of 0.13 seconds and a 2.5% word error rate, surpassing competitors like Grok Voice Transcribe 2.0 Streaming and ElevenLabs Scribe v2 Realtime.
The model provides partial text within 100 milliseconds and refines it as more context becomes available, enabling real-time applications such as customer service and live captions. Microsoft also highlighted its non-streaming counterpart, MAI-Transcribe-2, which has a 5.2% average word error rate and is suited for post-recording tasks.