Sarvam AI unveils saaras v4: a multilingual speech-to-text model for indian languages and global english
Sarvam AI has launched Saaras V4, a speech-to-text model that supports all 22 scheduled Indian languages along with global English accents. The model is available via Sarvam’s API using the model identifier 'saaras:v4'.
While the model’s weights are not publicly accessible, the company’s SageMaker self-hosting documentation currently supports Saaras v3. The new version employs an encoder-decoder architecture, with a 3B-parameter hybrid state-space language model, Sarvam-3B, as its decoder.
The model’s performance, including a 16.03% WER on the IndicContextEval benchmark, is reported by Sarvam as the lowest on record. Keyterm prompting, a new feature in V4, allows users to bias recognition with up to 50 terms. Sarv, the company, emphasizes that all metrics are vendor-reported, with no independent validation yet published.
Switching from Saaras v3 to V4 requires only a one-line change in the request format.