Technology

Sarvam AI launches Saaras V4 speech model supporting 22 Indian languages

An illustration representing artificial intelligence. [mikemacmarketing/Wikimedia Commons]

Indian artificial intelligence startup Sarvam AI has announced the launch of Saaras V4, a speech recognition model supporting 22 Indian languages and English, as the company broadens its scope to multilingual voice applications Business Standard reported.

The model has five output formats — transcription, translation, transliteration, verbatim and code-mixed text. Hence, the same audio can be processed in different formats, Sarvam AI said in its product announcement.

Saaras V4 has been trained to manage noisy audio, regional accents, dialectal variations and code-mixed speech. It also supports low-latency streaming, with Sarvam reporting a time-to-first-token of less than 150 milliseconds for real-time applications.

The model utilizes an in-house trained 3-billion-parameters hybrid state-space language model by Sarvam and an audio encoder. The company said it achieved its lowest average word error rate across seven English speech benchmarks and posted strong results in Indian language tests.

Saaras V4 can automatically identify languages, Sarvam said and the language-identification error rate was 5.22% across all 22 Indian languages, and 2.9% across the 10 most widely spoken languages, in testing.

The model can be accessed via Sarvam’s speech-to-text APIs for use cases like voice agents and call analytics.

Click to comment
To Top