Quick facts
- Best for
- Speech to Text, perfected for India
- Pricing
- Freemium
- Editor rating
- 4.5 / 5
- Community saves
- 0
About Sarvam AI Speech to Text API
Sarvam AI's Speech to Text API offers seamless and accurate transcription of speech to text in 22 Indian languages, including but not limited to Hindi, Bengali, Tamil, Telugu, Gujarati, Kannada, Malayalam, Marathi, Punjabi, Odia, and English with an Indian accent. It is powered by Sarvam AI's Saarika v2 model and supports speaker diarization, which is useful for identifying and labeling different speakers in audio. This feature is suitable, for instance, in meeting transcriptions, interviews, and call center analytics. Moreover, the tool also supports seamless code-mixing, efficiently handling a switch between Hindi, English, and regional languages mid-sentence. It also supports a plethora of audio formats including MP3, WAV, AAC, OGG, Opus, FLAC, M4A, AMR, WMA, and WebM. Capable of handling different acoustic conditions, Sarvam AI's API maintains its accuracy even in the presence of background noise, cross-talk,and poor connections. To accommodate diverse needs, it offers real-time and batch processing capabilities with three types of APIs - REST API for synchronous processing of files under 30 seconds, Batch API for files up to 1 hour with speaker diarization and timestamps, and Streaming API for real-time transcription via WebSocket. The API is designed to be developer-friendly with easy-to-integrate and scalable for enterprise applications. Supported featuresTranscription
Pros
- Supports 22 Indian languages
- Handles Indian English accent
- Speaker diarization feature
- Efficient code-mixing
- Supports multiple audio formats
- Handles noisy acoustic conditions
- Real-time and batch processing
- REST, Batch, Streaming APIs
- Developer friendly
- Scalable for enterprise applications
- Ideal for call center analytics
- Handy for meeting transcriptions
Cons
- Limited to Indian languages
- Only supports 22 languages
- No support for low-quality audios
- Processing time for long audios
- Lacks accuracy for regional dialects
- Transcriptions not context-aware
- Doesn't support large audio files
- REST API only for short audios
- Streaming API limited to Web
- Sockets
- No pretrained model customization
