Voice Assistant API in India
Build voice-activated AI assistants with speech-to-text + LLM. One endpoint, ₹ billing, no proprietary voice APIs.
Build Voice Assistants with Unoblox
Create voice bots for phone calls, IVR, voice commands—combine speech-to-text (STT) with unoblox LLM endpoint, then text-to-speech (TTS) for response. Billed entirely in rupees.
Voice assistant pipeline
- Listen: Turn speech into text (use Deepgram, AssemblyAI, or open-source Vosk).
- Understand: Send text to unoblox LLM endpoint.
- Respond: Turn the reply text into speech (Google TTS, ElevenLabs, or OSS).
- Act: Execute command (booking, transfer, payment).
Models for voice understanding
| Model | Input/Output (₹) | Latency | Use case |
|---|---|---|---|
| Qwen3 Max | ₹120.95 / ₹604.77 | 300–800ms | Complex queries |
| Claude Sonnet | ₹201.6 / ₹1008 | 200–600ms | Nuanced intent |
| Qwen3.8-27B | ₹16.32 / ₹48.96 | 100–400ms | Fast responses |
| Claude Opus | ₹504 / ₹2520 | 400–900ms | Premium reasoning |
Build a voice assistant in 3 steps
Step 1: Capture voice and turn it into text
Use a speech-to-text service (Deepgram API call):
POST https://api.deepgram.com/v1/listen?model=nova-2&language=hi-IN
Audio file → "Mere liye ek taxi book kar do"
Step 2: Send text to unoblox LLM
POST https://api.unoblox.ai/v1/messages
Authorization: Bearer ub-gw-...
{
"model": "qwen/qwen3-max",
"messages": [{"role": "user", "content": "User request (Hindi): Mere liye ek taxi book kar do. Respond in Hindi, briefly."}],
"max_tokens": 100
}
Response: "Main aapke liye ek taxi book kar deta hoon. Aapka current location kya hai?"
Step 3: Turn the response into speech + play
Use Google Cloud Text-to-Speech or ElevenLabs:
Text: "Main aapke liye ek taxi book kar deta hoon..."
Output: Audio MP3 → Play on phone/speaker
Real scenario: Hindi voice assistant for taxi booking
- Customer calls bot; system plays: "Namaste, unoblox taxi ke liye kya chahiye?"
- Customer says (Hindi): "Sector 10 se sector 20 ek taxi de do."
- Speech-to-text: Turns the audio into text.
- LLM call: Extract pickup, dropoff, time.
- Execute: Book via external API.
- Respond: "Aapki taxi 5 min mein aayegi. Booking ID: XYZ."
- Cost: ~100 tokens/call ≈ ₹20–50 on Qwen3 Max.
Multilingual voice support
Language support:
- English: All models (excellent).
- Hindi: Qwen3 Max, Claude Sonnet (good).
- Tamil, Telugu, Kannada: Test on your use case; may need fallback.
Approach:
- Auto-detect language from STT output.
- Prompt LLM to respond in same language.
- Use same language for TTS.
Intent extraction from voice
Ask the LLM to extract structured intent:
Given: "Book me a taxi from Sector 10 to Sector 20, arriving by 3 PM."
Extract JSON: {"intent": "taxi_booking", "pickup": "Sector 10", "dropoff": "Sector 20", "time": "3 PM"}
Frequently asked questions
Q: How do I handle speech recognition errors? A: Use confirmation loops. "I heard: Sector 10 to Sector 20. Is that correct?" Retry if user says no.
Q: Can the voice assistant handle noisy backgrounds? A: STT services (Deepgram, AssemblyAI) are noise-tolerant. Test in your environment; consider noise filtering if needed.
Q: What's the end-to-end latency (speech to speech)? A: ~2–5 seconds (STT ~1s, LLM ~0.5s, TTS ~1s). Acceptable for most use cases; queue requests if volume exceeds capacity.
Q: Can I build this for IVR (phone tree)? A: Yes. Connect to Twilio, Vonage, or open-source FreeSWITCH. Use their webhooks to send audio, receive text, call unoblox, return TTS.
Q: How do I handle sensitive commands (payment, transfer)? A: Use authentication (OTP, voice biometrics) before processing. Always confirm critical actions. Log all transactions.
Q: Does this support Hindi, Marathi, and regional languages? A: Claude Sonnet and Qwen3 Max handle Hindi well. Regional languages are underserved by current models. Recommend testing before production.
Get started in rupees → https://unoblox.ai/sign-in
Start building in rupees
Call every major model through one OpenAI-compatible endpoint, billed in ₹ on a GST invoice.