Streaming AI API Responses in India | unoblox
How to stream token-by-token AI API responses in India on unoblox's rupee-billed OpenAI-compatible endpoint, with model picks and FAQs.
Streaming turns one AI API call into a live feed of tokens instead of a single long wait — the difference between a reply that types itself out and one that freezes for several seconds before appearing. unoblox exposes streaming exactly the way the OpenAI SDK already expects, on one rupee-billed endpoint (https://api.unoblox.ai/v1), so teams building in India can ship real-time chat, coding copilots, and voice assistants without adopting a new client library or holding an international card.
Why streaming matters for user-facing apps
Anything a person watches while it happens — a support chatbot, a code-completion panel, a live translation box — feels faster and more trustworthy when the first words show up immediately. Waiting for a full response before showing anything makes even a quick model feel sluggish. Streaming does not change what the model computes, only how the output is delivered: tokens arrive as they are generated instead of all at once at the end.
How streaming works on unoblox
Set stream: true in your chat completion request and read the response as server-sent events, exactly like the OpenAI streaming format. No separate WebSocket, no proprietary event schema, and no change to how you authenticate — the same ub-gw-... key and the same base URL handle both streaming and non-streaming calls.
curl https://api.unoblox.ai/v1/chat/completions \
-H "Authorization: Bearer ub-gw-your-key-here" \
-H "Content-Type: application/json" \
-d '{"model":"deepseek-ai/deepseek-v4-flash","messages":[{"role":"user","content":"Summarise this in one line"}],"stream":true}'
Choosing a model for low-latency streaming
Any model in the catalog can stream, but the "feel" of a live response depends on picking a model sized for the job. A few starting points, priced per 1M tokens (input / output):
| Model | Input | Output | Good for |
|---|---|---|---|
| DeepSeek V4 Flash | ₹9.07 | ₹18.14 | fastest, cheapest for everyday chat streams |
| Qwen3 235B-A22B | ₹9.07 | ₹55.44 | heavier reasoning at flash-tier input cost |
| GPT-5 mini | ₹25.2 | ₹201.6 | general-purpose assistant streaming |
| Claude Sonnet | ₹201.6 | ₹1008 | higher-quality writing and coding streams |
Models outside this shortlist — including GPT-4o mini, o3, and o4-mini — all stream too; check /models for their live ₹ rates.
Billing on a token-by-token stream
Streaming does not carry a separate fee. Usage is metered the same way as a non-streaming call, by input and output tokens, and rolled into one consolidated ₹ invoice with GST each month. Turning stream on is a display choice for your client, not a pricing tier.
Frequently asked questions
Does streaming cost extra on unoblox?
No. Input and output tokens are metered identically whether or not stream is set, and everything lands on the same monthly ₹ GST invoice.
Do I need a different SDK to stream responses?
No — point your existing OpenAI SDK, or any OpenAI-compatible client, at https://api.unoblox.ai/v1 and set stream: true. The event format is unchanged.
Which models support streaming on unoblox?
Streaming works across the catalog, including the GPT, Claude, DeepSeek, Qwen, Llama, Gemma, Kimi, and Mistral families. See /models for the full list and current ₹ pricing.
Can I stream tool calls and JSON-mode responses too? Yes. Tool calls and structured outputs both arrive incrementally as chunks; you assemble the final call or object once the stream completes.
Does streaming reduce total generation time? Not meaningfully — it changes when you see output, not how long the model takes overall. The benefit is that the first tokens appear immediately instead of after the full reply is ready.
Is streaming a good fit for voice or real-time agents built for Indian users? Yes. Lower perceived latency is exactly what real-time chat and voice-assist flows need — pairing it with a fast, inexpensive model like DeepSeek V4 Flash gives the snappiest feel.
Get started in rupees → https://unoblox.ai/sign-in
More from unoblox
Developer Guides
AI API Rate Limits Explained (India Guide)
Developer Guides
How to Get an AI API Key in India
Developer Guides
Migrate from Azure OpenAI to unoblox (India)
Developer Guides
JSON Mode with the AI API in India | unoblox
Developer Guides
Tool Calling with the AI API in India | unoblox
Developer Guides
AI API for Agencies in India | unoblox
Start building in rupees
Call every major model through one OpenAI-compatible endpoint, billed in ₹ on a GST invoice.