DeepSeek V4 Flash API in India
DeepSeek V4 Flash in India: ₹9.07 input, ₹18.14 output per million tokens. The fast, low-cost tier for chat and RAG, billed in rupees.
DeepSeek V4 Flash is the speed-and-cost tier of the V4 family — built for high-volume inference rather than deep reasoning. On unoblox it's priced at ₹9.07 per million input tokens and ₹18.14 per million output tokens, billed monthly in rupees.
What Flash is optimized for
Flash trades some reasoning depth for throughput and price. That trade-off makes sense for chatbots, search ranking, classification, and any pipeline running thousands of requests an hour, where V3.2's deeper (and pricier) reasoning would be wasted.
Pricing in context
| Model | Input (₹/1M) | Output (₹/1M) | Fit |
|---|---|---|---|
| DeepSeek V4 Flash | ₹9.07 | ₹18.14 | High-volume chat, RAG, classification |
| DeepSeek V4.1 Flash | ₹20.16 | ₹60.48 | A step up in quality, still fast |
| DeepSeek V3.2 | ₹26.21 | ₹38.30 | Deep reasoning, proofs, research |
V4 Flash costs roughly 45% of V4.1 Flash's input price and about 30% of its output price. As a worked example: 50 million input tokens plus 10 million output tokens in a month costs about ₹635 on V4 Flash versus about ₹1,613 on V4.1 Flash — roughly ₹978 saved by staying on Flash for workloads that don't need the extra quality.
How to switch to Flash
- Get a key → sign up at https://unoblox.ai/sign-in (a couple of minutes).
- Swap the model string →
deepseek-ai/deepseek-v4-flash. - Ship and watch spend → per-request cost in rupees shows up in your dashboard immediately.
const OpenAI = require('openai');
const client = new OpenAI({
apiKey: process.env.UNOBLOX_KEY,
baseURL: 'https://api.unoblox.ai/v1',
});
const msg = await client.chat.completions.create({
model: 'deepseek-ai/deepseek-v4-flash',
messages: [{ role: 'user', content: 'Classify this support ticket...' }],
});
Why India teams pick Flash
- Low cost per token — ₹9.07 per million input tokens is one of the cheaper rates on the catalog.
- India-resident billing — a GST invoice every month, input-tax credit claimable, nothing to reconcile against a foreign statement.
- Scale without fixed fees — pay for what you use; no monthly minimum.
- One key for everything — move to V3.2, GPT, or Claude from the same integration when a request needs more depth.
Frequently asked questions
Is Flash good enough for production chatbots? Yes — it's built for exactly that. The trade-off is lighter reasoning depth, which rarely matters for chat, classification, or search.
Can I A/B test Flash against another model? Easily — changing the model string is a one-line edit, so you can run both in production and compare quality before committing.
Do you offer custom rates at high volume?
Published rates are on /models; teams spending significantly more than average can talk to us about custom terms.
Can I use Flash for embeddings or images?
Flash handles text and multi-turn chat. For vector search, call the native /v1/embeddings endpoint, which is priced separately.
Does a long prompt cost more per token? No. Input tokens are billed at the same per-token rate regardless of prompt length — there's no separate long-context surcharge.
Get started in rupees → https://unoblox.ai/sign-in
Start building in rupees
Call every major model through one OpenAI-compatible endpoint, billed in ₹ on a GST invoice.