Skip to content
Start freeGPT-5 · Claude · DeepSeek V4 · Qwen3 — in ₹One OpenAI-compatible endpointBilled in rupeesGST invoiceNo international cardGet started →Start freeGPT-5 · Claude · DeepSeek V4 · Qwen3 — in ₹One OpenAI-compatible endpointBilled in rupeesGST invoiceNo international cardGet started →
API Pricing Guides

AI Cost Per 1M Tokens in India (₹ Guide)

Compare real ₹ cost per 1 million tokens across GPT, Claude, DeepSeek, Qwen, Llama and Gemma on unoblox, billed monthly with GST invoicing.

"Cost per 1 million tokens" is the standard way AI providers price inference, and it's the number every India-based team ultimately needs in rupees to budget properly. unoblox publishes that number directly in ₹ for every priced model — no international card, just a per-token rupee rate and one monthly GST invoice.

₹ cost per 1M tokens, model by model

ModelInput (₹ / 1M tokens)Output (₹ / 1M tokens)
DeepSeek V4 Flash₹9.07₹18.14
Qwen3 235B-A22B₹9.07₹55.44
Llama 4 Scout₹10.08₹30.24
Gemma 4 31B₹13.10₹38.30
Qwen3.8-27B₹16.32₹48.96
DeepSeek V4.1 Flash₹20.16₹60.48
Llama 4 Maverick₹20.16₹80.64
GPT-5 mini₹25.2₹201.6
DeepSeek V3.2₹26.21₹38.30
Kimi K2.7₹68.5₹342.7
Qwen3 Max₹120.95₹604.77
GPT-5₹126₹1,008
Claude Sonnet₹201.6₹1,008
GPT-4.1₹201.6₹806.4
GPT-4o₹252₹1,008
Claude Opus₹504₹2,520
Qwen3 1.7B₹0 (free)₹0 (free)

This covers models with a published indicative rate today. Anything not listed — o3, GPT-5 nano, GPT-4o mini, Claude Haiku, Nemotron, and a few others — has a live rate on its own catalog page instead; see /models for the complete, current list.

How to actually estimate your monthly bill

The formula is the same for every model: (input tokens ÷ 1,000,000 × input rate) + (output tokens ÷ 1,000,000 × output rate). For example, a workload sending 500,000 input tokens and 150,000 output tokens a month through DeepSeek V4 Flash costs: (0.5 × ₹9.07) + (0.15 × ₹18.14) = ₹4.54 + ₹2.72 ≈ ₹7.26 for the whole month. Run the same workload through GPT-4o instead: (0.5 × ₹252) + (0.15 × ₹1,008) = ₹126 + ₹151.2 ≈ ₹277.2 — same traffic, a very different bill, purely from model choice.

Why input and output are priced differently

Generating tokens costs more compute than reading them, so output is priced higher than input across the catalog — often 2 to 5 times higher. A short-answer workload like classification is cheaper per call than a long-answer workload like drafting, even on the same model.

Calling any model at its published rate

Every model shares the same request shape — only the model field changes.

base_url: https://api.unoblox.ai/v1
api_key: ub-gw-xxxxxxxxxxxxxxxxxxxx
model: deepseek-ai/deepseek-v4-flash

Cheapest vs most capable — how to choose

Don't optimise on price alone. A cheaper model needing three retries can cost more in engineering effort than a pricier model that works first time. Prototype on a low-cost or free model like Qwen3 1.7B, measure quality against your task, and move up the price table only as far as accuracy actually demands.

Frequently asked questions

What does "per 1 million tokens" actually mean for a small project? Most real requests use a few hundred to a few thousand tokens, so 1 million is a large batch — hundreds or thousands of typical calls, not one purchase.

Are these rates final, or can they change? They're indicative. The live rate for each model is always shown on its own page at /models/<model-id> — treat that as the source of truth if this page and the model page ever disagree.

Does unoblox charge anything besides the per-token rate? No separate platform fee is added on top of the published per-token rate; usage across all models rolls into one monthly GST invoice.

Is the free model (Qwen3 1.7B) good enough for production traffic? It depends on your task. It's a genuinely usable model for prototyping, low-stakes traffic, and testing your integration before you commit budget to a paid model.

Why do some models on the catalog have no listed price here? We only print numbers we can stand behind at the time of writing. Models without a fixed rate on this page have a live rate on their own model page instead.

Can I mix cheap and expensive models in the same application? Yes — that's a common pattern. Route routine requests to a low-cost model and escalate only the harder cases to a premium one, all under the same key and invoice.

Get started in rupees → https://unoblox.ai/sign-in

cost per million tokens indiacostpricingtokens
ShareLinkedInWhatsAppTelegram

Start building in rupees

Call every major model through one OpenAI-compatible endpoint, billed in ₹ on a GST invoice.