LLM token cost in India (₹)
Input & output token pricing for every model on unoblox, in rupees. Real worked examples, a full rate table, and how to estimate your own monthly cost.
LLM token pricing in India — transparent ₹ rates, no hidden costs
Every LLM charges per token. unoblox prices all models in rupees, so you can budget in ₹ directly instead of estimating from a foreign price sheet. Below: how tokens are counted, the full rate table, and three worked monthly-cost examples with the arithmetic shown.
Token 101
A token is roughly 4 characters of English text. Your prompt plus the model's response together make up the billable tokens for a request.
- "Hello" ≈ 1 token.
- "What is machine learning?" ≈ 5 tokens.
- A 1,000-word blog post ≈ 1,300–1,400 tokens.
Cost per request = (input tokens × input rate per 1M) + (output tokens × output rate per 1M).
All rates at a glance (₹ per 1M tokens)
| Model | Input | Output | Notes |
|---|---|---|---|
| Qwen3 1.7B | ₹0 | ₹0 | Free tier |
| Qwen3 235B-A22B | ₹9.07 | ₹55.44 | Budget production |
| Qwen3 Max | ₹120.95 | ₹604.77 | Reasoning, vision |
| Llama 4 Scout | ₹10.08 | ₹30.24 | Cheapest quality option |
| Llama 4 Maverick | ₹20.16 | ₹80.64 | Reasoning, long-context |
| DeepSeek V4 Flash | ₹9.07 | ₹18.14 | Ultra-cheap |
| Claude Sonnet | ₹201.6 | ₹1,008 | Production, balanced |
| Claude Opus | ₹504 | ₹2,520 | Deep reasoning |
| GPT-5 mini | ₹25.2 | ₹201.6 | Cheap, capable |
| GPT-5 | ₹126 | ₹1,008 | Long-context reasoning |
| GPT-4o | ₹252 | ₹1,008 | Multimodal |
Rates not listed here are indicative and change occasionally — check the live rate on each model's page at /models before budgeting a large commitment.
Worked example 1: chatbot, 10K daily users, 5 messages/day
Each message runs ~200 input + 100 output tokens. That's 50,000 messages/day × 300 tokens = 15M tokens/day, or 300M input + 150M output tokens/month.
| Model | Monthly cost |
|---|---|
| Qwen3 235B | ₹11,037 |
| Llama Scout | ₹7,560 |
| Claude Sonnet | ₹2,11,680 |
| GPT-4o | ₹2,26,800 |
Worked example 2: RAG search, 100K searches/month
Each search sends ~500 tokens of context and returns ~200 tokens — 50M input + 20M output tokens/month.
| Model | Monthly cost |
|---|---|
| Qwen3 235B | ₹1,562 |
| Llama Scout | ₹1,109 |
| Claude Sonnet | ₹30,240 |
| GPT-4o | ₹32,760 |
Worked example 3: batch summarization, 10M documents/month
Each summary runs ~1,000 input + 300 output tokens — 10B input + 3B output tokens/month.
| Model | Monthly cost |
|---|---|
| Qwen3 235B | ₹2,57,020 |
| Llama Scout | ₹1,91,520 |
| Claude Sonnet | ₹50,40,000 |
| GPT-4o | ₹55,44,000 |
At batch scale, model choice is the single biggest lever on your bill — the gap between Llama Scout and Claude Sonnet alone is over ₹48 lakh a month at this volume.
Input vs output cost
Output tokens are usually priced 3–5× higher than input tokens across every provider. Claude Sonnet, for example, charges ₹201.6/M input vs ₹1,008/M output — exactly 5×. Practical implications:
- Tightening your prompt saves less than you'd think, since input is the cheaper side.
- Capping or truncating output (via
max_tokens, or stopping a stream early once you have what you need) usually saves more per request than prompt-shortening does.
How to estimate your own monthly cost
- Count average tokens per request (prompt + expected response length).
- Estimate daily request volume for the feature.
- Multiply: tokens/request × requests/day × 30 = monthly tokens.
- Split into input and output using your typical ratio.
- Apply the rate table above: (input tokens × input rate) + (output tokens × output rate).
Keeping a hard ceiling on spend
If you want a spending guardrail rather than just an estimate, set max_spend_monthly on the API key itself — requests past the cap return a clear error instead of silently continuing to bill.
Frequently asked questions
How are tokens counted exactly? Each model family uses its own tokenizer, so the same text can produce a slightly different token count across GPT, Claude, Llama, and Qwen — typically within a few percent of each other.
Can I see token usage per request?
Yes. Every API response includes usage.prompt_tokens and usage.completion_tokens.
Do I pay for failed requests? No. 4xx/5xx errors are never billed — only completed responses with real output tokens.
What if a model's rate changes? Rate changes are announced with notice; existing invoiced usage is never retroactively repriced.
Can I cap my monthly spend per key?
Yes, via max_spend_monthly in the key configuration — requests over the limit return an error instead of billing further.
Where do I see my usage history?
The dashboard at /user/usage breaks down daily and monthly usage by model, token count, and cost.
Get started in rupees → https://unoblox.ai/sign-in
More from unoblox
API Pricing Guides
Vision LLM API pricing in India
API Pricing Guides
Batch AI API pricing in India
API Pricing Guides
RAG pipeline cost in India
API Pricing Guides
AI agent running cost in India
API Pricing Guides
OpenAI o3 Pricing in India (Live ₹ Rate)
API Pricing Guides
Cheapest Vision LLM in India (₹ Pricing)
Start building in rupees
Call every major model through one OpenAI-compatible endpoint, billed in ₹ on a GST invoice.