Skip to content
Start freeGPT-5 · Claude · DeepSeek V4 · Qwen3 — in ₹One OpenAI-compatible endpointBilled in rupeesGST invoiceNo international cardGet started →Start freeGPT-5 · Claude · DeepSeek V4 · Qwen3 — in ₹One OpenAI-compatible endpointBilled in rupeesGST invoiceNo international cardGet started →
API Pricing Guides

LLM token cost in India (₹)

Input & output token pricing for every model on unoblox, in rupees. Real worked examples, a full rate table, and how to estimate your own monthly cost.

LLM token pricing in India — transparent ₹ rates, no hidden costs

Every LLM charges per token. unoblox prices all models in rupees, so you can budget in ₹ directly instead of estimating from a foreign price sheet. Below: how tokens are counted, the full rate table, and three worked monthly-cost examples with the arithmetic shown.

Token 101

A token is roughly 4 characters of English text. Your prompt plus the model's response together make up the billable tokens for a request.

  • "Hello" ≈ 1 token.
  • "What is machine learning?" ≈ 5 tokens.
  • A 1,000-word blog post ≈ 1,300–1,400 tokens.

Cost per request = (input tokens × input rate per 1M) + (output tokens × output rate per 1M).

All rates at a glance (₹ per 1M tokens)

ModelInputOutputNotes
Qwen3 1.7B₹0₹0Free tier
Qwen3 235B-A22B₹9.07₹55.44Budget production
Qwen3 Max₹120.95₹604.77Reasoning, vision
Llama 4 Scout₹10.08₹30.24Cheapest quality option
Llama 4 Maverick₹20.16₹80.64Reasoning, long-context
DeepSeek V4 Flash₹9.07₹18.14Ultra-cheap
Claude Sonnet₹201.6₹1,008Production, balanced
Claude Opus₹504₹2,520Deep reasoning
GPT-5 mini₹25.2₹201.6Cheap, capable
GPT-5₹126₹1,008Long-context reasoning
GPT-4o₹252₹1,008Multimodal

Rates not listed here are indicative and change occasionally — check the live rate on each model's page at /models before budgeting a large commitment.

Worked example 1: chatbot, 10K daily users, 5 messages/day

Each message runs ~200 input + 100 output tokens. That's 50,000 messages/day × 300 tokens = 15M tokens/day, or 300M input + 150M output tokens/month.

ModelMonthly cost
Qwen3 235B₹11,037
Llama Scout₹7,560
Claude Sonnet₹2,11,680
GPT-4o₹2,26,800

Worked example 2: RAG search, 100K searches/month

Each search sends ~500 tokens of context and returns ~200 tokens — 50M input + 20M output tokens/month.

ModelMonthly cost
Qwen3 235B₹1,562
Llama Scout₹1,109
Claude Sonnet₹30,240
GPT-4o₹32,760

Worked example 3: batch summarization, 10M documents/month

Each summary runs ~1,000 input + 300 output tokens — 10B input + 3B output tokens/month.

ModelMonthly cost
Qwen3 235B₹2,57,020
Llama Scout₹1,91,520
Claude Sonnet₹50,40,000
GPT-4o₹55,44,000

At batch scale, model choice is the single biggest lever on your bill — the gap between Llama Scout and Claude Sonnet alone is over ₹48 lakh a month at this volume.

Input vs output cost

Output tokens are usually priced 3–5× higher than input tokens across every provider. Claude Sonnet, for example, charges ₹201.6/M input vs ₹1,008/M output — exactly 5×. Practical implications:

  • Tightening your prompt saves less than you'd think, since input is the cheaper side.
  • Capping or truncating output (via max_tokens, or stopping a stream early once you have what you need) usually saves more per request than prompt-shortening does.

How to estimate your own monthly cost

  1. Count average tokens per request (prompt + expected response length).
  2. Estimate daily request volume for the feature.
  3. Multiply: tokens/request × requests/day × 30 = monthly tokens.
  4. Split into input and output using your typical ratio.
  5. Apply the rate table above: (input tokens × input rate) + (output tokens × output rate).

Keeping a hard ceiling on spend

If you want a spending guardrail rather than just an estimate, set max_spend_monthly on the API key itself — requests past the cap return a clear error instead of silently continuing to bill.

Frequently asked questions

How are tokens counted exactly? Each model family uses its own tokenizer, so the same text can produce a slightly different token count across GPT, Claude, Llama, and Qwen — typically within a few percent of each other.

Can I see token usage per request? Yes. Every API response includes usage.prompt_tokens and usage.completion_tokens.

Do I pay for failed requests? No. 4xx/5xx errors are never billed — only completed responses with real output tokens.

What if a model's rate changes? Rate changes are announced with notice; existing invoiced usage is never retroactively repriced.

Can I cap my monthly spend per key? Yes, via max_spend_monthly in the key configuration — requests over the limit return an error instead of billing further.

Where do I see my usage history? The dashboard at /user/usage breaks down daily and monthly usage by model, token count, and cost.

Get started in rupees → https://unoblox.ai/sign-in

llm token cost indiacost
ShareLinkedInWhatsAppTelegram

Start building in rupees

Call every major model through one OpenAI-compatible endpoint, billed in ₹ on a GST invoice.