AI API cost calculator (₹)
Plan monthly LLM spend in rupees. Compare input/output tokens across GPT, Claude, Llama, Qwen models. Budget forecasting.
AI API cost calculator — rupee-based budget planning
Building with LLMs? Predict monthly spend accurately by calculating token costs per model, then compare. unoblox's cost calculator shows real ₹ pricing across all major models — no hidden fees, transparent costs.
How token costs work
Every API call consumes two types of tokens:
- Input tokens: Your prompt (typically 200–2000 per request).
- Output tokens: The model's response (typically 100–1000 per request).
Cost = (Input tokens × input rate) + (Output tokens × output rate).
Rates vary by model. GPT-4o is ₹252/M input, ₹1008/M output. Llama Scout is ₹10.08/M input, ₹30.24/M output. Same 1M tokens, 25× price difference.
Step-by-step cost calculation
Scenario: A customer-support chatbot handling 100K daily customer questions.
-
Estimate daily tokens:
- Input per question: 300 tokens (context + instructions + customer message).
- Output per answer: 150 tokens (model response).
- Daily volume: 100K questions × 300 input + 100K × 150 output = 45M daily tokens.
- Monthly: 45M × 30 = 1.35B monthly tokens.
-
Pick a model and calculate cost** (900M input + 450M output — the 2:1 split implied by the 300:150 ratio above):
- Qwen3 1.7B (₹0/M): ₹0 / month ✓
- Llama Scout (₹10.08 input, ₹30.24 output): 900 × ₹10.08 + 450 × ₹30.24 = ₹9,072 + ₹13,608 = ₹22,680 / month.
- Claude Sonnet (₹201.6 input, ₹1,008 output): 900 × ₹201.6 + 450 × ₹1,008 = ₹181,440 + ₹453,600 = ₹635,040 / month.
- GPT-4o (₹252 input, ₹1,008 output): 900 × ₹252 + 450 × ₹1,008 = ₹226,800 + ₹453,600 = ₹680,400 / month.
Takeaway: The same workload costs ₹0 (free tier), ~₹23K (Llama Scout), or ~₹680K (GPT-4o). Model choice is often more important than optimization.
Model pricing reference (per 1M tokens)
| Model | Input (₹) | Output (₹) | ₹/monthly (1B input, 350M output) |
|---|---|---|---|
| Qwen3 1.7B | 0 | 0 | ₹0 |
| Qwen3 235B | 9.07 | 55.44 | ₹28,474 |
| Llama Scout | 10.08 | 30.24 | ₹20,664 |
| Llama Maverick | 20.16 | 80.64 | ₹48,384 |
| Claude Sonnet | 201.6 | 1,008 | ₹554,400 |
| Claude Opus | 504 | 2,520 | ₹1,386,000 |
| GPT-4o | 252 | 1,008 | ₹604,800 |
| GPT-4.1 | 201.6 | 806.4 | ₹483,840 |
| GPT-5 | 126 | 1,008 | ₹478,800 |
Monthly budget by use case
| Use case | Est. monthly tokens | Recommended model | ₹/month |
|---|---|---|---|
| Chatbot (1K daily users) | 50M | Llama Scout | ₹840 |
| RAG pipeline (100K docs) | 2B | Qwen3 235B | ₹49,053 |
| Batch summarization | 500M | Qwen3 1.7B | ₹0 |
| Production SaaS | 5B | Llama Maverick | ₹201,600 |
| Enterprise reasoning | 1B | Claude Sonnet | ₹470,400 |
Assumes a 2:1 input:output token mix, the same ratio used throughout this page.
How to reduce costs
-
Optimize prompts: Shorter, clearer instructions = fewer input tokens.
- Before: "You are a customer support bot. Please respond to the following query…" (25 tokens).
- After: "Support bot response:" (3 tokens). 8× token savings.
-
Use cheaper models first: Start on Qwen3 1.7B or Llama Scout. Upgrade only if output quality is insufficient.
-
Batch requests: Process 1000 questions at once instead of 1000 individual API calls — same token cost, lower latency overhead.
-
Cache context: If your app repeats the same 10K-token context (e.g., a company handbook), cache it and reuse.
-
Streaming: Only pay for tokens you generate. Stop early if the output is complete.
Using the unoblox calculator
Visit https://unoblox.ai/pricing to:
- Enter your estimated monthly tokens.
- Select 2–3 models to compare.
- See total ₹/month in real-time.
- Export a CSV for your finance team.
Frequently asked questions
How do I know if my token estimate is accurate?
Call https://api.unoblox.ai/v1/tokenize (demo) or enable usage logging in your SDK. Most apps underestimate tokens by 20–30%.
Do I get charged per API call or per token? Per token. 10 calls with 1M tokens each = same cost as 1 call with 10M tokens.
Can I set a monthly spend cap?
Yes. unoblox allows max_spend_monthly: ₹5000 per API key. Requests over the limit return 429 Limit Exceeded.
What about caching or context reuse? Claude supports prompt caching (25% cheaper). GPT has similar features. Check your model's docs.
Should I optimize for input or output tokens? Input is typically 3–5× more output. Optimize input first (prompt engineering). Output optimization happens naturally (better models = shorter answers).
How often does pricing change? unoblox posts 30-day notice for price changes. Current rates valid through Q4 2026.
Get started in rupees → https://unoblox.ai/sign-in
More from unoblox
API Pricing Guides
Vision LLM API pricing in India
API Pricing Guides
Batch AI API pricing in India
API Pricing Guides
RAG pipeline cost in India
API Pricing Guides
AI agent running cost in India
API Pricing Guides
OpenAI o3 Pricing in India (Live ₹ Rate)
API Pricing Guides
Cheapest Vision LLM in India (₹ Pricing)
Start building in rupees
Call every major model through one OpenAI-compatible endpoint, billed in ₹ on a GST invoice.