RAG pipeline cost in India
₹ cost math for a RAG pipeline serving India traffic: retrieval, generation tokens, and a worked per-model table using live unoblox pricing.
Retrieval-augmented generation has two metered legs — the retrieval step (embeddings and, optionally, reranking) and the generation step (an LLM call stuffed with retrieved chunks) — and in most RAG pipelines the generation leg dominates the bill, because every retrieved chunk becomes input tokens on every single query. Since unoblox prices generation by the token in rupees, that leg is straightforward to cost out once you fix an assumption for how much context you're retrieving.
The assumption this example uses
Assume 2,000 queries a day, each sending about 3,000 input tokens (retrieved chunks, the question, and a system prompt) and generating about 600 output tokens for the answer — a roughly 5:1 input:output ratio, typical of RAG where retrieved context, not the answer, dominates the prompt. That's 6.0M input and 1.2M output tokens a day; adjust the ratio to match your own chunk size and top-k before trusting a number for budgeting.
Generation cost by model
| Model id | ₹ / 1M input | ₹ / 1M output | ₹ / day at this volume |
|---|---|---|---|
deepseek-ai/deepseek-v4-flash | ₹9.07 | ₹18.14 | ₹76.19 |
qwen/qwen3-235b-a22b-instruct-2507 | ₹9.07 | ₹55.44 | ₹120.95 |
google/gemma-4-31b-it | ₹13.10 | ₹38.30 | ₹124.56 |
openai/gpt-4.1 | ₹201.60 | ₹806.40 | ₹2,177.28 |
openai/gpt-5-mini | ₹25.20 | ₹201.60 | ₹393.12 |
At this volume, the cheapest row above runs about ₹2,285.64/month, versus roughly ₹65,318.40/month for GPT-4.1 — the gap is almost entirely the input-token rate, since RAG prompts are input-heavy by construction. Claude Sonnet (anthropic/claude-sonnet-4-5/-4-6/-5) is priced at ₹201.60 / ₹1,008.00 per 1M tokens if you need it for the generation step. Anything not in this table — Claude Opus, o-series, Mistral Small and more — should be priced from its live /models page, not estimated from a neighbouring row.
The retrieval leg
unoblox also serves native /v1/embeddings and /v1/rerank endpoints for the retrieval side of the pipeline, metered separately from chat generation. Because that pricing isn't in the fixed ₹ list above, check the live rate for the specific embedding or rerank entry on /models before adding it to a total — don't assume it scales the same way generation does.
Where the money actually goes
Because input tokens scale with how much you retrieve, the two levers that move a RAG bill the most are top-k (how many chunks you retrieve per query) and chunk size — both are retrieval-pipeline decisions, not something unoblox's pricing changes. A smaller, well-ranked context window costs less on the generation leg regardless of which model answers it.
A minimal request
curl https://api.unoblox.ai/v1/chat/completions \
-H "Authorization: Bearer ub-gw-xxxxxxxxxxxxxxxxxxxxxxxxxxxx" \
-H "Content-Type: application/json" \
-d '{"model":"deepseek-ai/deepseek-v4-flash","messages":[{"role":"user","content":"Answer using only the provided context."}]}'
Frequently asked questions
Does the ₹ price list above include embeddings? No — it covers chat/generation models only. Embedding and rerank pricing lives on each endpoint's own /models page.
What input:output ratio should I actually use? Whatever matches your chunk size and top-k in practice — the 5:1 figure here is a worked example, not a fixed rule for every RAG pipeline.
Does a bigger context window always cost more? Yes — more retrieved tokens means more input tokens billed on every query, regardless of which model answers.
Is there a free model to prototype the generation step? Yes — qwen/qwen3-1.7b is listed at ₹0 while you're testing retrieval quality before paying for generation.
Does reranking add meaningfully to the bill? It adds a separate, smaller metered step; check its live rate on /models rather than assuming a fixed proportion of the generation cost.
How do I get an exact number for my own pipeline? Multiply your actual average input and output tokens per query by the chosen model's listed ₹ rate, then by your daily query volume — the table above shows the same arithmetic at one assumed volume.
Get started in rupees → https://unoblox.ai/sign-in
More from unoblox
API Pricing Guides
Vision LLM API pricing in India
API Pricing Guides
Batch AI API pricing in India
API Pricing Guides
AI agent running cost in India
API Pricing Guides
OpenAI o3 Pricing in India (Live ₹ Rate)
API Pricing Guides
Cheapest Vision LLM in India (₹ Pricing)
API Pricing Guides
Mistral Pricing in India | Live ₹ Rates
Start building in rupees
Call every major model through one OpenAI-compatible endpoint, billed in ₹ on a GST invoice.