Reranker pricing in India (₹)
Qwen3-Reranker-0.6B on unoblox's native /v1/rerank: ₹0 per 1M tokens, billed in rupees with a GST invoice. Cohere-compatible request format.
Native reranking on unoblox: ₹0 per 1M tokens
Reranking sits between retrieval and generation in a RAG pipeline: retrieval pulls a broad set of candidate documents, and a reranker re-scores that shortlist for actual relevance before it reaches the model. Qwen3-Reranker-0.6B runs on unoblox's native /v1/rerank endpoint (Cohere-compatible request format) at ₹0, itemized on your monthly rupee invoice like every other model.
How it fits into a RAG pipeline
POST https://api.unoblox.ai/v1/rerank
model: Qwen/Qwen3-Reranker-0.6B
request:
query: "best restaurants in Mumbai"
documents: [{id: "1", text: "North Indian..."}, ...]
top_n: 5
response:
results: [{index: 0, relevance_score: 0.87, document: {...}}, ...]
You retrieve a wider candidate set cheaply (embeddings or keyword search), then let the reranker narrow it to the handful of documents actually worth sending to the generation model — which is usually your most expensive step per token.
Two-stage retrieval, priced
| Stage | Purpose | Model | Cost |
|---|---|---|---|
| Embed / retrieve | Broad recall (e.g. top 100 candidates) | Qwen3-Embedding-0.6B | ₹0 |
| Rerank | Precision re-scoring (top 5–10) | Qwen3-Reranker-0.6B | ₹0 |
| Generate | Final answer | Any chat model | Varies by model — see /models |
The retrieval and reranking steps cost nothing; only the generation step consumes paid tokens, and it consumes far fewer of them because the reranker has already cut the context down to what's actually relevant.
Why add a reranking step at all
- Better relevance than embeddings alone. Embedding similarity is a fast, approximate signal; a reranker looks at the query and each candidate together, which is a meaningfully more precise (if slower) comparison — worth it for the top-K stage since it's cheap here.
- Smaller, cheaper generation prompts. Sending your generation model 5 well-chosen documents instead of 30 loosely-relevant ones cuts input tokens on the expensive step.
- Data stays on unoblox's infrastructure. Reranking runs on the same platform as the rest of your pipeline — no separate vendor, no separate invoice.
- Production-ready. Same uptime commitment as the rest of unoblox — see
/guides/ai-api-uptime-sla-india.
Quick start
import requests
resp = requests.post(
"https://api.unoblox.ai/v1/rerank",
headers={"Authorization": "Bearer ub-gw-YOUR_KEY"},
json={
"model": "Qwen/Qwen3-Reranker-0.6B",
"query": "refund policy for annual plans",
"documents": [doc["text"] for doc in candidates],
"top_n": 5,
},
)
ranked = resp.json()["results"]
Same ub-gw-… key as your chat completions — no separate account or credentials.
Frequently asked questions
What's the actual difference between embedding and reranking? Embedding encodes documents into vectors once, ahead of time, for fast approximate search across a large corpus. Reranking scores query and document together, at query time, on a small shortlist — slower per pair, but far more precise, which is why it's used only on the narrowed-down candidates.
Does it handle Hindi or other Indian-language queries? It's trained primarily on English and Chinese; Indian-language queries will work but with lower precision than English — worth testing on your own corpus before relying on it in production.
Can I rerank a large batch of documents in one call? Yes, but latency scales with the number of documents. For very large candidate sets, retrieve a smaller shortlist first (50–200 documents) rather than reranking thousands at once.
Is the relevance score a probability or just a ranking signal?
It's a 0–1 relevance score, useful for ordering and for picking a top_n cutoff — treat it as relative between documents in the same request rather than as an absolute confidence threshold.
Can I chain embeddings, reranking, and a chat model in one pipeline? Yes — all three sit behind the same API key with no re-authentication between calls.
Get started in rupees → https://unoblox.ai/sign-in
More from unoblox
API Pricing Guides
Vision LLM API pricing in India
API Pricing Guides
Batch AI API pricing in India
API Pricing Guides
RAG pipeline cost in India
API Pricing Guides
AI agent running cost in India
API Pricing Guides
OpenAI o3 Pricing in India (Live ₹ Rate)
API Pricing Guides
Cheapest Vision LLM in India (₹ Pricing)
Start building in rupees
Call every major model through one OpenAI-compatible endpoint, billed in ₹ on a GST invoice.