Skip to content
Start freeGPT-5 · Claude · DeepSeek V4 · Qwen3 — in ₹One OpenAI-compatible endpointBilled in rupeesGST invoiceNo international cardGet started →Start freeGPT-5 · Claude · DeepSeek V4 · Qwen3 — in ₹One OpenAI-compatible endpointBilled in rupeesGST invoiceNo international cardGet started →
API Pricing Guides

Reranker pricing in India (₹)

Qwen3-Reranker-0.6B on unoblox's native /v1/rerank: ₹0 per 1M tokens, billed in rupees with a GST invoice. Cohere-compatible request format.

Native reranking on unoblox: ₹0 per 1M tokens

Reranking sits between retrieval and generation in a RAG pipeline: retrieval pulls a broad set of candidate documents, and a reranker re-scores that shortlist for actual relevance before it reaches the model. Qwen3-Reranker-0.6B runs on unoblox's native /v1/rerank endpoint (Cohere-compatible request format) at ₹0, itemized on your monthly rupee invoice like every other model.

How it fits into a RAG pipeline

POST https://api.unoblox.ai/v1/rerank
model: Qwen/Qwen3-Reranker-0.6B

request:
  query: "best restaurants in Mumbai"
  documents: [{id: "1", text: "North Indian..."}, ...]
  top_n: 5

response:
  results: [{index: 0, relevance_score: 0.87, document: {...}}, ...]

You retrieve a wider candidate set cheaply (embeddings or keyword search), then let the reranker narrow it to the handful of documents actually worth sending to the generation model — which is usually your most expensive step per token.

Two-stage retrieval, priced

StagePurposeModelCost
Embed / retrieveBroad recall (e.g. top 100 candidates)Qwen3-Embedding-0.6B₹0
RerankPrecision re-scoring (top 5–10)Qwen3-Reranker-0.6B₹0
GenerateFinal answerAny chat modelVaries by model — see /models

The retrieval and reranking steps cost nothing; only the generation step consumes paid tokens, and it consumes far fewer of them because the reranker has already cut the context down to what's actually relevant.

Why add a reranking step at all

  • Better relevance than embeddings alone. Embedding similarity is a fast, approximate signal; a reranker looks at the query and each candidate together, which is a meaningfully more precise (if slower) comparison — worth it for the top-K stage since it's cheap here.
  • Smaller, cheaper generation prompts. Sending your generation model 5 well-chosen documents instead of 30 loosely-relevant ones cuts input tokens on the expensive step.
  • Data stays on unoblox's infrastructure. Reranking runs on the same platform as the rest of your pipeline — no separate vendor, no separate invoice.
  • Production-ready. Same uptime commitment as the rest of unoblox — see /guides/ai-api-uptime-sla-india.

Quick start

import requests

resp = requests.post(
    "https://api.unoblox.ai/v1/rerank",
    headers={"Authorization": "Bearer ub-gw-YOUR_KEY"},
    json={
        "model": "Qwen/Qwen3-Reranker-0.6B",
        "query": "refund policy for annual plans",
        "documents": [doc["text"] for doc in candidates],
        "top_n": 5,
    },
)
ranked = resp.json()["results"]

Same ub-gw-… key as your chat completions — no separate account or credentials.

Frequently asked questions

What's the actual difference between embedding and reranking? Embedding encodes documents into vectors once, ahead of time, for fast approximate search across a large corpus. Reranking scores query and document together, at query time, on a small shortlist — slower per pair, but far more precise, which is why it's used only on the narrowed-down candidates.

Does it handle Hindi or other Indian-language queries? It's trained primarily on English and Chinese; Indian-language queries will work but with lower precision than English — worth testing on your own corpus before relying on it in production.

Can I rerank a large batch of documents in one call? Yes, but latency scales with the number of documents. For very large candidate sets, retrieve a smaller shortlist first (50–200 documents) rather than reranking thousands at once.

Is the relevance score a probability or just a ranking signal? It's a 0–1 relevance score, useful for ordering and for picking a top_n cutoff — treat it as relative between documents in the same request rather than as an absolute confidence threshold.

Can I chain embeddings, reranking, and a chat model in one pipeline? Yes — all three sit behind the same API key with no re-authentication between calls.

Get started in rupees → https://unoblox.ai/sign-in

rerank pricing indiacost
ShareLinkedInWhatsAppTelegram

Start building in rupees

Call every major model through one OpenAI-compatible endpoint, billed in ₹ on a GST invoice.