Skip to content
Start freeGPT-5 · Claude · DeepSeek V4 · Qwen3 — in ₹One OpenAI-compatible endpointBilled in rupeesGST invoiceNo international cardGet started →Start freeGPT-5 · Claude · DeepSeek V4 · Qwen3 — in ₹One OpenAI-compatible endpointBilled in rupeesGST invoiceNo international cardGet started →
Model API Guides

Reranker API in India — Qwen3 Live at ₹0

Native /v1/rerank endpoint on unoblox: Qwen3-Reranker-0.6B live and free. Boost RAG precision with one OpenAI-compatible, rupee-billed key.

Native Reranking API for RAG and Search

unoblox's native /v1/rerank endpoint gives you Cohere-compatible reranking, currently powered by Qwen3-Reranker-0.6B — live today at ₹0. Improve retrieval precision in your RAG pipeline without leaving the gateway or managing a separate vendor.

Why Rerank on unoblox

  • Same key, same endpoint: no extra sign-up — use your existing ub-gw-… key.
  • Free reranker: ₹0 for Qwen3-Reranker-0.6B, so you can scale without worrying about a per-call bill.
  • Cohere-compatible shape: the familiar /v1/rerank request and response format that LangChain, LlamaIndex, and the Vercel AI SDK already know how to call.
  • Rupee invoicing: any future paid reranking still lands on the same monthly GST invoice as the rest of your usage.

unoblox's Native, ₹0 Endpoints

EndpointModelPrice
/v1/rerankQwen3-Reranker-0.6B₹0
/v1/embeddingsQwen3-Embedding-0.6B₹0

Both run on unoblox's own infrastructure alongside the main chat-completions gateway, under the same key.

Reranking in Action

  1. Retrieve: use embeddings to pull your top candidate documents from a vector database.
  2. Rerank: send the query and candidates to /v1/rerank; get back scored, sorted results.
  3. Generate: feed the top few reranked documents to your LLM for the final answer.
curl -X POST https://api.unoblox.ai/v1/rerank \
  -H "Authorization: Bearer ub-gw-…" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "Qwen/Qwen3-Reranker-0.6B",
    "query": "customer billing issue",
    "documents": [{"id": "doc1", "text": "..."}],
    "top_n": 5
  }'

Why Reranking Improves RAG

Vector search alone is approximate — it ranks by embedding similarity, which doesn't always match true relevance. A reranker re-scores your shortlist with a model trained specifically on query-document relevance, so the documents you actually feed to your LLM are the ones most likely to answer the question. That means fewer irrelevant chunks in context, and a shorter, cheaper prompt to your chat model.

Frequently asked questions

Q: How is reranking different from embeddings? A: Embeddings are fast but approximate — good for narrowing millions of documents to a shortlist. Reranking is slower but precise — good for ordering that shortlist by true relevance.

Q: Is the Qwen reranker accurate? A: It's trained specifically on query-document relevance data and holds up well against simpler similarity-only ranking.

Q: Can I rerank in Hindi or other Indian languages? A: Yes — the model handles multilingual queries and documents.

Q: How fast is it? A: Fast enough for real-time RAG — reranking a shortlist of documents adds a small, well-under-a-second step to your pipeline in typical use.

Q: Do you support Cohere Rerank v2 semantics? A: Yes, the endpoint speaks Cohere v2's request and response shape.

Q: Will pricing change if I outgrow the free tier? A: If usage-based pricing is introduced for reranking, it will show live on the model's /models page and bill straight into your existing rupee invoice — no separate account.

Get started in rupees → https://unoblox.ai/sign-in

rerankerllmindiarupeesragapi
ShareLinkedInWhatsAppTelegram

Start building in rupees

Call every major model through one OpenAI-compatible endpoint, billed in ₹ on a GST invoice.