Reranker API in India — Qwen3 Live at ₹0
Native /v1/rerank endpoint on unoblox: Qwen3-Reranker-0.6B live and free. Boost RAG precision with one OpenAI-compatible, rupee-billed key.
Native Reranking API for RAG and Search
unoblox's native /v1/rerank endpoint gives you Cohere-compatible reranking, currently powered by Qwen3-Reranker-0.6B — live today at ₹0. Improve retrieval precision in your RAG pipeline without leaving the gateway or managing a separate vendor.
Why Rerank on unoblox
- Same key, same endpoint: no extra sign-up — use your existing
ub-gw-…key. - Free reranker: ₹0 for Qwen3-Reranker-0.6B, so you can scale without worrying about a per-call bill.
- Cohere-compatible shape: the familiar
/v1/rerankrequest and response format that LangChain, LlamaIndex, and the Vercel AI SDK already know how to call. - Rupee invoicing: any future paid reranking still lands on the same monthly GST invoice as the rest of your usage.
unoblox's Native, ₹0 Endpoints
| Endpoint | Model | Price |
|---|---|---|
/v1/rerank | Qwen3-Reranker-0.6B | ₹0 |
/v1/embeddings | Qwen3-Embedding-0.6B | ₹0 |
Both run on unoblox's own infrastructure alongside the main chat-completions gateway, under the same key.
Reranking in Action
- Retrieve: use embeddings to pull your top candidate documents from a vector database.
- Rerank: send the query and candidates to
/v1/rerank; get back scored, sorted results. - Generate: feed the top few reranked documents to your LLM for the final answer.
curl -X POST https://api.unoblox.ai/v1/rerank \
-H "Authorization: Bearer ub-gw-…" \
-H "Content-Type: application/json" \
-d '{
"model": "Qwen/Qwen3-Reranker-0.6B",
"query": "customer billing issue",
"documents": [{"id": "doc1", "text": "..."}],
"top_n": 5
}'
Why Reranking Improves RAG
Vector search alone is approximate — it ranks by embedding similarity, which doesn't always match true relevance. A reranker re-scores your shortlist with a model trained specifically on query-document relevance, so the documents you actually feed to your LLM are the ones most likely to answer the question. That means fewer irrelevant chunks in context, and a shorter, cheaper prompt to your chat model.
Frequently asked questions
Q: How is reranking different from embeddings? A: Embeddings are fast but approximate — good for narrowing millions of documents to a shortlist. Reranking is slower but precise — good for ordering that shortlist by true relevance.
Q: Is the Qwen reranker accurate? A: It's trained specifically on query-document relevance data and holds up well against simpler similarity-only ranking.
Q: Can I rerank in Hindi or other Indian languages? A: Yes — the model handles multilingual queries and documents.
Q: How fast is it? A: Fast enough for real-time RAG — reranking a shortlist of documents adds a small, well-under-a-second step to your pipeline in typical use.
Q: Do you support Cohere Rerank v2 semantics? A: Yes, the endpoint speaks Cohere v2's request and response shape.
Q: Will pricing change if I outgrow the free tier?
A: If usage-based pricing is introduced for reranking, it will show live on the model's /models page and bill straight into your existing rupee invoice — no separate account.
Get started in rupees → https://unoblox.ai/sign-in
Start building in rupees
Call every major model through one OpenAI-compatible endpoint, billed in ₹ on a GST invoice.