Knowledge Base AI in India | unoblox
Build a knowledge-base assistant with unoblox's native embeddings, rerank, and chat models — one rupee-billed endpoint, retrieval pipeline inside.
A knowledge-base assistant that answers from your own documents, instead of a model's general training, needs three pieces working together: a way to search your documents for relevant passages, a way to rank those passages by actual relevance, and a chat model to turn the best passages into a plain-language answer. unoblox serves all three natively — /v1/embeddings, /v1/rerank, and /v1/chat/completions — behind one rupee-billed key, so a retrieval pipeline doesn't mean stitching together three separate vendors.
What a grounded knowledge-base assistant needs
Asking a general-purpose model a question about your internal policy document without giving it the document first just invites a confident-sounding guess. The standard fix — retrieval-augmented generation — searches your own content for the passages most likely to answer the question, and only then asks the model to answer using those passages, explicitly instructed not to go beyond them.
The retrieval pipeline
- Embed your documents in advance, chunked into passages, using
/v1/embeddings. - Embed the user's question the same way at query time.
- Search your vector store for the closest-matching chunks.
- Rerank the top candidates with
/v1/rerankto push the truly relevant passages to the top — embeddings alone are a coarse first pass. - Generate the final answer with a chat model, given only the reranked passages plus the question.
curl https://api.unoblox.ai/v1/embeddings \
-H "Authorization: Bearer ub-gw-your-key-here" \
-H "Content-Type: application/json" \
-d '{"model":"qwen/qwen3-1.7b","input":"How do I reset my password?"}'
(Check /models for the current embedding and rerank model ids and their ₹ pricing — this endpoint shape stays the same regardless of which one you pick.)
Picking a chat model for the answer step
| Stage | Endpoint | Notes |
|---|---|---|
| Embedding | /v1/embeddings | run once per document chunk, then per query |
| Reranking | /v1/rerank | reorders candidates by relevance before generation |
| Answer generation | /v1/chat/completions | DeepSeek V4 Flash (₹9.07 / ₹18.14) for high-volume Q&A, or Claude Sonnet (₹201.6 / ₹1008) where answer quality on nuanced questions matters more than cost |
Keeping answers grounded
Instruct the answer-generation step explicitly to say it doesn't know rather than guess when the retrieved passages don't cover the question — this single instruction prevents most of the confident-but-wrong answers a knowledge-base bot can produce. It's also worth showing users which source passage an answer came from, so they can verify it themselves rather than trusting the bot blindly.
Frequently asked questions
Do I need a separate vector database, or does unoblox provide one? unoblox provides the embedding and rerank models; storing and searching the resulting vectors is up to your own vector database or search index.
Is my internal knowledge base processed on Indian infrastructure? It depends on which models you choose — unoblox's own small hosted models process end-to-end on Indian infrastructure, while third-party chat models process on that provider's infrastructure, with the India benefit there being rupee billing and one GST invoice rather than data residency.
Why do I need reranking if I already have embeddings-based search? Embedding similarity is a fast, approximate first pass; reranking re-scores a shortlist more carefully, which usually improves the passages that actually reach the answer-generation step.
Can this work with Hindi or regional-language documents? The underlying models handle multiple languages, but we don't publish per-language accuracy figures — test against a sample of your own documents before relying on it for a specific language.
How is this billed? Each step — embedding, rerank, and chat generation — is billed per token in ₹ under the same account, consolidated on one monthly GST invoice.
What's the free Qwen3 1.7B model good for here? It's a reasonable, no-cost starting point for the embedding step or for lightweight answer generation while you're prototyping the pipeline, before deciding whether a larger model is worth the additional per-token cost.
Get started in rupees → https://unoblox.ai/sign-in
Start building in rupees
Call every major model through one OpenAI-compatible endpoint, billed in ₹ on a GST invoice.