AI Product Search in India
Build rupee-billed semantic product search with unoblox's native embeddings, reranking, and chat models — one API, one key, transparent pricing.
Shoppers increasingly expect to type a plain-language query — a cotton kurta under fifteen hundred rupees, size M — and get relevant results, not just keyword matches. Building that kind of search means combining unoblox's native embeddings and reranking endpoints with a chat model for query understanding, all billed in rupees through one endpoint.
The anatomy of AI-powered product search
Semantic product search generally has three parts: an embedding step that turns product titles and descriptions into vectors for similarity search, a reranking step that reorders the top candidates by relevance to the actual query, and, optionally, a chat model that parses a natural-language query into structured filters (price ceiling, size, category) before the search even runs. unoblox exposes all three as native endpoints.
Where each unoblox endpoint fits
/v1/embeddings— embed your product catalog once, then embed each incoming query for similarity search against it/v1/rerank— reorder your top candidate results by relevance before showing them to the shopper/v1/chat/completions— parse a natural-language query into structured filters, or generate a short explanation of why a result matched
The embedding and reranking models in the catalog change as new ones are added — check /models for the current model and its live rupee rate rather than assuming a fixed price.
Chat models for query parsing
| Model | ₹ input / ₹ output (per 1M tokens) | Notes |
|---|---|---|
deepseek-ai/deepseek-v4-flash | ₹9.07 / ₹18.14 | Cheapest option — good for parsing high query volumes. |
qwen/qwen3-235b-a22b-instruct-2507 | ₹9.07 / ₹55.44 | Stronger structured-extraction quality at a similar input rate. |
openai/gpt-5-mini | ₹25.2 / ₹201.6 | A balanced mid-tier option. |
qwen/qwen3-1.7b | FREE (₹0) | Free for prototyping your query-parsing prompts. |
Sample query-parsing call
curl https://api.unoblox.ai/v1/chat/completions \
-H "Authorization: Bearer ub-gw-xxxxxxxxxxxxxxxx" \
-H "Content-Type: application/json" \
-d '{
"model": "deepseek-ai/deepseek-v4-flash",
"messages": [{"role": "user", "content": "Extract category, max price in rupees, and size from this search query as JSON: cotton kurta under fifteen hundred rupees size M"}]
}'
Budgeting the pipeline
Because every step — embedding, reranking, and query parsing — is billed per million tokens in rupees with a single monthly GST invoice, you can estimate the full pipeline's cost by multiplying expected query volume by each step's rate, without juggling separate vendor bills for each part.
Frequently asked questions
Do I need three separate vendors for embeddings, reranking, and chat? No — unoblox exposes all three as native endpoints behind one key and one rupee invoice.
Which embedding or reranking model does unoblox use? This changes as the catalog is updated — check /models for the current model and its live rate.
Can a chat model replace embeddings for search? Not efficiently at catalog scale — embeddings plus reranking are built for fast similarity search over large catalogs; use a chat model for the query-understanding step, not for scoring every product directly.
Is there a free way to prototype query parsing? Yes, qwen/qwen3-1.7b is free (₹0).
Is pricing in rupees across all three endpoints? Yes — embeddings, reranking, and chat completions are all priced per million tokens in rupees.
Do I need a new SDK for the chat portion? No — it's OpenAI-compatible, so existing chat-completions code works by pointing at https://api.unoblox.ai/v1.
Get started in rupees → https://unoblox.ai/sign-in
Start building in rupees
Call every major model through one OpenAI-compatible endpoint, billed in ₹ on a GST invoice.