Skip to content
Start freeGPT-5 · Claude · DeepSeek V4 · Qwen3 — in ₹One OpenAI-compatible endpointBilled in rupeesGST invoiceNo international cardGet started →Start freeGPT-5 · Claude · DeepSeek V4 · Qwen3 — in ₹One OpenAI-compatible endpointBilled in rupeesGST invoiceNo international cardGet started →
Use Cases

Semantic Search API in India

Build vector similarity search on India's ₹-native AI gateway. Free Qwen3 embeddings, one API key, GST invoice — no international card needed.

Semantic Search API in India

Semantic search finds the right answer by matching meaning, not exact keywords. Pair unoblox's embeddings endpoint with your own vector store to build search or retrieval that understands what a user actually meant — billed entirely in rupees from an Indian entity.

How it works

unoblox exposes a native /v1/embeddings endpoint, compatible with any OpenAI SDK. You turn your documents and incoming queries into vectors, store them, and rank results by similarity.

StepActionModel / endpoint
1Embed your corpusQwen3-Embedding-0.6B (₹0)
2Store vectorsYour choice: Postgres+pgvector, Pinecone, Weaviate
3Embed the query, rank by cosine similarityRetrieve top-k results

Step-by-step

Step 1: Embed your documents

POST https://api.unoblox.ai/v1/embeddings
Authorization: Bearer ub-gw-...
Content-Type: application/json

{"model": "qwen/qwen3-embedding-0.6b", "input": ["Your document text here"]}

Step 2: Store the vectors in Postgres with pgvector, Pinecone, or Weaviate — embeddings are free, so your only ongoing cost is the database itself.

Step 3: Query — embed the user's question the same way, then find nearest neighbors by cosine distance and return the matching documents.

Real rupee pricing

Qwen3-Embedding-0.6B costs ₹0 on unoblox, up to the freemium cap. No international card, nothing to reconcile across currencies — your monthly invoice is in rupees, GST claimable via input tax credit.

Getting better recall

  • Chunk sensibly: 256–512 tokens per chunk, with a small overlap, so you don't split a sentence's meaning across two vectors.
  • Hybrid search: blend keyword search (BM25) with vector similarity — pure semantic search can miss exact model numbers, IDs, or codes.
  • Re-embed on update: refresh a document's vector whenever its source text changes; stale vectors quietly hurt recall over time.
  • Evaluate on your own data: run a small labelled test set through your pipeline before trusting recall numbers from any public benchmark.

Build and deploy faster

There's no need to self-host an embedding model or pay a foreign provider per token. One endpoint, one API key, rupee billing — integrate with LangChain, LlamaIndex, or raw Python in under an hour.

Frequently asked questions

Q: How good are Qwen3-Embedding vectors for Indian-language content? A: Strong on English; for Hindi, Tamil, and other Indian languages, test on your own corpus first. If cross-language precision matters, compare results against a Claude Sonnet re-ranking pass.

Q: Can I scale to millions of embeddings? A: Yes — the API is stateless, so your real bottleneck is vector storage and query latency in your chosen database, not the embeddings endpoint itself.

Q: Do I need a separate vector database? A: Yes. unoblox returns vectors; you still need Postgres+pgvector, Pinecone, or Milvus to index and search them.

Q: What does it cost to embed 10 million tokens a month? A: ₹0 on the freemium tier, up to the current cap — check /models for the exact limit. Beyond that, your cost is the vector database, not the embedding call.

Q: Should I use semantic search alone, or combine it with keyword search? A: Combine them for production. Semantic search handles paraphrases and intent; keyword search still wins on exact codes, SKUs, and rare terms.

Q: Is this ready for production traffic today? A: Yes — the endpoint is live and stateless. Start with a small test corpus, check recall on your own queries, then scale up.

Get started in rupees → https://unoblox.ai/sign-in

semantic searchembeddingsraguse-cases
ShareLinkedInWhatsAppTelegram

Start building in rupees

Call every major model through one OpenAI-compatible endpoint, billed in ₹ on a GST invoice.