Embeddings API in India — Qwen3 Free
Native embeddings API for RAG and semantic search in India. Qwen3-Embedding-0.6B is free at ₹0, OpenAI-compatible /v1/embeddings.
unoblox exposes a native /v1/embeddings endpoint for building retrieval and search in India — no separate account, no foreign invoice, just the same ub-gw-… key you already use for chat models. The first embedding model live on it, Qwen3-Embedding-0.6B, is priced at ₹0.
Why a native embeddings endpoint matters
Most RAG stacks need three things from one place: chat, embeddings, and reranking. unoblox covers all three natively — /v1/chat/completions, /v1/embeddings, and /v1/rerank — on one key, one GST invoice, one dashboard. You don't stitch together a separate embeddings provider just to build a search pipeline.
What's live today
| Model | Dimensions | Cost (₹/1M tokens) | Status |
|---|---|---|---|
| Qwen3-Embedding-0.6B | 1024 | ₹0 (free) | Live |
| Additional embedding models | — | See /models | Rolling out |
At ₹0, every embedding request settles to zero regardless of how many texts you send — useful for prototyping a RAG pipeline before you commit spend to the chat model that sits on top of it.
How to call it
from openai import OpenAI
client = OpenAI(api_key="ub-gw-...", base_url="https://api.unoblox.ai/v1")
result = client.embeddings.create(
model="qwen/qwen3-embedding-0.6b",
input=["What is unoblox?", "How is billing calculated?"],
)
vectors = [row.embedding for row in result.data]
Swap in your own texts, write the vectors to your store of choice — Pinecone, Weaviate, pgvector, or anything that takes a plain float array — and you have a working retrieval index.
Common use cases
- RAG pipelines — embed documents at indexing time, embed the user's question at query time, retrieve the closest matches, then hand them to a chat model for the final answer.
- Semantic search — index product or article descriptions and rank results by meaning rather than exact keyword match.
- Deduplication — flag near-duplicate support tickets, listings, or documents by embedding distance.
- Recommendations — embed user history and catalog items, then recommend by vector proximity.
Frequently asked questions
Is a free model good enough for production?
Qwen3-Embedding-0.6B is a genuine production embedding model, not a limited trial tier — it stays ₹0 as long as it's marked free on /models, and you're billed normally the moment you call a paid model on the same key.
Does it work with LangChain or LlamaIndex?
Yes. Both frameworks support OpenAI-style embeddings clients — point the base URL at unoblox and the model string at qwen/qwen3-embedding-0.6b.
What's a sensible batch size? Batch multiple texts into one request for throughput; send them one at a time only when you need the lowest possible latency on a single query.
Can I combine embeddings with the native reranker?
Yes — embed and retrieve a shortlist first, then call /v1/rerank to reorder that shortlist by relevance before it reaches your chat model.
Will unoblox add paid embedding models?
Yes, more are rolling out; /models always shows the current line-up and live rupee rates, so pricing here never needs to be guessed at.
Do embeddings count against the same balance as chat models? Yes — one balance, one wallet, topped up through Razorpay, shared across chat, embeddings, and reranking.
Get started in rupees → https://unoblox.ai/sign-in
Start building in rupees
Call every major model through one OpenAI-compatible endpoint, billed in ₹ on a GST invoice.