Skip to content
Start freeGPT-5 · Claude · DeepSeek V4 · Qwen3 — in ₹One OpenAI-compatible endpointBilled in rupeesGST invoiceNo international cardGet started →Start freeGPT-5 · Claude · DeepSeek V4 · Qwen3 — in ₹One OpenAI-compatible endpointBilled in rupeesGST invoiceNo international cardGet started →
Model API Guides

Embeddings API in India — Qwen3 Free

Native embeddings API for RAG and semantic search in India. Qwen3-Embedding-0.6B is free at ₹0, OpenAI-compatible /v1/embeddings.

unoblox exposes a native /v1/embeddings endpoint for building retrieval and search in India — no separate account, no foreign invoice, just the same ub-gw-… key you already use for chat models. The first embedding model live on it, Qwen3-Embedding-0.6B, is priced at ₹0.

Why a native embeddings endpoint matters

Most RAG stacks need three things from one place: chat, embeddings, and reranking. unoblox covers all three natively — /v1/chat/completions, /v1/embeddings, and /v1/rerank — on one key, one GST invoice, one dashboard. You don't stitch together a separate embeddings provider just to build a search pipeline.

What's live today

ModelDimensionsCost (₹/1M tokens)Status
Qwen3-Embedding-0.6B1024₹0 (free)Live
Additional embedding models—See /modelsRolling out

At ₹0, every embedding request settles to zero regardless of how many texts you send — useful for prototyping a RAG pipeline before you commit spend to the chat model that sits on top of it.

How to call it

from openai import OpenAI

client = OpenAI(api_key="ub-gw-...", base_url="https://api.unoblox.ai/v1")
result = client.embeddings.create(
    model="qwen/qwen3-embedding-0.6b",
    input=["What is unoblox?", "How is billing calculated?"],
)
vectors = [row.embedding for row in result.data]

Swap in your own texts, write the vectors to your store of choice — Pinecone, Weaviate, pgvector, or anything that takes a plain float array — and you have a working retrieval index.

Common use cases

  1. RAG pipelines — embed documents at indexing time, embed the user's question at query time, retrieve the closest matches, then hand them to a chat model for the final answer.
  2. Semantic search — index product or article descriptions and rank results by meaning rather than exact keyword match.
  3. Deduplication — flag near-duplicate support tickets, listings, or documents by embedding distance.
  4. Recommendations — embed user history and catalog items, then recommend by vector proximity.

Frequently asked questions

Is a free model good enough for production? Qwen3-Embedding-0.6B is a genuine production embedding model, not a limited trial tier — it stays ₹0 as long as it's marked free on /models, and you're billed normally the moment you call a paid model on the same key.

Does it work with LangChain or LlamaIndex? Yes. Both frameworks support OpenAI-style embeddings clients — point the base URL at unoblox and the model string at qwen/qwen3-embedding-0.6b.

What's a sensible batch size? Batch multiple texts into one request for throughput; send them one at a time only when you need the lowest possible latency on a single query.

Can I combine embeddings with the native reranker? Yes — embed and retrieve a shortlist first, then call /v1/rerank to reorder that shortlist by relevance before it reaches your chat model.

Will unoblox add paid embedding models? Yes, more are rolling out; /models always shows the current line-up and live rupee rates, so pricing here never needs to be guessed at.

Do embeddings count against the same balance as chat models? Yes — one balance, one wallet, topped up through Razorpay, shared across chat, embeddings, and reranking.

Get started in rupees → https://unoblox.ai/sign-in

embeddingsllmindiarupeesragapi
ShareLinkedInWhatsAppTelegram

Start building in rupees

Call every major model through one OpenAI-compatible endpoint, billed in ₹ on a GST invoice.