Embeddings pricing in India (₹)
Qwen3-Embedding-0.6B native API in India at ₹0 per 1M tokens. OpenAI-compatible /v1/embeddings, billed on GST invoice.
Vector embeddings at zero cost in India
Embeddings have never been cheaper. Qwen3-Embedding-0.6B is free (₹0 per 1M tokens) on unoblox, billed in rupees on your monthly GST invoice from an Indian entity. No subscription, no card, no surprise fees.
Native /v1/embeddings endpoint
Point your Python / Node.js / LangChain client directly at our base URL:
base_url = "https://api.unoblox.ai/v1"
model = "qwen/qwen3-embedding-0.6b"
A full request, OpenAI-SDK style:
from openai import OpenAI
client = OpenAI(base_url="https://api.unoblox.ai/v1", api_key="ub-gw-...")
resp = client.embeddings.create(
model="qwen/qwen3-embedding-0.6b",
input=["unoblox bills every model in rupees."]
)
vector = resp.data[0].embedding # 1024-dim list of floats
Async calls, batched arrays of input text, and standard OpenAI-SDK patterns all work unchanged. Use it for RAG, semantic search, clustering, or as the retrieval layer under an AI agent.
When to use Qwen3 Embedding
| Use case | Dimension | Speed | Best for |
|---|---|---|---|
| Semantic search | 1024 | <50ms | E-commerce, docs, knowledge bases |
| RAG chunk retrieval | 1024 | <50ms | LLM context enrichment |
| Clustering | 1024 | Batch | User segmentation |
| Reranking companion | 1024 | Fast | Two-stage retrieval |
Note the 1024-dimensional vector when you provision your vector database — that's the fixed native output size for this model on unoblox.
Why free embeddings matter
- Cost: ₹0. Build at scale without worrying about per-token fees.
- Sovereignty. Data stays in India; no US/EU pipeline.
- Invoice. Every 1M vector tokens on your GST bill — claimable as input-tax credit (ITC) if registered.
- No limits. Soft freemium caps only; upgrade pricing comes later.
Pairing embeddings with reranking
Embeddings alone get you fast approximate search; adding a reranker sharpens the final result. unoblox also serves Qwen3-Reranker-0.6B natively at ₹0 on /v1/rerank — a typical two-stage RAG pipeline embeds your corpus once (free), retrieves the top 20–50 candidates by vector similarity, then reranks them down to the top 3–5 before they reach your LLM prompt. Both stages run on the same base URL and API key as everything else on unoblox.
Frequently asked questions
Q: Why is it free? Qwen3-Embedding-0.6B is a compact, efficient open-weight model that we host in India. We offer it at ₹0 on the free tier so teams can build RAG and semantic search without per-token embedding costs — you only pay, in rupees, for the chat models you call.
Q: Do I need an API key? Yes. Generate one free at /sign-in → Playground → Keys. Lasts 30 days; rotate anytime.
Q: Can I use it in production? Absolutely. It runs under the same 99.95% SLA as every other model on unoblox, with logs retained 90–180 days and DPDP-compliant processing — free tier isn't a lesser tier for reliability.
Q: How many vectors can I store? No hard limit on storage on our side — you're generating vectors, not storing them with us. The soft free-tier cap is on monthly embedding requests (₹250/month equivalent), and it scales transparently the moment you upgrade.
Q: Does it support batch embeddings? Yes. Send arrays of texts; get vectors back in order. Native OpenAI format.
Q: What about image embeddings? Qwen3-Embedding is text-only. Vision models (Claude, GPT-4o) handle images; rerank text results.
Get started in rupees → https://unoblox.ai/sign-in
More from unoblox
API Pricing Guides
Vision LLM API pricing in India
API Pricing Guides
Batch AI API pricing in India
API Pricing Guides
RAG pipeline cost in India
API Pricing Guides
AI agent running cost in India
API Pricing Guides
OpenAI o3 Pricing in India (Live ₹ Rate)
API Pricing Guides
Cheapest Vision LLM in India (₹ Pricing)
Start building in rupees
Call every major model through one OpenAI-compatible endpoint, billed in ₹ on a GST invoice.