Skip to content
Start freeGPT-5 · Claude · DeepSeek V4 · Qwen3 — in ₹One OpenAI-compatible endpointBilled in rupeesGST invoiceNo international cardGet started →Start freeGPT-5 · Claude · DeepSeek V4 · Qwen3 — in ₹One OpenAI-compatible endpointBilled in rupeesGST invoiceNo international cardGet started →
API Pricing Guides

Embeddings pricing in India (₹)

Qwen3-Embedding-0.6B native API in India at ₹0 per 1M tokens. OpenAI-compatible /v1/embeddings, billed on GST invoice.

Vector embeddings at zero cost in India

Embeddings have never been cheaper. Qwen3-Embedding-0.6B is free (₹0 per 1M tokens) on unoblox, billed in rupees on your monthly GST invoice from an Indian entity. No subscription, no card, no surprise fees.

Native /v1/embeddings endpoint

Point your Python / Node.js / LangChain client directly at our base URL:

base_url = "https://api.unoblox.ai/v1"
model = "qwen/qwen3-embedding-0.6b"

A full request, OpenAI-SDK style:

from openai import OpenAI
client = OpenAI(base_url="https://api.unoblox.ai/v1", api_key="ub-gw-...")

resp = client.embeddings.create(
    model="qwen/qwen3-embedding-0.6b",
    input=["unoblox bills every model in rupees."]
)
vector = resp.data[0].embedding  # 1024-dim list of floats

Async calls, batched arrays of input text, and standard OpenAI-SDK patterns all work unchanged. Use it for RAG, semantic search, clustering, or as the retrieval layer under an AI agent.

When to use Qwen3 Embedding

Use caseDimensionSpeedBest for
Semantic search1024<50msE-commerce, docs, knowledge bases
RAG chunk retrieval1024<50msLLM context enrichment
Clustering1024BatchUser segmentation
Reranking companion1024FastTwo-stage retrieval

Note the 1024-dimensional vector when you provision your vector database — that's the fixed native output size for this model on unoblox.

Why free embeddings matter

  • Cost: ₹0. Build at scale without worrying about per-token fees.
  • Sovereignty. Data stays in India; no US/EU pipeline.
  • Invoice. Every 1M vector tokens on your GST bill — claimable as input-tax credit (ITC) if registered.
  • No limits. Soft freemium caps only; upgrade pricing comes later.

Pairing embeddings with reranking

Embeddings alone get you fast approximate search; adding a reranker sharpens the final result. unoblox also serves Qwen3-Reranker-0.6B natively at ₹0 on /v1/rerank — a typical two-stage RAG pipeline embeds your corpus once (free), retrieves the top 20–50 candidates by vector similarity, then reranks them down to the top 3–5 before they reach your LLM prompt. Both stages run on the same base URL and API key as everything else on unoblox.

Frequently asked questions

Q: Why is it free? Qwen3-Embedding-0.6B is a compact, efficient open-weight model that we host in India. We offer it at ₹0 on the free tier so teams can build RAG and semantic search without per-token embedding costs — you only pay, in rupees, for the chat models you call.

Q: Do I need an API key? Yes. Generate one free at /sign-in → Playground → Keys. Lasts 30 days; rotate anytime.

Q: Can I use it in production? Absolutely. It runs under the same 99.95% SLA as every other model on unoblox, with logs retained 90–180 days and DPDP-compliant processing — free tier isn't a lesser tier for reliability.

Q: How many vectors can I store? No hard limit on storage on our side — you're generating vectors, not storing them with us. The soft free-tier cap is on monthly embedding requests (₹250/month equivalent), and it scales transparently the moment you upgrade.

Q: Does it support batch embeddings? Yes. Send arrays of texts; get vectors back in order. Native OpenAI format.

Q: What about image embeddings? Qwen3-Embedding is text-only. Vision models (Claude, GPT-4o) handle images; rerank text results.

Get started in rupees → https://unoblox.ai/sign-in

embeddingscostfree
ShareLinkedInWhatsAppTelegram

Start building in rupees

Call every major model through one OpenAI-compatible endpoint, billed in ₹ on a GST invoice.