Skip to content
Start freeGPT-5 · Claude · DeepSeek V4 · Qwen3 — in ₹One OpenAI-compatible endpointBilled in rupeesGST invoiceNo international cardGet started →Start freeGPT-5 · Claude · DeepSeek V4 · Qwen3 — in ₹One OpenAI-compatible endpointBilled in rupeesGST invoiceNo international cardGet started →
Use Cases

Build RAG with an AI API in India (₹ rupees)

RAG in India: embed documents, retrieve with semantic search, augment with LLM. See ₹ costs, model choice (Qwen or Claude), and 3-step setup.

Build RAG with an AI API in India

Retrieval-augmented generation (RAG) lets your AI assistant "read" documents and answer questions grounded in real data. In India, unoblox handles the full pipeline: embeddings (free with Qwen3), retrieval (your DB), and generation (Claude or Qwen). Billed in ₹ rupees.

Here's the 3-step setup + real ₹ costs.

Why unoblox for RAG

  • Free embeddings: Qwen3-Embedding-0.6B (₹0 per 1M tokens)
  • Production LLMs: Qwen3 Max or Claude Sonnet at one endpoint
  • OpenAI-compatible: Works with LangChain, LlamaIndex
  • One key, one invoice: Manage embeddings + generation in rupees
  • GST invoice: Input-tax-credit claimable

Recommended models

  • Qwen3-Embedding-0.6B: Free, strong general-purpose retrieval — validate recall on your own corpus before scaling
  • Qwen3 Max (generation): ₹120.95 input / ₹604.77 output per 1M tokens
  • Claude Sonnet (generation): ₹201.6 input / ₹1,008 output per 1M tokens

Step 1: Embed documents

Take a PDF or document and turn it into a vector. Store it in Pinecone or Postgres.

from openai import OpenAI

client = OpenAI(
    base_url="https://api.unoblox.ai/v1",
    api_key="ub-gw-..."
)

response = client.embeddings.create(
    model="qwen/qwen3-embedding-0.6b",
    input="Your document text..."
)

vector = response.data[0].embedding
# Store (doc_id, vector) in your database

Step 2: Retrieve relevant documents

When a user asks a question, embed the query and find similar documents.

# Embed the query
query_response = client.embeddings.create(
    model="qwen/qwen3-embedding-0.6b",
    input="What are the refund policies?"
)
query_vector = query_response.data[0].embedding

# Find similar documents (using Postgres + pgvector)
results = db.query(
    "SELECT content FROM docs ORDER BY embedding <-> %s LIMIT 5",
    (query_vector,)
)

Step 3: Augment with LLM generation

Pass retrieved documents to the LLM for a grounded answer.

context = "\n".join([r[0] for r in results])

response = client.chat.completions.create(
    model="qwen/qwen-3-max",
    messages=[
        {
            "role": "system",
            "content": "Answer based on the documents below.\n\n" + context
        },
        {"role": "user", "content": "What are the refund policies?"}
    ]
)

print(response.choices[0].message.content)

Real ₹ cost: 10,000 documents

Setup (one-time): embedding 10,000 docs at ~300 tokens average is about 3M tokens — free on Qwen3-Embedding-0.6B within the freemium cap; check /models for the current cap and any rate beyond it.

Per-query costs (ongoing):

  • Embed the query: ₹0 (Qwen3-Embedding-0.6B, freemium)
  • Retrieve: free (your own database)
  • Generate the answer: ~₹0.02/query on Qwen3 Max · ~₹0.20/query on Claude Sonnet

Monthly, at 1,000 queries:

  • Qwen3 Max generation: ≈₹20–30/month
  • Claude Sonnet generation: ≈₹200–300/month

Production tips

  1. Chunk documents smartly: 256–512 tokens per chunk (overlap by 50)
  2. Hybrid search: Combine BM25 (keyword) + semantic (vector)
  3. Monitor latency: Embedding: 10ms; retrieval depends on DB
  4. Re-embed quarterly: Refresh as documents update
  5. Track ₹ spend: Dashboard shows cost per query

Frequently asked questions

Q: What size documents can I embed? Up to 8,000 tokens per embedding. Split larger docs into chunks.

Q: Qwen or another embedding model? Qwen3-Embedding-0.6B is free and a solid default for most RAG builds. If you need a different embedding model, check live pricing and specs on /models before committing — don't assume a rate that isn't listed there.

Q: How do I handle real-time document updates? Re-embed updated chunks and upsert into your vector DB. No need to re-embed all 10k.

Q: Can I use RAG with Claude? Yes. Claude excels at reasoning over long context. Embed with Qwen (₹0), generate with Claude (₹0.20/query).

Q: How many documents can I store? Unlimited in your DB (Pinecone, Postgres, Weaviate). Costs scale with embedding count and query volume.

Q: Does unoblox support vector DBs? Yes. Pinecone, Weaviate, Milvus, Postgres+pgvector all work. Use their SDKs; unoblox provides embeddings and generation.

Get started in rupees

Get started → https://unoblox.ai/sign-in

RAGembeddingsretrievaluse-case
ShareLinkedInWhatsAppTelegram

Start building in rupees

Call every major model through one OpenAI-compatible endpoint, billed in ₹ on a GST invoice.