Skip to content
Start freeGPT-5 · Claude · DeepSeek V4 · Qwen3 — in ₹One OpenAI-compatible endpointBilled in rupeesGST invoiceNo international cardGet started →Start freeGPT-5 · Claude · DeepSeek V4 · Qwen3 — in ₹One OpenAI-compatible endpointBilled in rupeesGST invoiceNo international cardGet started →
Integrations

LlamaIndex with an India AI API

Build LlamaIndex RAG pipelines on unoblox — free Qwen embeddings plus Claude, GPT-4o, or DeepSeek for answers, billed monthly in rupees with GST.

LlamaIndex (formerly GPT Index) is the lightweight RAG framework — load documents, embed them, query, generate answers. unoblox makes it ₹-native: keep the LlamaIndex code unchanged, swap in cheaper or India-priced models, and bill everything in rupees.

Why LlamaIndex + unoblox

  • Minimal: simpler than LangChain for pure RAG workflows
  • Model-agnostic: works with any OpenAI-compatible provider, including unoblox
  • Free embeddings: Qwen3-Embedding-0.6B is ₹0 on unoblox's native /v1/embeddings
  • Fast retrieval: built for large document sets
  • ₹ billing: one GST invoice, no international card needed

Two-line setup

from llama_index.core import VectorStoreIndex, SimpleDirectoryReader
from llama_index.llms.openai import OpenAI
from llama_index.embeddings.openai import OpenAIEmbedding

llm = OpenAI(
  api_key="ub-gw-YOUR_API_KEY",
  api_base="https://api.unoblox.ai/v1",
  model="openai/gpt-4o"
)

embeddings = OpenAIEmbedding(
  api_key="ub-gw-YOUR_API_KEY",
  api_base="https://api.unoblox.ai/v1",
  model="Qwen/Qwen3-Embedding-0.6B"
)

documents = SimpleDirectoryReader(input_dir="/path/to/docs").load_data()
index = VectorStoreIndex.from_documents(documents, embed_model=embeddings)
query_engine = index.as_query_engine(llm=llm)
response = query_engine.query("What's the revenue trend?")

Cost breakdown (₹ per 1M tokens)

ComponentModel₹ per 1M inputPurpose
EmbeddingsQwen3-Embedding-0.6BFREEvectorize documents once
LLM answerQwen3.8-27B₹16.32fast retrieval + summary
LLM answerClaude Opus₹504nuanced, long-form analysis

Embedding a large document set costs ₹0 up front; each query then costs only the LLM's per-token rate on the (small) retrieved context plus the answer.

Three-model RAG example

# Budget mode: free embeddings, cheap LLM
llm_budget = OpenAI(
  api_key="ub-gw-...",
  api_base="https://api.unoblox.ai/v1",
  model="deepseek-ai/deepseek-v4-flash"  # ₹9.07 / ₹18.14
)

# Balanced mode
llm_balanced = OpenAI(
  api_key="ub-gw-...",
  api_base="https://api.unoblox.ai/v1",
  model="openai/gpt-4o"  # ₹252 / ₹1008
)

# Premium mode
llm_premium = OpenAI(
  api_key="ub-gw-...",
  api_base="https://api.unoblox.ai/v1",
  model="anthropic/claude-opus-5"  # ₹504 / ₹2520
)

documents = SimpleDirectoryReader("/data").load_data()
index = VectorStoreIndex.from_documents(documents)
query = "Analyze Q3 financial performance trends"
engine = index.as_query_engine(llm=llm_premium if "complex" in query.lower() else llm_budget)
answer = engine.query(query)

One unoblox key, three LLMs, cost matched to query difficulty.

Common questions

Does LlamaIndex cache embeddings? Yes. The first run embeds every document once; later queries reuse the stored vectors — no re-embedding, no repeat cost.

Can I use LlamaIndex with reranking? Yes. unoblox's free Qwen3-Reranker-0.6B works with LlamaIndex's reranking postprocessors via /v1/rerank.

What's the difference between api_base and base_url? LlamaIndex's OpenAI wrapper uses api_base; the raw OpenAI SDK uses base_url. Both point at https://api.unoblox.ai/v1 — check which class you're instantiating.

Can I stream responses? Yes. query_engine.aquery() (async) works with unoblox streaming; consume the response incrementally as tokens arrive.

How do I export answers to a file? LlamaIndex's response object exposes .response (text), .source_nodes (citations), and .metadata (token counts) — serialize whichever fields your logging layer needs.

Is there a Node.js version of LlamaIndex? Yes, npm install llamaindex. Configuration mirrors the Python client — one base URL, one key, the same model ids.

Get started in rupees → https://unoblox.ai/sign-in

llamaindexragintegrationsembeddings
ShareLinkedInWhatsAppTelegram

Start building in rupees

Call every major model through one OpenAI-compatible endpoint, billed in ₹ on a GST invoice.