LlamaIndex with an India AI API
Build LlamaIndex RAG pipelines on unoblox — free Qwen embeddings plus Claude, GPT-4o, or DeepSeek for answers, billed monthly in rupees with GST.
LlamaIndex (formerly GPT Index) is the lightweight RAG framework — load documents, embed them, query, generate answers. unoblox makes it ₹-native: keep the LlamaIndex code unchanged, swap in cheaper or India-priced models, and bill everything in rupees.
Why LlamaIndex + unoblox
- Minimal: simpler than LangChain for pure RAG workflows
- Model-agnostic: works with any OpenAI-compatible provider, including unoblox
- Free embeddings: Qwen3-Embedding-0.6B is ₹0 on unoblox's native
/v1/embeddings - Fast retrieval: built for large document sets
- ₹ billing: one GST invoice, no international card needed
Two-line setup
from llama_index.core import VectorStoreIndex, SimpleDirectoryReader
from llama_index.llms.openai import OpenAI
from llama_index.embeddings.openai import OpenAIEmbedding
llm = OpenAI(
api_key="ub-gw-YOUR_API_KEY",
api_base="https://api.unoblox.ai/v1",
model="openai/gpt-4o"
)
embeddings = OpenAIEmbedding(
api_key="ub-gw-YOUR_API_KEY",
api_base="https://api.unoblox.ai/v1",
model="Qwen/Qwen3-Embedding-0.6B"
)
documents = SimpleDirectoryReader(input_dir="/path/to/docs").load_data()
index = VectorStoreIndex.from_documents(documents, embed_model=embeddings)
query_engine = index.as_query_engine(llm=llm)
response = query_engine.query("What's the revenue trend?")
Cost breakdown (₹ per 1M tokens)
| Component | Model | ₹ per 1M input | Purpose |
|---|---|---|---|
| Embeddings | Qwen3-Embedding-0.6B | FREE | vectorize documents once |
| LLM answer | Qwen3.8-27B | ₹16.32 | fast retrieval + summary |
| LLM answer | Claude Opus | ₹504 | nuanced, long-form analysis |
Embedding a large document set costs ₹0 up front; each query then costs only the LLM's per-token rate on the (small) retrieved context plus the answer.
Three-model RAG example
# Budget mode: free embeddings, cheap LLM
llm_budget = OpenAI(
api_key="ub-gw-...",
api_base="https://api.unoblox.ai/v1",
model="deepseek-ai/deepseek-v4-flash" # ₹9.07 / ₹18.14
)
# Balanced mode
llm_balanced = OpenAI(
api_key="ub-gw-...",
api_base="https://api.unoblox.ai/v1",
model="openai/gpt-4o" # ₹252 / ₹1008
)
# Premium mode
llm_premium = OpenAI(
api_key="ub-gw-...",
api_base="https://api.unoblox.ai/v1",
model="anthropic/claude-opus-5" # ₹504 / ₹2520
)
documents = SimpleDirectoryReader("/data").load_data()
index = VectorStoreIndex.from_documents(documents)
query = "Analyze Q3 financial performance trends"
engine = index.as_query_engine(llm=llm_premium if "complex" in query.lower() else llm_budget)
answer = engine.query(query)
One unoblox key, three LLMs, cost matched to query difficulty.
Common questions
Does LlamaIndex cache embeddings? Yes. The first run embeds every document once; later queries reuse the stored vectors — no re-embedding, no repeat cost.
Can I use LlamaIndex with reranking?
Yes. unoblox's free Qwen3-Reranker-0.6B works with LlamaIndex's reranking postprocessors via /v1/rerank.
What's the difference between api_base and base_url?
LlamaIndex's OpenAI wrapper uses api_base; the raw OpenAI SDK uses base_url. Both point at https://api.unoblox.ai/v1 — check which class you're instantiating.
Can I stream responses?
Yes. query_engine.aquery() (async) works with unoblox streaming; consume the response incrementally as tokens arrive.
How do I export answers to a file?
LlamaIndex's response object exposes .response (text), .source_nodes (citations), and .metadata (token counts) — serialize whichever fields your logging layer needs.
Is there a Node.js version of LlamaIndex?
Yes, npm install llamaindex. Configuration mirrors the Python client — one base URL, one key, the same model ids.
Get started in rupees → https://unoblox.ai/sign-in
Start building in rupees
Call every major model through one OpenAI-compatible endpoint, billed in ₹ on a GST invoice.