GPU inference hosting in India
Deploy Qwen, Llama, Gemma on unoblox's India-hosted GPUs. Private models, local data residency, no external API dependency.
Host your own LLM in India — same endpoint, full control
Sometimes API routing isn't enough. You need your model on your infrastructure, physically in India—for privacy, for control over vendor lock-in, or for a model or fine-tune that isn't available as a shared endpoint. unoblox lets you deploy open-source LLMs (Qwen, Llama, Gemma, or your own fine-tune) on India-hosted GPUs and reach them through the same ub-gw-… key you already use.
What's hosted now
Production-ready
- Qwen3-1.7B (free, rate-limited)—lightweight, multilingual, good for internal automation
- Qwen3-Embedding-0.6B—semantic search, RAG embeddings
- Qwen3-Reranker-0.6B—rerank search results for quality
On the roadmap
- Llama 4 Maverick and Llama 4 Scout (open-weight reasoning)
- Gemma 4 31B (coding & math)
- Your custom fine-tune (bring your own LoRA)
How deployment works
1. Onboard a model
Request it via Dashboard → Model Deploy, and provide:
- Model name (a HuggingFace repo URL, e.g.,
Qwen/Qwen3-235B-Instruct, or your own fine-tuned weights) - Inference parameters (max sequence length, tensor parallelism, etc.)
- Your target cost per token (you set the margin)
2. We handle the hosting
Load balancing, GPU allocation, and scaling are unoblox's job—you don't manage containers or orchestration.
3. Use it via the standard endpoint
import openai
client = openai.OpenAI(
api_key="ub-gw-YOUR_KEY",
base_url="https://api.unoblox.ai/v1"
)
# Your hosted model, same SDK
response = client.chat.completions.create(
model="custom/my-fine-tuned-qwen",
messages=[{"role": "user", "content": "..."}]
)
How pricing works
| Cost center | Your control | unoblox's role |
|---|---|---|
| GPU allocation | You choose the model and scale | Charged per token, not per GPU-hour |
| Ingress data | Your model, your traffic | Included, no egress charge |
| Logs & audit | 90-day retention | Included |
Example: you deploy a custom fine-tune on reserved GPUs. unoblox agrees a per-token cost with you up front, based on the model size and hardware it needs; you decide what to charge your own users or internal teams on top, and keep the difference. Capacity stays reserved for your workload rather than shared with other customers.
Frequently asked questions
Q: Can I use this for customer-facing features without competing for shared rate limits?
Yes. You get a reserved GPU allocation, so your traffic doesn't compete with other customers' usage.
Q: What if my model gets popular and I need more GPUs?
Request scale-up via Dashboard → Capacity. We provision additional capacity, typically within hours rather than days, and billing updates proportionally.
Q: Can I fine-tune the model inside unoblox?
Not yet—that's on our roadmap. Today: fine-tune externally, upload the weights, and we deploy them for you.
Q: Is data residency guaranteed if I host here?
Yes, for models you deploy this way—the model, its weights, and all inference logs stay in India. The only traffic leaving is your own encrypted API calls in and responses out.
Q: What about version control for model updates?
Each deployment is tagged with a commit SHA. Roll back to a prior version in one click; unoblox keeps recent versions available for that.
Q: Can other customers see my custom model or data?
No. Each deployment is fully isolated—only your own API keys can reach your model.
Get started in rupees → https://unoblox.ai/sign-in
More from unoblox
Developer Guides
AI API Rate Limits Explained (India Guide)
Developer Guides
How to Get an AI API Key in India
Developer Guides
Migrate from Azure OpenAI to unoblox (India)
Developer Guides
JSON Mode with the AI API in India | unoblox
Developer Guides
Streaming AI API Responses in India | unoblox
Developer Guides
Tool Calling with the AI API in India | unoblox
Start building in rupees
Call every major model through one OpenAI-compatible endpoint, billed in ₹ on a GST invoice.