Skip to content
Start freeGPT-5 · Claude · DeepSeek V4 · Qwen3 — in ₹One OpenAI-compatible endpointBilled in rupeesGST invoiceNo international cardGet started →Start freeGPT-5 · Claude · DeepSeek V4 · Qwen3 — in ₹One OpenAI-compatible endpointBilled in rupeesGST invoiceNo international cardGet started →
Developer Guides

GPU inference hosting in India

Deploy Qwen, Llama, Gemma on unoblox's India-hosted GPUs. Private models, local data residency, no external API dependency.

Host your own LLM in India — same endpoint, full control

Sometimes API routing isn't enough. You need your model on your infrastructure, physically in India—for privacy, for control over vendor lock-in, or for a model or fine-tune that isn't available as a shared endpoint. unoblox lets you deploy open-source LLMs (Qwen, Llama, Gemma, or your own fine-tune) on India-hosted GPUs and reach them through the same ub-gw-… key you already use.

What's hosted now

Production-ready

  • Qwen3-1.7B (free, rate-limited)—lightweight, multilingual, good for internal automation
  • Qwen3-Embedding-0.6B—semantic search, RAG embeddings
  • Qwen3-Reranker-0.6B—rerank search results for quality

On the roadmap

  • Llama 4 Maverick and Llama 4 Scout (open-weight reasoning)
  • Gemma 4 31B (coding & math)
  • Your custom fine-tune (bring your own LoRA)

How deployment works

1. Onboard a model

Request it via Dashboard → Model Deploy, and provide:

  • Model name (a HuggingFace repo URL, e.g., Qwen/Qwen3-235B-Instruct, or your own fine-tuned weights)
  • Inference parameters (max sequence length, tensor parallelism, etc.)
  • Your target cost per token (you set the margin)

2. We handle the hosting

Load balancing, GPU allocation, and scaling are unoblox's job—you don't manage containers or orchestration.

3. Use it via the standard endpoint

import openai

client = openai.OpenAI(
    api_key="ub-gw-YOUR_KEY",
    base_url="https://api.unoblox.ai/v1"
)

# Your hosted model, same SDK
response = client.chat.completions.create(
    model="custom/my-fine-tuned-qwen",
    messages=[{"role": "user", "content": "..."}]
)

How pricing works

Cost centerYour controlunoblox's role
GPU allocationYou choose the model and scaleCharged per token, not per GPU-hour
Ingress dataYour model, your trafficIncluded, no egress charge
Logs & audit90-day retentionIncluded

Example: you deploy a custom fine-tune on reserved GPUs. unoblox agrees a per-token cost with you up front, based on the model size and hardware it needs; you decide what to charge your own users or internal teams on top, and keep the difference. Capacity stays reserved for your workload rather than shared with other customers.

Frequently asked questions

Q: Can I use this for customer-facing features without competing for shared rate limits?

Yes. You get a reserved GPU allocation, so your traffic doesn't compete with other customers' usage.

Q: What if my model gets popular and I need more GPUs?

Request scale-up via Dashboard → Capacity. We provision additional capacity, typically within hours rather than days, and billing updates proportionally.

Q: Can I fine-tune the model inside unoblox?

Not yet—that's on our roadmap. Today: fine-tune externally, upload the weights, and we deploy them for you.

Q: Is data residency guaranteed if I host here?

Yes, for models you deploy this way—the model, its weights, and all inference logs stay in India. The only traffic leaving is your own encrypted API calls in and responses out.

Q: What about version control for model updates?

Each deployment is tagged with a commit SHA. Roll back to a prior version in one click; unoblox keeps recent versions available for that.

Q: Can other customers see my custom model or data?

No. Each deployment is fully isolated—only your own API keys can reach your model.

Get started in rupees → https://unoblox.ai/sign-in

GPU-inferencedeploymentdata-residencyguides
ShareLinkedInWhatsAppTelegram

Start building in rupees

Call every major model through one OpenAI-compatible endpoint, billed in ₹ on a GST invoice.