Skip to content
Start freeGPT-5 · Claude · DeepSeek V4 · Qwen3 — in ₹One OpenAI-compatible endpointBilled in rupeesGST invoiceNo international cardGet started →Start freeGPT-5 · Claude · DeepSeek V4 · Qwen3 — in ₹One OpenAI-compatible endpointBilled in rupeesGST invoiceNo international cardGet started →
Model API Guides

Qwen3 235B API in India — Open-Weight Pricing

Qwen3 235B, Alibaba's open-weight flagship, on unoblox at ₹9.07 input / ₹55.44 output per 1M tokens. Fine-tune it yourself or run it managed.

Qwen3 235B is the open-weight reasoning workhorse from Alibaba's Qwen team. At ₹9.07 per million input tokens, it costs a fraction of comparable closed models while holding its own on reasoning and coding. Download it, fine-tune it, or call it through unoblox — your choice.

Why Qwen3 235B

Qwen3 235B is a large mixture-of-experts model built for reasoning, coding, multilingual tasks, and long-context analysis, with a genuine 256K-token context window. Unlike closed APIs, the weights are open — you can download and run it on your own hardware if you need full control.

Pricing: Managed vs. Self-Hosted

OptionSetupCostControl
Managed (unoblox)Minutes₹9.07 in / ₹55.44 out per 1M tokensInstant scaling, no ops
Self-hosted GPUHours to daysHardware cost onlyFull control, needs a multi-GPU box

Qwen3 235B is a large model — self-hosting means provisioning multiple high-memory GPUs, so most teams start managed and self-host only once volume justifies the hardware.

When to Choose Managed vs. Self-Hosted

Use unoblox if: you want instant scaling, no GPU management, monthly GST billing, and predictable costs.

Self-host if: you have sensitive data that must stay on your own infrastructure, need sub-100ms local inference, or plan to fine-tune heavily and serve continuously.

How to Use Qwen3 235B on unoblox

  1. Create an account → https://unoblox.ai/sign-in.
  2. Copy your API key → "API Keys" section.
  3. Point your SDK → base_url = https://api.unoblox.ai/v1, model = qwen/qwen3-235b-a22b-instruct-2507.

Python quickstart:

from openai import OpenAI
client = OpenAI(api_key="ub-gw-...", base_url="https://api.unoblox.ai/v1")
response = client.chat.completions.create(
    model="qwen/qwen3-235b-a22b-instruct-2507",
    messages=[{"role": "user", "content": "Generate a novel algorithm for..."}],
    max_tokens=2000
)
print(response.choices[0].message.content)

Fine-Tuning & Customization

Qwen3 235B is openly licensed. Download it from Hugging Face, fine-tune on your own data (medical documents, legal contracts, internal knowledge), then serve it locally or fall back to unoblox's managed tier for burst capacity.

Frequently asked questions

Q: Can I download Qwen3 235B? A: Yes — it's open-weight, downloadable from Hugging Face with no registration needed.

Q: What's the latency like on unoblox? A: Streaming responses start quickly and tokens flow continuously; exact speed varies with prompt size, output length, and load.

Q: Is 235B better than Qwen3 Max for reasoning? A: They're both strong. Max edges ahead on the hardest reasoning tasks; for most production workloads, 235B is more than capable and costs over 13x less on input.

Q: Can I A/B test 235B against other open-weight models? A: Yes — switch the model id in one line to compare against Llama or Mistral, both also live on unoblox.

Q: Do you offer discounts for high-volume usage? A: Yes — contact sales for ₹10L+/month commitments to discuss custom rates.

Get started in rupees → https://unoblox.ai/sign-in

qwenllm
ShareLinkedInWhatsAppTelegram

Start building in rupees

Call every major model through one OpenAI-compatible endpoint, billed in ₹ on a GST invoice.