Qwen3 235B API in India — Open-Weight Pricing
Qwen3 235B, Alibaba's open-weight flagship, on unoblox at ₹9.07 input / ₹55.44 output per 1M tokens. Fine-tune it yourself or run it managed.
Qwen3 235B is the open-weight reasoning workhorse from Alibaba's Qwen team. At ₹9.07 per million input tokens, it costs a fraction of comparable closed models while holding its own on reasoning and coding. Download it, fine-tune it, or call it through unoblox — your choice.
Why Qwen3 235B
Qwen3 235B is a large mixture-of-experts model built for reasoning, coding, multilingual tasks, and long-context analysis, with a genuine 256K-token context window. Unlike closed APIs, the weights are open — you can download and run it on your own hardware if you need full control.
Pricing: Managed vs. Self-Hosted
| Option | Setup | Cost | Control |
|---|---|---|---|
| Managed (unoblox) | Minutes | ₹9.07 in / ₹55.44 out per 1M tokens | Instant scaling, no ops |
| Self-hosted GPU | Hours to days | Hardware cost only | Full control, needs a multi-GPU box |
Qwen3 235B is a large model — self-hosting means provisioning multiple high-memory GPUs, so most teams start managed and self-host only once volume justifies the hardware.
When to Choose Managed vs. Self-Hosted
Use unoblox if: you want instant scaling, no GPU management, monthly GST billing, and predictable costs.
Self-host if: you have sensitive data that must stay on your own infrastructure, need sub-100ms local inference, or plan to fine-tune heavily and serve continuously.
How to Use Qwen3 235B on unoblox
- Create an account → https://unoblox.ai/sign-in.
- Copy your API key → "API Keys" section.
- Point your SDK →
base_url = https://api.unoblox.ai/v1,model = qwen/qwen3-235b-a22b-instruct-2507.
Python quickstart:
from openai import OpenAI
client = OpenAI(api_key="ub-gw-...", base_url="https://api.unoblox.ai/v1")
response = client.chat.completions.create(
model="qwen/qwen3-235b-a22b-instruct-2507",
messages=[{"role": "user", "content": "Generate a novel algorithm for..."}],
max_tokens=2000
)
print(response.choices[0].message.content)
Fine-Tuning & Customization
Qwen3 235B is openly licensed. Download it from Hugging Face, fine-tune on your own data (medical documents, legal contracts, internal knowledge), then serve it locally or fall back to unoblox's managed tier for burst capacity.
Frequently asked questions
Q: Can I download Qwen3 235B? A: Yes — it's open-weight, downloadable from Hugging Face with no registration needed.
Q: What's the latency like on unoblox? A: Streaming responses start quickly and tokens flow continuously; exact speed varies with prompt size, output length, and load.
Q: Is 235B better than Qwen3 Max for reasoning? A: They're both strong. Max edges ahead on the hardest reasoning tasks; for most production workloads, 235B is more than capable and costs over 13x less on input.
Q: Can I A/B test 235B against other open-weight models?
A: Yes — switch the model id in one line to compare against Llama or Mistral, both also live on unoblox.
Q: Do you offer discounts for high-volume usage? A: Yes — contact sales for ₹10L+/month commitments to discuss custom rates.
Get started in rupees → https://unoblox.ai/sign-in
Start building in rupees
Call every major model through one OpenAI-compatible endpoint, billed in ₹ on a GST invoice.