NVIDIA Nemotron 3.5 Lightning API in India
Nemotron 3.5 Lightning: ₹5.78 input, ₹15.41 output per 1M tokens in India. NVIDIA's lightweight open model for fast, cheap agent work.
Nemotron 3.5 Lightning is NVIDIA's compact open model for fast, inexpensive agent and assistant work, live on unoblox as nvidia/nemotron-3.5-lightning. It is among the lowest-priced models in our catalogue and bills in rupees.
What Nemotron 3.5 Lightning is
NVIDIA released Nemotron 3.5 Lightning on 11 August 2026. It has about 30 billion total parameters with 3 billion active, in a hybrid design that interleaves Mamba-2 and mixture-of-experts layers with a few attention layers. NVIDIA published it under the OpenMDW-1.1 licence with weights, training data and recipes. unoblox lists a context length of 262,144 tokens, with tool calling and text in, text out.
Where it fits
- Routing and classification at very low cost per call.
- Agent sub-steps like tool selection and argument filling.
- Prototyping where you want to iterate cheaply before moving a workload to a larger model.
Pick Lightning for speed and price. When quality falls short, move the hard cases to Nemotron 3 Ultra in the catalogue or another larger model. Because the endpoint is the same, escalation is a model-string change.
₹ pricing for Nemotron 3.5 Lightning
On unoblox, Nemotron 3.5 Lightning costs ₹5.78 per million input tokens and ₹15.41 per million output tokens. That displayed price is what you are billed per token; nothing is added on top. Cached input tokens are listed at ₹2.89 per million. The model's context length on unoblox is 262,144 tokens.
| Model | Input (₹ / 1M tokens) | Output (₹ / 1M tokens) |
|---|---|---|
| Nemotron 3.5 Lightning | ₹5.78 | ₹15.41 |
| Nemotron 3 Ultra 550B | ₹48.14 | ₹211.82 |
| gpt-oss 20B | ₹2.89 | ₹13.48 |
| Qwen3.6 35B A3B | ₹9.63 | ₹91.47 |
As a worked example, a month with 20 million input tokens and 4 million output tokens comes to about ₹177.24 on Nemotron 3.5 Lightning (₹115.60 for input plus ₹61.64 for output). Live rates are always on the model page, and each request is rounded up to the nearest paisa.
Calling Nemotron 3.5 Lightning from the unoblox endpoint
unoblox is OpenAI-compatible, so any OpenAI SDK works by changing the base URL and key.
curl https://api.unoblox.ai/v1/chat/completions \
-H "Authorization: Bearer YOUR_UNOBLOX_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "nvidia/nemotron-3.5-lightning",
"messages": [{"role": "user", "content": "Summarise this support ticket in two lines."}],
"max_tokens": 512
}'
from openai import OpenAI
client = OpenAI(
api_key="YOUR_UNOBLOX_API_KEY",
base_url="https://api.unoblox.ai/v1",
)
resp = client.chat.completions.create(
model="nvidia/nemotron-3.5-lightning",
messages=[{"role": "user", "content": "Summarise this support ticket in two lines."}],
max_tokens=512,
)
print(resp.choices[0].message.content)
Get a key at https://unoblox.ai/sign-in and top up your wallet in rupees.
Notes on parameters, billing and data
Use max_tokens for output limits. The model runs at an external provider and not in India, so we make no data-residency claim. The catalogue does not list a reasoning flag for it. unoblox provides rupee pricing, GST invoicing and one endpoint.
Frequently asked questions
Is Nemotron 3.5 Lightning available in India?
Yes, as nvidia/nemotron-3.5-lightning on https://api.unoblox.ai/v1.
How big is it? About 30 billion parameters in total with 3 billion active, per NVIDIA.
What context length is listed? 262,144 tokens.
Does it support tool calling? Yes, the catalogue lists tools and tool_choice.
Is it open? NVIDIA released it with weights, data and recipes under OpenMDW-1.1.
Does it run in India? No claim is made; the model runs at its provider.
Get started in rupees → https://unoblox.ai/sign-in
Start building in rupees
Call every major model through one OpenAI-compatible endpoint, billed in ₹ on a GST invoice.