Node.js AI API quickstart (₹ pricing, India)
Call LLMs from Node.js via OpenAI SDK or fetch. One endpoint, one key from unoblox, billed in ₹ monthly with GST.
Node.js + unoblox: LLM calls billed in rupees
Node.js developers can call unoblox's /v1 endpoint with the official OpenAI SDK or plain fetch — no custom client needed. Swap the base URL and key, deploy to Vercel, Lambda, or your own server, and get one GST-compliant ₹ invoice every month.
Get started in 2 minutes
npm install openai
import OpenAI from "openai";
const openai = new OpenAI({
baseURL: "https://api.unoblox.ai/v1",
apiKey: process.env.UNOBLOX_API_KEY // ub-gw-... from unoblox.ai/sign-in
});
const response = await openai.chat.completions.create({
model: "qwen/qwen3-235b-a22b-instruct-2507",
messages: [
{ role: "user", content: "Explain GST in 50 words" }
]
});
console.log(response.choices[0].message.content);
Streaming (real-time tokens)
const stream = await openai.chat.completions.create({
model: "qwen/qwen3-235b-a22b-instruct-2507",
messages: [{ role: "user", content: "..." }],
stream: true
});
for await (const chunk of stream) {
process.stdout.write(chunk.choices[0]?.delta?.content || "");
}
Handling errors and retries
import OpenAI, { APIError } from "openai";
try {
const res = await openai.chat.completions.create({ model: "openai/gpt-4o", messages });
} catch (err) {
if (err instanceof APIError && err.status === 429) {
// back off and retry — unoblox returns standard OpenAI error shapes
}
throw err;
}
Pricing per 1M tokens (input / output)
| Model | ₹ cost | Best for |
|---|---|---|
| Qwen3.8-27B | ₹16.32 / ₹48.96 | fast, vision |
| DeepSeek V3.2 | ₹26.21 / ₹38.30 | deep reasoning |
| Qwen3 1.7B | ₹0 / ₹0 | prototyping, rate-limited |
| Claude Sonnet | ₹201.6 / ₹1008 | production quality |
Deployment notes
- Vercel Edge: works with
runtime: "edge"— keep the OpenAI client's defaultfetchtransport - AWS Lambda: pass the key via environment variables; set the function timeout to at least 30s for longer generations
- Express.js: a
POST /api/chatroute that forwards the body to unoblox and pipes back JSON or a stream - Discord or Slack bots: run on your own server, call unoblox per message, log tokens for cost tracking
FAQ
Do I need to buffer the whole response before returning it? No. Forward the stream directly to your HTTP response as chunks arrive — Node's stream APIs handle the backpressure.
Can I use TypeScript types?
Yes. The OpenAI SDK ships full types — ChatCompletion, ChatCompletionMessageParam, and so on — with editor autocomplete out of the box.
How do I retry failed requests?
Catch APIError, check err.status, and retry with exponential backoff. The SDK doesn't auto-retry by default.
Can I run this on Cloudflare Workers?
Yes — plain fetch against https://api.unoblox.ai/v1/chat/completions is the safest option in Workers' restricted runtime.
What's a reasonable timeout? Most unoblox responses complete well under 30s; configure your HTTP client's timeout a little above that for longer generations.
Can I run multiple models in parallel?
Yes. Fire several create() calls with Promise.all() — each is billed separately, by its own model's ₹ rate.
Get started in rupees → https://unoblox.ai/sign-in
Start building in rupees
Call every major model through one OpenAI-compatible endpoint, billed in ₹ on a GST invoice.