AI API Rate Limits Explained (India Guide)
How AI API rate limits (RPM/TPM) work, why they exist, and how to design around them on unoblox, billed in rupees with GST invoicing.
Every production AI integration eventually runs into a rate limit — a 429 instead of a completion. Understanding how limits work, and designing for them up front, saves a scramble later. This guide covers the concepts generically and how they apply on unoblox, where every model sits behind one key billed in rupees.
The two numbers that matter: RPM and TPM
Rate limits are almost always expressed as two separate numbers:
- RPM (requests per minute) — how many individual API calls you can make in a minute, regardless of size.
- TPM (tokens per minute) — how many total tokens (input plus output) you can process in a minute, regardless of how many requests that's split across.
A workload making many small requests can hit the RPM ceiling before the TPM one; a workload making few but very large requests can do the opposite. Knowing which one your traffic pattern is likely to hit first tells you which number to actually watch.
Why rate limits exist at all
Rate limits protect shared infrastructure from being overwhelmed by any single account, and they protect you as a caller from an unbounded, runaway loop (a retry bug, for instance) generating an unexpectedly large bill in a short window. They're a safety mechanism in both directions, not just a restriction.
Free-tier vs paid-tier limits
Exact limits vary by account type and can change as the platform grows, so confirm current numbers in your own dashboard rather than from an article. The general shape is stable: a free or low-commitment tier (like calling the ₹0 Qwen3 1.7B model) generally carries a lower ceiling than a funded, paid account calling premium models, since the limit partly manages shared capacity fairly across all users. If your application is scaling past your current tier, check your account settings rather than treat it as a permanent ceiling.
What happens when you hit a limit
When a request would exceed your current RPM or TPM allowance, the API returns an HTTP 429 instead of a completion. The right response isn't to retry immediately in a tight loop — that just trips the limit again — but to back off: wait briefly, retry, and increase the wait on repeated failures (exponential backoff). Most OpenAI-compatible SDKs support this out of the box.
Designing your application around limits
- Batch where you can. Fewer, larger requests are often more efficient against an RPM ceiling than many tiny ones.
- Queue burst traffic. If your application can experience sudden spikes (a marketing push, a viral moment), a queue that smooths requests out over time is more resilient than firing everything at once.
- Mix models by criticality. Route latency-tolerant background work to one model and keep your rate-limit headroom free for user-facing requests on another.
- Watch for 429s in monitoring, not just errors in production. Treating a 429 as a normal, expected signal — not just an outage — makes it easier to tune your retry logic correctly.
base_url: https://api.unoblox.ai/v1
api_key: ub-gw-xxxxxxxxxxxxxxxxxxxx
model: openai/gpt-5-mini
Frequently asked questions
What is the exact RPM/TPM limit on my unoblox key? Limits are shown in your account dashboard and can vary by plan and by model, so check there for the current figure rather than relying on a fixed number from this guide.
Do rate limits apply per key or per account? Generally per key, though an account can hold multiple keys. Check your dashboard for how limits are structured on your specific account.
Does the free Qwen3 1.7B model have the same limits as paid models? Free and paid tiers are typically structured differently to manage shared capacity fairly — confirm the current figures for the specific model and tier you're using in your dashboard.
How should my code handle a 429 response? Back off and retry with increasing wait times between attempts (exponential backoff), rather than retrying immediately in a loop. Most OpenAI-compatible SDKs support this out of the box.
Can I request a higher rate limit? If your usage is genuinely outgrowing your current tier, that's generally something to raise through your account rather than work around in code — check your dashboard or account settings for how to do that.
Are rate limits the same across every model on unoblox? Not necessarily — different models can carry different limits depending on upstream capacity. Check the specific model and your account dashboard rather than assuming one number applies everywhere.
Get started in rupees → https://unoblox.ai/sign-in
More from unoblox
Developer Guides
How to Get an AI API Key in India
Developer Guides
Migrate from Azure OpenAI to unoblox (India)
Developer Guides
JSON Mode with the AI API in India | unoblox
Developer Guides
Streaming AI API Responses in India | unoblox
Developer Guides
Tool Calling with the AI API in India | unoblox
Developer Guides
AI API for Agencies in India | unoblox
Start building in rupees
Call every major model through one OpenAI-compatible endpoint, billed in ₹ on a GST invoice.