Skip to content
Start freeGPT-5 · Claude · DeepSeek V4 · Qwen3 — in ₹One OpenAI-compatible endpointBilled in rupeesGST invoiceNo international cardGet started →Start freeGPT-5 · Claude · DeepSeek V4 · Qwen3 — in ₹One OpenAI-compatible endpointBilled in rupeesGST invoiceNo international cardGet started →
API Pricing Guides

Llama pricing in India (₹)

Meta Llama 4 Maverick (₹20.16/M input) & Scout (₹10.08/M) on unoblox — open-weight, fast, billed straight in rupees with a GST invoice.

Llama pricing in India — fast, open-weight, ₹-native

Meta's Llama 4 models are priced on unoblox in rupees: Llama 4 Maverick and Llama 4 Scout cost a fraction of Claude or GPT, with the same monthly GST invoice and rupee-only billing as every other model on the platform.

Llama 4 pricing (per 1M tokens)

ModelInput (₹)Output (₹)Typical latencyStrengths
Llama 4 Maverick₹20.16₹80.64FastReasoning, code, long-context
Llama 4 Scout₹10.08₹30.24FastestClassification, simple generation

Why Llama is worth a look before defaulting to GPT or Claude

Llama 4 is open-weight and heavily optimized for inference speed, which is why unoblox can price it well below proprietary models:

  • Maverick at ₹20.16/M input is 10× cheaper than Claude Sonnet's ₹201.6/M, and 25× cheaper than Claude Opus's ₹504/M — while remaining strong on reasoning and code.
  • Scout at ₹10.08/M input is the cheapest quality option on the platform outside the free tier, built for high-volume, well-defined tasks.

Worked example — 40M input + 10M output tokens a month:

  • Llama Scout: (40 × ₹10.08) + (10 × ₹30.24) = ₹706
  • Claude Sonnet: (40 × ₹201.6) + (10 × ₹1,008) = ₹18,144

Same volume, roughly 26× apart. The output quality gap for structured tasks (classification, tagging, templated generation) is small enough that most teams don't notice it in production.

Use cases where Llama is the practical default

TaskBest modelWhy
Document summarizationMaverickLong-context, cost-efficient
Ticket classificationScoutFast, cheap, accurate enough
Code generationMaverickStrong reasoning, reliable output
Sentiment analysisScoutSub-second, near-zero marginal cost
Conversational agentsMaverickCoherent multi-turn responses

Cost breakdown: Llama vs alternatives

Monthly spend at 30M input tokens + 10M output tokens:

ModelMonthly cost
Llama 4 Scout₹605
Qwen3 235B-A22B₹826
Claude Sonnet₹16,128
GPT-4o₹17,640

At this volume, Llama Scout runs roughly 27–29× cheaper than Claude Sonnet or GPT-4o for workloads where structured, high-volume output matters more than nuanced reasoning.

Calling Llama through unoblox

from openai import OpenAI
client = OpenAI(api_key="ub-gw-YOUR_KEY", base_url="https://api.unoblox.ai/v1")
response = client.chat.completions.create(
    model="meta-llama/llama-4-scout-17b-16e-instruct",
    messages=[{"role": "user", "content": "Hello"}]
)

Same monthly rupee invoice, same claimable GST input credit, no separate account needed.

Frequently asked questions

Is open-weight Llama as good as proprietary models? For classification, summarization, and templated generation, the practical gap is small at a fraction of the cost. For open-ended research or edge-case reasoning, proprietary models still have an edge.

How does Llama's latency compare to Claude or GPT? Llama 4 is generally fast, since it's optimized for efficient serving. Proprietary models can be more consistent under heavy multi-step reasoning, but for typical chat and classification workloads the difference is rarely noticeable.

Can I run Llama in production, not just prototypes? Yes — Llama 4 is stable and widely deployed; unoblox serves it on the same infrastructure and SLA as every other model.

Maverick or Scout — which one? Scout for high-volume, simple tasks where cost matters most. Maverick when you need longer context or stronger reasoning and can absorb roughly double the cost.

Do I need my own GPU capacity to use Llama? No. unoblox hosts inference; you call a standard API endpoint.

Is there a free option before I commit to Llama pricing? Qwen3 1.7B is ₹0 for prototyping. Llama Scout at ₹10.08/M input is the cheapest paid step up once you need more consistent quality.

Get started in rupees → https://unoblox.ai/sign-in

llama pricing indiacost
ShareLinkedInWhatsAppTelegram

Start building in rupees

Call every major model through one OpenAI-compatible endpoint, billed in ₹ on a GST invoice.