Google Gemma API in India
Gemma 4 31B in India: ₹13.10 input, ₹38.30 output per million tokens. Open-weight efficiency from Google, billed in rupees, GST.
Google's Gemma 4 31B is a compact, open-weight model built for efficiency rather than frontier scale. On unoblox it's priced at ₹13.10 per million input tokens and ₹38.30 per million output tokens — billed monthly in rupees, no international card, no licensing negotiation with Google.
Pricing and model profile
Per 1M tokens:
- Input: ₹13.10
- Output: ₹38.30
- Parameters: 31B (open weights)
Gemma sits in the "small but capable" category — Google designed it to be light enough to self-host on modest hardware while staying strong on instruction-following, which is why it shows up so often behind chatbots and classification pipelines rather than research assistants.
Ideal use cases
Chatbots and conversational AI — Gemma's instruction-tuning is solid, and its low per-token cost makes it comfortable for high-volume customer interactions.
Content summarization — good at pulling key points from long documents, articles, and email threads.
Classification and tagging — document classification, intent detection, and topic labeling all play to its strengths.
Semantic search and filtering — rank or filter search results in e-commerce or internal knowledge bases.
Batch processing — open weights carry no licensing fee, so running Gemma across large batches stays predictable on cost.
Quick integration
from openai import OpenAI
client = OpenAI(api_key="ub-gw-...", base_url="https://api.unoblox.ai/v1")
response = client.chat.completions.create(
model="google/gemma-4-31b",
messages=[{"role": "user", "content": "Summarize this article..."}],
)
Python, Node.js, and Go all work through the same OpenAI-compatible client; streaming is enabled by default.
Gemma for Indian startups
Lean deployment — 31B parameters means it's realistic to self-host on modest hardware if you ever want to leave the managed API.
Cost predictability — at ₹13.10 per million input tokens, forecasting a year of heavy usage is straightforward, with no surprise line item on the invoice.
Google's backing — Gemma is maintained by Google, so the weights and safety tuning come from the same team behind Gemini.
Frequently asked questions
Is Gemma as capable as Claude or GPT? Not on deep, multi-step reasoning — but for classification, summarization, and instruction-following it's competitive, and it costs roughly 15–19x less than GPT-4o or Claude Sonnet per token.
Can I fine-tune Gemma through unoblox? Fine-tuning isn't exposed through the API today. Since the weights are open, you can fine-tune it yourself and self-host if you need a custom version.
Does Gemma support vision? This version is text-only. For image or document vision, use GPT-4o or a Claude model on the same key.
How well does Gemma handle Hindi and other Indian languages? Multilingual support is decent, including Hindi, though Qwen's models are generally stronger on Asian-language coverage if that's your primary use case.
Can I deploy Gemma locally instead of using the API? Yes — it's open-source. Download the weights and run them on your own GPU or CPU; unoblox is there for when you'd rather not run the infrastructure yourself.
Get started in rupees → https://unoblox.ai/sign-in
Start building in rupees
Call every major model through one OpenAI-compatible endpoint, billed in ₹ on a GST invoice.