Skip to content
Start freeGPT-5 · Claude · DeepSeek V4 · Qwen3 — in ₹One OpenAI-compatible endpointBilled in rupeesGST invoiceNo international cardGet started →Start freeGPT-5 · Claude · DeepSeek V4 · Qwen3 — in ₹One OpenAI-compatible endpointBilled in rupeesGST invoiceNo international cardGet started →
Use Cases

Document & Invoice Extraction API

Extract structured data from PDFs, invoices, forms using vision models. One API, rupee billing, no monthly SaaS fees.

Extract Data From Documents & Invoices

Turn PDFs, scans, and photos into structured JSON—invoice numbers, amounts, dates, signatures—without custom OCR pipelines. Use unoblox's vision-capable models, billed in rupees.

Why vision over traditional OCR

Traditional OCR struggles with layouts, handwriting, and context. Vision models (Claude, Qwen) understand structure, read tables, and extract meaning. Faster iteration, fewer exceptions.

Models for document extraction

ModelCost (input/output)Best forContext
Claude Opus₹504 / ₹2520High-accuracy legal, complex layouts200K+ tokens
Claude Sonnet₹201.6 / ₹1008Fast extraction, most invoices200K+ tokens
Qwen3.8-27B₹16.32 / ₹48.96Budget, open-weight262K tokens
GPT-5₹126 / ₹1008Reasoning over ambiguous docs400K tokens

3-step extraction workflow

Step 1: Upload document as base64

POST https://api.unoblox.ai/v1/chat/completions
Authorization: Bearer ub-gw-...

{
  "model": "anthropic/claude-opus-4-5",
  "messages": [
    {"role": "user", "content": [
      {"type": "text", "text": "Extract invoice number, date, amount, vendor name as JSON."},
      {"type": "image_url", "image_url": {"url": "data:image/png;base64,...base64..."}}
    ]}
  ]
}

Step 2: Parse JSON response

{"invoice_number": "INV-2026-001", "date": "2026-09-19", "amount": "₹50,000", "vendor": "Acme Inc."}

Step 3: Validate and store — Schema-validate the JSON, retry on parse errors, store in your database.

Real-world pricing for 1,000 invoices

  • Claude Sonnet: ~₹200–300 (depending on document length)
  • Qwen3.8-27B: ~₹50–80
  • No SaaS fees, no per-document charges, just token usage

Frequently asked questions

Q: How do I handle handwritten or poor-quality scans? A: Claude Opus excels at degraded images; Qwen3 is faster but less tolerant. Test both on your worst-case docs; fallback to Opus for edge cases.

Q: Can I extract tables from multi-page PDFs? A: Yes, if the PDF fits into token limits (~200k for Claude models). For very long documents, split into pages and batch-process.

Q: What about GSTIN, PAN, or bank account validation? A: The API extracts raw text; add post-processing to validate checksums (GSTIN, IFSC) or regex patterns.

Q: Do you store the documents I upload? A: No. unoblox and upstream providers process images stateless; no retention. See compliance docs for DPDP/GDPR details.

Q: Which format should I send—PDF or image? A: Turn PDFs into PNG or JPG first (using PyPDF2 or ImageMagick), then send as a base64 image. Simpler and more reliable than raw PDF parsing.

Q: Can I extract from handwritten forms? A: Yes, Claude Opus reads cursive well. Qwen3 is faster but less reliable on cursive. Test on a sample.

Get started in rupees → https://unoblox.ai/sign-in

document extractioninvoiceocruse-cases
ShareLinkedInWhatsAppTelegram

Start building in rupees

Call every major model through one OpenAI-compatible endpoint, billed in ₹ on a GST invoice.