Skip to content
Start freeGPT-5 · Claude · DeepSeek V4 · Qwen3 — in ₹One OpenAI-compatible endpointBilled in rupeesGST invoiceNo international cardGet started →Start freeGPT-5 · Claude · DeepSeek V4 · Qwen3 — in ₹One OpenAI-compatible endpointBilled in rupeesGST invoiceNo international cardGet started →
Guides

DeepSeek-OCR: recommended settings and quickstart

Turn scanned pages, forms and invoices into structured markdown with DeepSeek-OCR on unoblox. Recommended settings, a Python, Node and curl quickstart, and how to read the layout output.

DeepSeek-OCR turns a page image into structured text. You get headings, paragraphs, checkboxes and full HTML tables, plus the position of every block on the page. It is self-hosted in India and is part of the unoblox Freemium tier.

Recommended settings

Use these on every request. They are DeepSeek's documented usage, and we tested them on dense Indian KYC and brokerage forms.

SettingValueWhy
modeldeepseek-ai/deepseek-ocrThe model slug.
Prompt (text part)<|grounding|>Convert the document to markdown.DeepSeek's official document prompt. It also avoids the repetition loops a plain prompt can fall into on forms with rows of empty boxes.
skip_special_tokensfalseKeeps table cells (<td>, <tr>) and the layout tags. Without it, table cells run together.
frequency_penalty0.3Stops the model repeating itself on very dense tables.
temperature0Deterministic output.
max_tokens4000 (up to about 6500)In layout mode a dense page can need 1,500–4,000 output tokens.

Send one page per request, with the image part first and the text part second.

Quickstart — Python

import base64, os
from openai import OpenAI

client = OpenAI(base_url="https://api.unoblox.ai/v1", api_key=os.environ["UNOBLOX_API_KEY"])

with open("page.png", "rb") as f:
    b64 = base64.b64encode(f.read()).decode()

res = client.chat.completions.create(
    model="deepseek-ai/deepseek-ocr",
    temperature=0,
    max_tokens=4000,
    frequency_penalty=0.3,
    extra_body={"skip_special_tokens": False},
    messages=[{"role": "user", "content": [
        {"type": "image_url", "image_url": {"url": f"data:image/png;base64,{b64}"}},
        {"type": "text", "text": "<|grounding|>Convert the document to markdown."},
    ]}],
)
print(res.choices[0].message.content)

Quickstart — Node.js

import OpenAI from "openai";
import { readFileSync } from "node:fs";

const client = new OpenAI({
  baseURL: "https://api.unoblox.ai/v1",
  apiKey: process.env.UNOBLOX_API_KEY,
});

const b64 = readFileSync("page.png").toString("base64");

const res = await client.chat.completions.create({
  model: "deepseek-ai/deepseek-ocr",
  temperature: 0,
  max_tokens: 4000,
  frequency_penalty: 0.3,
  skip_special_tokens: false, // passed through to the model
  messages: [{
    role: "user",
    content: [
      { type: "image_url", image_url: { url: `data:image/png;base64,${b64}` } },
      { type: "text", text: "<|grounding|>Convert the document to markdown." },
    ],
  }],
} as any);

console.log(res.choices[0].message.content);

Quickstart — curl

B64=$(base64 < page.png | tr -d '\n')
curl https://api.unoblox.ai/v1/chat/completions \
  -H "Authorization: Bearer $UNOBLOX_API_KEY" \
  -H "Content-Type: application/json" \
  --data-binary @- <<EOF
{"model":"deepseek-ai/deepseek-ocr","temperature":0,"max_tokens":4000,"frequency_penalty":0.3,"skip_special_tokens":false,
 "messages":[{"role":"user","content":[
   {"type":"image_url","image_url":{"url":"data:image/png;base64,$B64"}},
   {"type":"text","text":"<|grounding|>Convert the document to markdown."}]}]}
EOF

PDFs

DeepSeek-OCR reads images. To process a PDF, convert each page to a PNG at 150 dpi (for example with Poppler: pdftoppm -r 150 -png input.pdf page). Then send one request per page.

Reading the output

In layout mode, each block of the page comes back like this:

<|ref|>sub_title<|/ref|><|det|>[[55, 81, 317, 98]]<|/det|>
## Option form for issue of DIS booklet
  • The label before the box is the block type: title, sub_title, text, table and so on.
  • The four numbers are the block's position on the page (left, top, right, bottom), on a 0–999 grid relative to the page's width and height. Use them to highlight, crop or map form fields.
  • Tables come back as HTML: <table><tr><td>…</td></tr></table>. Most markdown renderers display these as-is.

If you only want the text, strip the layout tags:

import re

def to_plain(text: str) -> str:
    return re.sub(r"<\|ref\|>.*?<\|/ref\|><\|det\|>.*?<\|/det\|>", "", text, flags=re.S).strip()

Limits

  • Context: 8,192 tokens per request. The page image takes roughly 300–950 of those, and the rest is available for output.
  • Very dense tables: occasionally the model repeats a table row until it reaches max_tokens. If finish_reason is length, retry that page at a lower resolution (for example 100 dpi).
  • Inputs: PNG or JPEG images, one page per request. PDFs need converting first (see above).

Pricing and data policy

  • Billed in ₹ at the rate shown on the model page. A typical page costs a fraction of a paisa, and the model is part of the Freemium tier.
  • Your pages are never used for training. Request logging is metadata-only.
deepseekocrguideindia
ShareLinkedInWhatsAppTelegram

Start building in rupees

Call every major model through one OpenAI-compatible endpoint, billed in ₹ on a GST invoice.