DeepSeek-OCR: recommended settings and quickstart
Turn scanned pages, forms and invoices into structured markdown with DeepSeek-OCR on unoblox. Recommended settings, a Python, Node and curl quickstart, and how to read the layout output.
DeepSeek-OCR turns a page image into structured text. You get headings, paragraphs, checkboxes and full HTML tables, plus the position of every block on the page. It is self-hosted in India and is part of the unoblox Freemium tier.
Recommended settings
Use these on every request. They are DeepSeek's documented usage, and we tested them on dense Indian KYC and brokerage forms.
| Setting | Value | Why |
|---|---|---|
model | deepseek-ai/deepseek-ocr | The model slug. |
| Prompt (text part) | <|grounding|>Convert the document to markdown. | DeepSeek's official document prompt. It also avoids the repetition loops a plain prompt can fall into on forms with rows of empty boxes. |
skip_special_tokens | false | Keeps table cells (<td>, <tr>) and the layout tags. Without it, table cells run together. |
frequency_penalty | 0.3 | Stops the model repeating itself on very dense tables. |
temperature | 0 | Deterministic output. |
max_tokens | 4000 (up to about 6500) | In layout mode a dense page can need 1,500–4,000 output tokens. |
Send one page per request, with the image part first and the text part second.
Quickstart — Python
import base64, os
from openai import OpenAI
client = OpenAI(base_url="https://api.unoblox.ai/v1", api_key=os.environ["UNOBLOX_API_KEY"])
with open("page.png", "rb") as f:
b64 = base64.b64encode(f.read()).decode()
res = client.chat.completions.create(
model="deepseek-ai/deepseek-ocr",
temperature=0,
max_tokens=4000,
frequency_penalty=0.3,
extra_body={"skip_special_tokens": False},
messages=[{"role": "user", "content": [
{"type": "image_url", "image_url": {"url": f"data:image/png;base64,{b64}"}},
{"type": "text", "text": "<|grounding|>Convert the document to markdown."},
]}],
)
print(res.choices[0].message.content)
Quickstart — Node.js
import OpenAI from "openai";
import { readFileSync } from "node:fs";
const client = new OpenAI({
baseURL: "https://api.unoblox.ai/v1",
apiKey: process.env.UNOBLOX_API_KEY,
});
const b64 = readFileSync("page.png").toString("base64");
const res = await client.chat.completions.create({
model: "deepseek-ai/deepseek-ocr",
temperature: 0,
max_tokens: 4000,
frequency_penalty: 0.3,
skip_special_tokens: false, // passed through to the model
messages: [{
role: "user",
content: [
{ type: "image_url", image_url: { url: `data:image/png;base64,${b64}` } },
{ type: "text", text: "<|grounding|>Convert the document to markdown." },
],
}],
} as any);
console.log(res.choices[0].message.content);
Quickstart — curl
B64=$(base64 < page.png | tr -d '\n')
curl https://api.unoblox.ai/v1/chat/completions \
-H "Authorization: Bearer $UNOBLOX_API_KEY" \
-H "Content-Type: application/json" \
--data-binary @- <<EOF
{"model":"deepseek-ai/deepseek-ocr","temperature":0,"max_tokens":4000,"frequency_penalty":0.3,"skip_special_tokens":false,
"messages":[{"role":"user","content":[
{"type":"image_url","image_url":{"url":"data:image/png;base64,$B64"}},
{"type":"text","text":"<|grounding|>Convert the document to markdown."}]}]}
EOF
PDFs
DeepSeek-OCR reads images. To process a PDF, convert each page to a PNG at 150 dpi (for example with Poppler: pdftoppm -r 150 -png input.pdf page). Then send one request per page.
Reading the output
In layout mode, each block of the page comes back like this:
<|ref|>sub_title<|/ref|><|det|>[[55, 81, 317, 98]]<|/det|>
## Option form for issue of DIS booklet
- The label before the box is the block type:
title,sub_title,text,tableand so on. - The four numbers are the block's position on the page (left, top, right, bottom), on a 0–999 grid relative to the page's width and height. Use them to highlight, crop or map form fields.
- Tables come back as HTML:
<table><tr><td>…</td></tr></table>. Most markdown renderers display these as-is.
If you only want the text, strip the layout tags:
import re
def to_plain(text: str) -> str:
return re.sub(r"<\|ref\|>.*?<\|/ref\|><\|det\|>.*?<\|/det\|>", "", text, flags=re.S).strip()
Limits
- Context: 8,192 tokens per request. The page image takes roughly 300–950 of those, and the rest is available for output.
- Very dense tables: occasionally the model repeats a table row until it reaches
max_tokens. Iffinish_reasonislength, retry that page at a lower resolution (for example 100 dpi). - Inputs: PNG or JPEG images, one page per request. PDFs need converting first (see above).
Pricing and data policy
- Billed in ₹ at the rate shown on the model page. A typical page costs a fraction of a paisa, and the model is part of the Freemium tier.
- Your pages are never used for training. Request logging is metadata-only.
Start building in rupees
Call every major model through one OpenAI-compatible endpoint, billed in ₹ on a GST invoice.