Skip to content

Guides

Use unoblox inside Claude Code

Claude Code can use a catalog model through unoblox’s Anthropic-compatible Messages endpoint. Set the gateway origin, authenticate with your unoblox key, and pin a model returned by the live catalog.

Configure Claude Code

Set ANTHROPIC_BASE_URL to the public gateway origin without /v1. Claude Code appends the Messages path. Use either ANTHROPIC_API_KEY (the x-api-key form) or ANTHROPIC_AUTH_TOKEN (the Authorization: Bearer form). The gateway accepts anthropic-version and does not use it to change request behavior.

x-api-key authentication
export ANTHROPIC_BASE_URL="https://api.unoblox.ai"
export ANTHROPIC_API_KEY="$UNOBLOX_API_KEY"
export ANTHROPIC_MODEL="<model-id>"

# Pin the catalog model explicitly for a non-interactive run.
claude -p --model "$ANTHROPIC_MODEL" "Summarize this repository."
Bearer authentication
export ANTHROPIC_BASE_URL="https://api.unoblox.ai"
export ANTHROPIC_AUTH_TOKEN="$UNOBLOX_API_KEY"
export ANTHROPIC_MODEL="<model-id>"

claude -p --model "$ANTHROPIC_MODEL" "Explain the selected files."

Pin a catalog model

Claude Code commonly sends a claude-* model name. unoblox does not map that name to a provider automatically. Set ANTHROPIC_MODEL or pass --model with an exact unoblox catalog id. Find one with GET /v1/models; see the Models guide for the catalog response shape.

A claude-* model name returns UB-GW-242:

400 response message
'<model>' is an Anthropic model name, not a unoblox catalog model id. To use unoblox from Claude Code, pin a unoblox catalog model id: set `--model <unoblox-model-id>` on the CLI, or export `ANTHROPIC_MODEL=<unoblox-model-id>` (see GET /v1/models for available ids).

Pin an unoblox catalog id before sending the request. The literal model value is shown in place of <model>.

Thinking and tool use

Function tool calls are translated to Anthropic tool_use blocks and tool results are accepted on the next Messages request. Streaming preserves the Anthropic event sequence for text and function-tool blocks.

Thinking is off when omitted, and { type: "disabled" } keeps it off. To request reasoning with a fixed budget, send { type: "enabled", budget_tokens: <n> } with budget_tokens at least 1024 and less than max_tokens. Recent Claude Code releases default to { type: "adaptive" } instead — no budget to set; the model decides whether and how much to think. Both forms route only to reasoning-capable candidates and both are represented the same way: returned reasoning as Anthropic thinking blocks, non-stream and streamed. output_config.effort (below) is accepted alongside either form but does not yet steer thinking depth.

If no routed candidate supports reasoning, either form returns UB-GW-244:

400 response message
The routed model doesn't support extended thinking. Remove 'thinking' from the request, or pin a reasoning-capable model.

Limits and unsupported fields

SurfaceBehavior
POST /v1/messagesRequires model, max_tokens, and a non-empty messages array.
POST /v1/messages/count_tokensReturns an input-token estimate; it has no billing effect.
ToolsFunction tools and tool-result follow-ups are supported.
ThinkingOmitted or disabled keeps reasoning off. Enabled or adaptive thinking requires a reasoning-capable routed candidate.
speedNot supported. Any presence returns UB-GW-045; it is not a documented Anthropic Messages field.

Count-token request-shape failures return UB-GW-243 with one of these literal messages. See the API reference for the full Messages request shape and response events.

UB-GW-243 400 response messages
The request body is not valid JSON.
The request body must be a JSON object.
The 'model' field is required.
'messages' must be a non-empty array.

speed is the one field this endpoint still refuses outright:

400 response message
This field isn't supported on this endpoint and can't be silently ignored: 'speed'.

Accepted but not applied

These five fields are real, documented Anthropic Messages fields. The gateway accepts a well-formed value — it is never forwarded to the upstream model provider under its own name, and it never causes a 400 by itself — but none of their real semantics run against a self-hosted SGLang/vLLM-style endpoint today. Every request that used one or more of them carries a comma-separated x-unoblox-unsupported-fields response header naming exactly which ones, so this is never a silent drop.

FieldAcceptedApplied?Effect
context_managementYesNoServer-side context editing (clearing old thinking/tool-use blocks) is not executed; the conversation is sent to the provider unedited.
metadataYesNometadata.user_id is accepted but never forwarded to the provider or used for abuse detection on this endpoint.
output_configYesNoeffort and format (structured-output JSON Schema) are accepted but do not steer generation or constrain the response shape.
containerYesNoCode-execution container reuse is not implemented; the value is accepted and ignored.
service_tierYesNoPriority-tier routing is not implemented; every request routes through the same catalog selection regardless of this value.

A value that does not match the documented shape for one of these fields (for example service_tier: "default", which is neither "auto" nor "standard_only") returns UB-GW-247 rather than being silently accepted.