Skip to content

Provider guide

Bring your SGLang or vLLM endpoint

Register a public OpenAI-compatible endpoint for a catalog model. Admission is evidence-based: the registry validates its network target, probes it, and keeps production routing and payout readiness separate.

What to expose

Use a public HTTPS origin that serves OpenAI Chat Completions. You can register either a base such as https://inference.example.com/v1 or a full operation URL; the registry derives /v1/chat/completions from a base for a text endpoint. It may also try /v1/models to resolve your upstream model identifier, but a failed or unsupported model-list response does not by itself admit or reject an endpoint.

The endpoint must be reachable from the registry. Loopback, private, link-local, cloud-metadata and other non-globally-routable addresses are refused, including DNS results that contain one of those addresses. The accepted routable IP is pinned between validation and a routed call, which blocks DNS-rebinding redirects. The health scheduler periodically revalidates DNS and may update that pin only to a newly resolved globally-routable address.

Do not register an internal address.

The browser check is advisory. Server-side SSRF validation and DNS pinning are the authoritative gate.

Start a compatible server

SGLang and vLLM can both expose an OpenAI-compatible Chat Completions surface. Bind the server behind a public HTTPS reverse proxy, verify that the configured model can produce a normal visible assistant response, then register the proxy URL from the provider console.

vLLM — illustrative local process
vllm serve <your-model> \
  --host 0.0.0.0 --port 8000

# Register the public HTTPS origin, for example:
# https://inference.example.com/v1
SGLang — illustrative local process
python -m sglang.launch_server \
  --model-path <your-model> --host 0.0.0.0 --port 30000

# Put a public HTTPS reverse proxy in front, then register its /v1 origin.

If your upstream needs credentials, enter the complete authorization value in the write-only field. It is used for admission, health probes and routed inference; it is not returned by the API. Declare only the request parameters your endpoint actually supports.

Admission and health

Registration performs a correctness probe and a bounded latency/load burst through the pinned transport. For a text model, the correctness probe asks for one short visible word. A reasoning-only response is not a successful text admission response. A failed check does not create an admitted endpoint.

After registration, the health scheduler continues to probe the endpoint. Check the provider console's Health page for the reported state and counters. A serviceable health state alone is not a promise of production traffic: routing also applies provider eligibility, model policy, price, capacity and request requirements.

Compatibility boundaries

The current onboarding contract is OpenAI-compatible Chat Completions plus the supported-parameter declaration. It does not register or forward chat_template_kwargs, and it does not expose an enable_thinking onboarding switch. Do not depend on either option being accepted or propagated by this path.

Reasoning capability is only appropriate when the serving endpoint and its declared capabilities support it. This guide does not promise thinking-block translation, template controls, or any provider-specific extension beyond the currently admitted OpenAI-compatible surface.

Pricing, KYC and payouts

Enter per-token endpoint pricing during registration. Pricing is USD per token for the OpenRouter-shaped prompt/completion keys (USD per unit for image, request, cache_read and cache_write, where they apply) — the same unit the catalog listing uses. Eligible served usage can appear in provider earnings, but a displayed balance is not a transfer confirmation. Production eligibility, KYC, a verified payout rail, contracts, tax treatment and payout workflow readiness are evaluated separately.

Submit and review verification from the provider console, then use the Payoutspage for reported earnings and payout-worklist evidence. This documentation does not promise payout timing, settlement, tax certificates, or bank receipt.