Skip to content

Guides

jev · System One

jev is a System One model: give it something to judge and a set of typed questions, and it returns structured, calibrated answers — a probability, a choice, or a score — instead of prose. It is the decision layer you put around a chat model, not a chatbot itself.

At a glance

FieldDetail
EndpointPOST /v1/systemone
Modeltypesafe/jev
BillingInput tokens only — output tokens are free.
Launch priceFree at launch (zero rupees). Standard input-token pricing and the platform fee are covered in Pricing & credits.
Context64k tokens, text only.

Decisions, not text

jev does not write text or code. It makes decisions.

There is no prose to read back — the output is the verdict: a yes/no probability, a chosen option, or a score, each with a confidence. If you need something written, that's a chat model.

jev is what you put around chat models to route, gate, classify and score — the calls you'd otherwise (expensively) make a big LLM do.

jev vs. a chat model

The request and the response are both shaped differently from a normal completion:

A chat modeljev (System One)
Flowprompt → model → free textstate + typed questions → jev → structured answers + confidence
ShapeOpen-ended. You read the output and hope it followed instructions.Constrained to your schema. Every answer is machine-readable and calibrated.
Best forGenerating prose, code, or a summary a human will read.A plain yes/no, pick-one, or score your code branches on.

The three question types

Name each question and pick a type. Mix as many as you want in one call — they run in parallel, so extra questions cost you almost nothing more.

TypeAnswersUse forExample
noulA yes/no question, answered as a probability.Gates & flags — is it spam? does it need a human? is it on-topic?"Does this express urgency?" → noul: 0.98
choicePick exactly one of the options you define.Routing & classification — which team, which handler, which label.billing / technical / sales → choice: "billing" (0.90)
scoreRate against ordered levels you describe.Severity & sentiment — how urgent, how angry, how good.calm · civil · very angry → score: 1.0 (of 2)

Confidence is a second axis.

The answer tells you what; the confidence tells you whether to trust it. A common pattern: act automatically above a confidence threshold, route to a human below it.

Call jev end to end

One customer message, three questions, one call. Send it to unoblox with your own ub-gw-… key — unoblox handles the upstream.

POST/v1/systemone
curl
# curl
curl -sS https://api.unoblox.ai/v1/systemone \
  -H "Authorization: Bearer $UNOBLOX_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "typesafe/jev",
    "state": "My Stripe payout has failed for 3 days and I am losing sales. Help ASAP.",
    "questions": {
      "urgency":     { "type":"noul",   "instructions":"Does this express urgency?" },
      "department":  { "type":"choice", "instructions":"Which team handles this",
                       "criteria":{ "billing":"Payments", "technical":"Bugs", "sales":"Pricing" } },
      "frustration": { "type":"score",  "instructions":"How frustrated the customer is",
                       "criteria":["Calm","Frustrated but civil","Very angry"] }
    }
  }'
200 · response
{
  "model": "jev-1.13.0",
  "answers": {
    "urgency": {
      "type":"noul", "noul":0.98
    },
    "department": {
      "type":"choice",
      "choice":"billing",
      "confidence":0.90,
      "probabilities":{ "billing":0.90,
        "technical":0.09, "sales":0.01 }
    },
    "frustration": {
      "type":"score",
      "score":1.0, "confidence":0.84
    }
  },
  "usage": { "input_tokens":312, "output_tokens":48 }
}

Your code reads the answer directly.

answers.department.choice === "billing" and answers.urgency.noul > 0.9 and routes — no parsing prose, no "please respond in JSON," no retries on malformed output.

Errors and retries

/v1/systemone fails the same way as every other endpoint — a typed JSON envelope and a stable error_code (see Errors) — plus these systemone-specific codes:

CodeHTTPMessageWhat to do
UB-GW-252500Request decoding failed.Check that the request body is valid JSON and within size limits, then retry.
UB-GW-253503This model isn't approved for systemone inference, or is unavailable.Confirm the model id is a systemone-capable model in good standing, or try again shortly.
UB-GW-254503Native systemone inference is disabled.System One isn't enabled on this deployment yet — contact support.
UB-GW-255400systemone models must be used via /v1/systemone.Call this model through POST /v1/systemone, not /v1/chat/completions.
UB-GW-256400systemone request exceeds 1 MiB.Reduce the combined size of state and questions to under 1 MiB.
UB-GW-257400This request's state and questions exceed the workspace's configured prompt-size limit.Shorten state and questions to fit your workspace's configured prompt-size limit.
UB-GW-258503This request's usage could not be determined and requires manual review.The provider answered without a usage figure, so the result was withheld. Retry the request; if it keeps happening, contact support with the request id.
UB-GW-259503The result is unavailable, though usage was recorded for billing purposes.The provider processed the request but its answer could not be returned; that attempt's input tokens are metered. Retry — a retry is a new request and is metered on its own.
UB-GW-260503This feature is temporarily unavailable. Please try again later.Wait and retry with backoff.
UB-GW-261400The request body must be a JSON object.Send a single JSON object as the request body, not an array or scalar.
UB-GW-262400A non-empty "model" is required.Include a non-empty model field naming the System One model to call.
UB-GW-263400A non-empty "state" is required. / "state" must not be empty.Include a non-empty state field describing what jev should judge.
UB-GW-264400A non-empty "questions" object is required. / "questions" must not be empty.Include a questions object with at least one named question.
UB-GW-265400Each question must be a JSON object. / Each question needs a "type". / Each question's "type" must be one of noul, choice, or score. / Each question needs "instructions".Give every question an object shape with a valid type (noul, choice, or score) and instructions.

How upstream conditions surface: a malformed or oversized request — bad JSON, a missing or empty state/questions, an unrecognized question type — fails fast as one of the 4xx validation codes above; fix the request before retrying. A workspace or key over its call-rate limit gets the platform-wide 429 rate_limit_exceeded (see Errors) with a Retry-After header — wait, then retry. If the model provider itself rejects the request — for example a question whose criteria it cannot accept — its 4xx status is passed through with code UB-GW-037 and a sanitized message; nothing is metered, so fix the request before retrying. If the provider is overloaded or unreachable and no other endpoint can serve the request, you get UB-GW-036 — retry with exponential backoff, as you would for UB-GW-258/259/260. UB-GW-253/254 mean the model or the feature is not currently enabled, so an immediate retry will not help.

retry with backoff
import time, requests

def call_systemone(payload, max_retries=3):
    for attempt in range(max_retries):
        resp = requests.post(
            "https://api.unoblox.ai/v1/systemone",
            headers={"Authorization": f"Bearer {UNOBLOX_KEY}"},
            json=payload,
        )
        if resp.ok:
            return resp.json()
        if resp.status_code == 429:
            time.sleep(float(resp.headers.get("retry-after", "1")))
        elif resp.status_code >= 500:
            time.sleep(2 ** attempt)      # exponential backoff
        else:
            resp.raise_for_status()       # 400 — fix the request, don't retry
    resp.raise_for_status()

What people build with it

Anywhere you're making a big model answer a small, structured question, jev does it faster, cheaper, and with a probability you can threshold on.

Use caseQuestion type(s)What it does
Intent routingchoiceClassify an incoming request and send it to the right handler — deterministic code, a specialist model, or a human.
Guardrails & moderationnoulGate every LLM input, output, or tool call — “is this a jailbreak / off-topic / unsafe?” — at a fraction of an LLM's cost.
Ticket & lead triagechoice + scoreTeam, priority and sentiment in one call, so the queue sorts itself before a person ever looks at it.
LLM output checksnoul“Did the model actually answer the question?” “Is this citation supported?” Catch error modes without a second big call.

jev, or a chat model?

Reach for a chat model when…Reach for jev when…
You need text, code, or a summary written.You need a yes/no, a pick-one, or a score.
The output is open-ended.Your code consumes the answer.
A human will read the result.You want it calibrated, fast, and cheap.
→ /v1/chat/completions→ /v1/systemone

For the chat side of that table, see Quickstart and Routing.

They compose.

Let jev decide and route, then let a chat model generate. A jev call in front of an LLM often removes the need for the LLM call at all.