Guides
jev · System One
jev is a System One model: give it something to judge and a set of typed questions, and it returns structured, calibrated answers — a probability, a choice, or a score — instead of prose. It is the decision layer you put around a chat model, not a chatbot itself.
At a glance
| Field | Detail |
|---|---|
| Endpoint | POST /v1/systemone |
| Model | typesafe/jev |
| Billing | Input tokens only — output tokens are free. |
| Launch price | Free at launch (zero rupees). Standard input-token pricing and the platform fee are covered in Pricing & credits. |
| Context | 64k tokens, text only. |
Decisions, not text
jev does not write text or code. It makes decisions.
There is no prose to read back — the output is the verdict: a yes/no probability, a chosen option, or a score, each with a confidence. If you need something written, that's a chat model.
jev is what you put around chat models to route, gate, classify and score — the calls you'd otherwise (expensively) make a big LLM do.
jev vs. a chat model
The request and the response are both shaped differently from a normal completion:
| A chat model | jev (System One) | |
|---|---|---|
| Flow | prompt → model → free text | state + typed questions → jev → structured answers + confidence |
| Shape | Open-ended. You read the output and hope it followed instructions. | Constrained to your schema. Every answer is machine-readable and calibrated. |
| Best for | Generating prose, code, or a summary a human will read. | A plain yes/no, pick-one, or score your code branches on. |
The three question types
Name each question and pick a type. Mix as many as you want in one call — they run in parallel, so extra questions cost you almost nothing more.
| Type | Answers | Use for | Example |
|---|---|---|---|
| noul | A yes/no question, answered as a probability. | Gates & flags — is it spam? does it need a human? is it on-topic? | "Does this express urgency?" → noul: 0.98 |
| choice | Pick exactly one of the options you define. | Routing & classification — which team, which handler, which label. | billing / technical / sales → choice: "billing" (0.90) |
| score | Rate against ordered levels you describe. | Severity & sentiment — how urgent, how angry, how good. | calm · civil · very angry → score: 1.0 (of 2) |
Confidence is a second axis.
The answer tells you what; the confidence tells you whether to trust it. A common pattern: act automatically above a confidence threshold, route to a human below it.
Call jev end to end
One customer message, three questions, one call. Send it to unoblox with your own ub-gw-… key — unoblox handles the upstream.
# curl
curl -sS https://api.unoblox.ai/v1/systemone \
-H "Authorization: Bearer $UNOBLOX_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "typesafe/jev",
"state": "My Stripe payout has failed for 3 days and I am losing sales. Help ASAP.",
"questions": {
"urgency": { "type":"noul", "instructions":"Does this express urgency?" },
"department": { "type":"choice", "instructions":"Which team handles this",
"criteria":{ "billing":"Payments", "technical":"Bugs", "sales":"Pricing" } },
"frustration": { "type":"score", "instructions":"How frustrated the customer is",
"criteria":["Calm","Frustrated but civil","Very angry"] }
}
}'{
"model": "jev-1.13.0",
"answers": {
"urgency": {
"type":"noul", "noul":0.98
},
"department": {
"type":"choice",
"choice":"billing",
"confidence":0.90,
"probabilities":{ "billing":0.90,
"technical":0.09, "sales":0.01 }
},
"frustration": {
"type":"score",
"score":1.0, "confidence":0.84
}
},
"usage": { "input_tokens":312, "output_tokens":48 }
}Your code reads the answer directly.
answers.department.choice === "billing" and answers.urgency.noul > 0.9 and routes — no parsing prose, no "please respond in JSON," no retries on malformed output.
Errors and retries
/v1/systemone fails the same way as every other endpoint — a typed JSON envelope and a stable error_code (see Errors) — plus these systemone-specific codes:
| Code | HTTP | Message | What to do |
|---|---|---|---|
| UB-GW-252 | 500 | Request decoding failed. | Check that the request body is valid JSON and within size limits, then retry. |
| UB-GW-253 | 503 | This model isn't approved for systemone inference, or is unavailable. | Confirm the model id is a systemone-capable model in good standing, or try again shortly. |
| UB-GW-254 | 503 | Native systemone inference is disabled. | System One isn't enabled on this deployment yet — contact support. |
| UB-GW-255 | 400 | systemone models must be used via /v1/systemone. | Call this model through POST /v1/systemone, not /v1/chat/completions. |
| UB-GW-256 | 400 | systemone request exceeds 1 MiB. | Reduce the combined size of state and questions to under 1 MiB. |
| UB-GW-257 | 400 | This request's state and questions exceed the workspace's configured prompt-size limit. | Shorten state and questions to fit your workspace's configured prompt-size limit. |
| UB-GW-258 | 503 | This request's usage could not be determined and requires manual review. | The provider answered without a usage figure, so the result was withheld. Retry the request; if it keeps happening, contact support with the request id. |
| UB-GW-259 | 503 | The result is unavailable, though usage was recorded for billing purposes. | The provider processed the request but its answer could not be returned; that attempt's input tokens are metered. Retry — a retry is a new request and is metered on its own. |
| UB-GW-260 | 503 | This feature is temporarily unavailable. Please try again later. | Wait and retry with backoff. |
| UB-GW-261 | 400 | The request body must be a JSON object. | Send a single JSON object as the request body, not an array or scalar. |
| UB-GW-262 | 400 | A non-empty "model" is required. | Include a non-empty model field naming the System One model to call. |
| UB-GW-263 | 400 | A non-empty "state" is required. / "state" must not be empty. | Include a non-empty state field describing what jev should judge. |
| UB-GW-264 | 400 | A non-empty "questions" object is required. / "questions" must not be empty. | Include a questions object with at least one named question. |
| UB-GW-265 | 400 | Each question must be a JSON object. / Each question needs a "type". / Each question's "type" must be one of noul, choice, or score. / Each question needs "instructions". | Give every question an object shape with a valid type (noul, choice, or score) and instructions. |
How upstream conditions surface: a malformed or oversized request — bad JSON, a missing or empty state/questions, an unrecognized question type — fails fast as one of the 4xx validation codes above; fix the request before retrying. A workspace or key over its call-rate limit gets the platform-wide 429 rate_limit_exceeded (see Errors) with a Retry-After header — wait, then retry. If the model provider itself rejects the request — for example a question whose criteria it cannot accept — its 4xx status is passed through with code UB-GW-037 and a sanitized message; nothing is metered, so fix the request before retrying. If the provider is overloaded or unreachable and no other endpoint can serve the request, you get UB-GW-036 — retry with exponential backoff, as you would for UB-GW-258/259/260. UB-GW-253/254 mean the model or the feature is not currently enabled, so an immediate retry will not help.
import time, requests
def call_systemone(payload, max_retries=3):
for attempt in range(max_retries):
resp = requests.post(
"https://api.unoblox.ai/v1/systemone",
headers={"Authorization": f"Bearer {UNOBLOX_KEY}"},
json=payload,
)
if resp.ok:
return resp.json()
if resp.status_code == 429:
time.sleep(float(resp.headers.get("retry-after", "1")))
elif resp.status_code >= 500:
time.sleep(2 ** attempt) # exponential backoff
else:
resp.raise_for_status() # 400 — fix the request, don't retry
resp.raise_for_status()What people build with it
Anywhere you're making a big model answer a small, structured question, jev does it faster, cheaper, and with a probability you can threshold on.
| Use case | Question type(s) | What it does |
|---|---|---|
| Intent routing | choice | Classify an incoming request and send it to the right handler — deterministic code, a specialist model, or a human. |
| Guardrails & moderation | noul | Gate every LLM input, output, or tool call — “is this a jailbreak / off-topic / unsafe?” — at a fraction of an LLM's cost. |
| Ticket & lead triage | choice + score | Team, priority and sentiment in one call, so the queue sorts itself before a person ever looks at it. |
| LLM output checks | noul | “Did the model actually answer the question?” “Is this citation supported?” Catch error modes without a second big call. |
jev, or a chat model?
| Reach for a chat model when… | Reach for jev when… |
|---|---|
| You need text, code, or a summary written. | You need a yes/no, a pick-one, or a score. |
| The output is open-ended. | Your code consumes the answer. |
| A human will read the result. | You want it calibrated, fast, and cheap. |
| → /v1/chat/completions | → /v1/systemone |
For the chat side of that table, see Quickstart and Routing.
They compose.
Let jev decide and route, then let a chat model generate. A jev call in front of an LLM often removes the need for the LLM call at all.