Decisions
Realtime classification with POST /v1/decisions — ask up to 16 choice, score, or yes/no questions about one input and get calibrated probabilities back in a single call.
Decisions
POST /v1/decisions answers structured questions about one input in a single realtime call. Instead of prompting a chat model and parsing its text, you describe the questions and their options. Back come probabilities for every option: routing, triage, moderation, intent detection, and tagging, with no output to parse.
Decisions run on Clef (Cloudflare/clef), a 27B model built for this task. It reads the input and every question together and scores all options at once. It generates no text, so you pay for input tokens only.
Request
curl https://api.sference.com/v1/decisions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $SFERENCE_API_KEY" \
-d '{
"model": "Cloudflare/clef",
"state": "I was charged twice for my subscription this month. Please refund the duplicate charge.",
"questions": {
"route": {
"type": "choice",
"criteria": {
"billing": "Payments, invoices, refunds",
"technical": "Bugs and outages",
"sales": "Pricing and new plans"
}
},
"urgency": { "type": "score", "criteria": ["Low", "Medium", "High"] },
"refund": { "type": "noul", "instructions": "Does the customer ask for a refund?" }
}
}'state: the input to decide about. It is either text or any JSON value, such as a ticket object or a conversation. The whole request is limited to 1 MiB, and the model has a 64K-token context.questions: 1–16 named questions, all answered in the same call. Each may carry optionalinstructionsthat frame the question.
| Type | criteria | Answer |
|---|---|---|
choice | Object of 1–256 labels → descriptions | choice (the most likely label), confidence, and probabilities per label |
score | Ordered list of 1–256 levels, lowest first | score (expected level index, 0-based), confidence, legend, and probabilities per level |
noul | Optional {"true": …, "false": …} descriptions | noul: the probability that the answer is yes |
Response
{
"model": "Cloudflare/clef",
"answers": {
"route": {
"type": "choice",
"choice": "billing",
"confidence": 0.9909,
"probabilities": { "billing": 0.9909, "technical": 0.0055, "sales": 0.0036 }
},
"urgency": {
"type": "score",
"score": 1.5588,
"confidence": 0.5892,
"legend": { "0": "Low", "1": "Medium", "2": "High" },
"probabilities": { "0": 0.0304, "1": 0.3804, "2": 0.5892 }
},
"refund": { "type": "noul", "noul": 0.9928 }
},
"usage": { "input_tokens": 323, "output_tokens": 0 }
}The probabilities are the model's full distribution, not a single sampled answer. Pick your own thresholds per question: for example, auto-route above 0.9 and send anything lower to a human.
Python SDK and CLI
from sference_sdk import ChoiceQuestion, NoulQuestion, ScoreQuestion, SferenceClient
client = SferenceClient() # reads SFERENCE_API_KEY
decision = client.create_decision(
model="Cloudflare/clef",
state="I was charged twice this month, please refund me.",
questions={
"route": ChoiceQuestion(criteria={"billing": "Payments", "support": "Everything else"}),
"urgency": ScoreQuestion(criteria=["Low", "Medium", "High"]),
"refund": NoulQuestion(instructions="Does the customer ask for a refund?"),
},
)
print(decision.answers["route"].choice, decision.answers["refund"].noul)sference decisions create --model Cloudflare/clef \
--state "I was charged twice, please refund me" \
--questions '{"route": {"type": "choice", "criteria": {"billing": "Payments", "support": "Other"}}}'AsyncSferenceClient.create_decision is the async equivalent. Questions may also be plain dicts with a type. With the CLI, --state-json sends a JSON state, and both options accept @file (or @- for stdin).
Limits and errors
Decisions are realtime only. They don't support batches, background responses, or the flex tier.
| Status | Meaning |
|---|---|
| 400 | Invalid questions (count or criteria bounds), or a model that does not support decisions |
| 402 | Insufficient credit or spending limit reached |
| 413 | Request body over 1 MiB |
| 429 | No decision capacity right now. Retry with backoff. |
| 503 | Decisions temporarily unavailable. Honour Retry-After. |
| 504 | Deadline passed while the request was running. Not charged. |
Only successful decisions are billed, at $0.24 per 1M input tokens. The input count includes the rendered questions.
Responses & streams
Realtime and async responses, flex processing, server-sent events, and streaming IDs under /v1/responses and /v1/streams.
Anthropic Messages API
POST /v1/messages on sference: an Anthropic Messages-compatible endpoint for open models. Auth, streaming SSE, tool use, extended thinking, vision, usage, and the exact compatibility surface.