Guides

Decisions

Realtime classification with POST /v1/decisions — ask up to 16 choice, score, or yes/no questions about one input and get calibrated probabilities back in a single call.

Decisions

POST /v1/decisions answers structured questions about one input in a single realtime call. Instead of prompting a chat model and parsing its text, you describe the questions and their options. Back come probabilities for every option: routing, triage, moderation, intent detection, and tagging, with no output to parse.

Decisions run on Clef (Cloudflare/clef), a 27B model built for this task. It reads the input and every question together and scores all options at once. It generates no text, so you pay for input tokens only.

Request

curl https://api.sference.com/v1/decisions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $SFERENCE_API_KEY" \
  -d '{
    "model": "Cloudflare/clef",
    "state": "I was charged twice for my subscription this month. Please refund the duplicate charge.",
    "questions": {
      "route": {
        "type": "choice",
        "criteria": {
          "billing": "Payments, invoices, refunds",
          "technical": "Bugs and outages",
          "sales": "Pricing and new plans"
        }
      },
      "urgency": { "type": "score", "criteria": ["Low", "Medium", "High"] },
      "refund": { "type": "noul", "instructions": "Does the customer ask for a refund?" }
    }
  }'
  • state: the input to decide about. It is either text or any JSON value, such as a ticket object or a conversation. The whole request is limited to 1 MiB, and the model has a 64K-token context.
  • questions: 1–16 named questions, all answered in the same call. Each may carry optional instructions that frame the question.
TypecriteriaAnswer
choiceObject of 1–256 labels → descriptionschoice (the most likely label), confidence, and probabilities per label
scoreOrdered list of 1–256 levels, lowest firstscore (expected level index, 0-based), confidence, legend, and probabilities per level
noulOptional {"true": …, "false": …} descriptionsnoul: the probability that the answer is yes

Response

{
  "model": "Cloudflare/clef",
  "answers": {
    "route": {
      "type": "choice",
      "choice": "billing",
      "confidence": 0.9909,
      "probabilities": { "billing": 0.9909, "technical": 0.0055, "sales": 0.0036 }
    },
    "urgency": {
      "type": "score",
      "score": 1.5588,
      "confidence": 0.5892,
      "legend": { "0": "Low", "1": "Medium", "2": "High" },
      "probabilities": { "0": 0.0304, "1": 0.3804, "2": 0.5892 }
    },
    "refund": { "type": "noul", "noul": 0.9928 }
  },
  "usage": { "input_tokens": 323, "output_tokens": 0 }
}

The probabilities are the model's full distribution, not a single sampled answer. Pick your own thresholds per question: for example, auto-route above 0.9 and send anything lower to a human.

Python SDK and CLI

from sference_sdk import ChoiceQuestion, NoulQuestion, ScoreQuestion, SferenceClient

client = SferenceClient()  # reads SFERENCE_API_KEY

decision = client.create_decision(
    model="Cloudflare/clef",
    state="I was charged twice this month, please refund me.",
    questions={
        "route": ChoiceQuestion(criteria={"billing": "Payments", "support": "Everything else"}),
        "urgency": ScoreQuestion(criteria=["Low", "Medium", "High"]),
        "refund": NoulQuestion(instructions="Does the customer ask for a refund?"),
    },
)
print(decision.answers["route"].choice, decision.answers["refund"].noul)
sference decisions create --model Cloudflare/clef \
  --state "I was charged twice, please refund me" \
  --questions '{"route": {"type": "choice", "criteria": {"billing": "Payments", "support": "Other"}}}'

AsyncSferenceClient.create_decision is the async equivalent. Questions may also be plain dicts with a type. With the CLI, --state-json sends a JSON state, and both options accept @file (or @- for stdin).

Limits and errors

Decisions are realtime only. They don't support batches, background responses, or the flex tier.

StatusMeaning
400Invalid questions (count or criteria bounds), or a model that does not support decisions
402Insufficient credit or spending limit reached
413Request body over 1 MiB
429No decision capacity right now. Retry with backoff.
503Decisions temporarily unavailable. Honour Retry-After.
504Deadline passed while the request was running. Not charged.

Only successful decisions are billed, at $0.24 per 1M input tokens. The input count includes the rendered questions.