# Decisions

> Realtime classification with POST /v1/decisions — ask up to 16 choice, score, or yes/no questions about one input and get calibrated probabilities back in a single call.

Source: https://sference.com/docs/guides/decisions

`POST /v1/decisions` answers structured questions about one input in a single realtime call. Instead of prompting a chat model and parsing its text, you describe the **questions** and their options. Back come **probabilities** for every option: routing, triage, moderation, intent detection, and tagging, with no output to parse.

Decisions run on **[Clef](https://sference.com/docs/models#decisions-models)** (`Cloudflare/clef`), a 27B model built for this task. It reads the input and every question together and scores all options at once. It generates no text, so you pay for **input tokens only**.

## Request

```bash
curl https://api.sference.com/v1/decisions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $SFERENCE_API_KEY" \
  -d '{
    "model": "Cloudflare/clef",
    "state": "I was charged twice for my subscription this month. Please refund the duplicate charge.",
    "questions": {
      "route": {
        "type": "choice",
        "criteria": {
          "billing": "Payments, invoices, refunds",
          "technical": "Bugs and outages",
          "sales": "Pricing and new plans"
        }
      },
      "urgency": { "type": "score", "criteria": ["Low", "Medium", "High"] },
      "refund": { "type": "noul", "instructions": "Does the customer ask for a refund?" }
    }
  }'
```

- **`state`**: the input to decide about. It is either text or any JSON value, such as a ticket object or a conversation. The whole request is limited to 1 MiB, and the model has a 64K-token context.
- **`questions`**: 1–16 named questions, all answered in the same call. Each may carry optional **`instructions`** that frame the question.

| Type | `criteria` | Answer |
| --- | --- | --- |
| **`choice`** | Object of 1–256 labels → descriptions | `choice` (the most likely label), `confidence`, and `probabilities` per label |
| **`score`** | Ordered list of 1–256 levels, lowest first | `score` (expected level index, 0-based), `confidence`, `legend`, and `probabilities` per level |
| **`noul`** | Optional `{"true": …, "false": …}` descriptions | `noul`: the probability that the answer is yes |

## Response

```json
{
  "model": "Cloudflare/clef",
  "answers": {
    "route": {
      "type": "choice",
      "choice": "billing",
      "confidence": 0.9909,
      "probabilities": { "billing": 0.9909, "technical": 0.0055, "sales": 0.0036 }
    },
    "urgency": {
      "type": "score",
      "score": 1.5588,
      "confidence": 0.5892,
      "legend": { "0": "Low", "1": "Medium", "2": "High" },
      "probabilities": { "0": 0.0304, "1": 0.3804, "2": 0.5892 }
    },
    "refund": { "type": "noul", "noul": 0.9928 }
  },
  "usage": { "input_tokens": 323, "output_tokens": 0 }
}
```

The probabilities are the model's full distribution, not a single sampled answer. Pick your own thresholds per question: for example, auto-route above 0.9 and send anything lower to a human.

## Python SDK and CLI

```python
from sference_sdk import ChoiceQuestion, NoulQuestion, ScoreQuestion, SferenceClient

client = SferenceClient()  # reads SFERENCE_API_KEY

decision = client.create_decision(
    model="Cloudflare/clef",
    state="I was charged twice this month, please refund me.",
    questions={
        "route": ChoiceQuestion(criteria={"billing": "Payments", "support": "Everything else"}),
        "urgency": ScoreQuestion(criteria=["Low", "Medium", "High"]),
        "refund": NoulQuestion(instructions="Does the customer ask for a refund?"),
    },
)
print(decision.answers["route"].choice, decision.answers["refund"].noul)
```

```bash
sference decisions create --model Cloudflare/clef \
  --state "I was charged twice, please refund me" \
  --questions '{"route": {"type": "choice", "criteria": {"billing": "Payments", "support": "Other"}}}'
```

`AsyncSferenceClient.create_decision` is the async equivalent. Questions may also be plain dicts with a `type`. With the CLI, `--state-json` sends a JSON state, and both options accept `@file` (or `@-` for stdin).

## Limits and errors

Decisions are **realtime only**. They don't support batches, background responses, or the flex tier.

| Status | Meaning |
| --- | --- |
| **400** | Invalid questions (count or criteria bounds), or a model that does not support decisions |
| **402** | Insufficient credit or spending limit reached |
| **413** | Request body over 1 MiB |
| **429** | No decision capacity right now. Retry with backoff. |
| **503** | Decisions temporarily unavailable. Honour `Retry-After`. |
| **504** | Deadline passed while the request was running. **Not charged.** |

Only successful decisions are billed, at $0.24 per 1M input tokens. The input count includes the rendered questions.
