# Clef 27B

> A 27B decisions model for realtime classification, routing, and scoring: up to 16 questions answered about one input in a single call, billed on input tokens only.

Source: https://sference.com/models/Cloudflare/clef

Catalog id: `Cloudflare/clef` · Maker: Cloudflare · Modality: decisions · Released: 2026-09-30

## Pricing

USD per 1M tokens, realtime priority: billed on input tokens only: $0.24 per 1M input tokens.
Decisions generate no text, so there is no output rate to bill.

## Specs

- Context window: 64K tokens
- Endpoints: `POST /v1/decisions` (realtime only; no batch, no stream)

## Portfolio comparison

Every model in the live catalog (this one in bold):

| Model | Modality | Context | Input | Output | Vision |
| --- | --- | --- | --- | --- | --- |
| `moonshotai/Kimi-K3` | text | 1M | $3.00 | $15.00 | — |
| `zai-org/GLM-5.3` | text | 1M | $1.20 | $4.20 | — |
| `deepseek-ai/DeepSeek-V4.1-Flash` | text | 1M | $0.50 | $1.50 | ✓ |
| `zai-org/GLM-5.3-Flash` | text | 1M | $0.20 | $0.60 | ✓ |
| `deepseek-ai/DeepSeek-V4-Flash-0731` | text | 1M | $0.28 | $0.56 | — |
| **`Cloudflare/clef`** | decisions | 64K | $0.24 | — | — |

## Quickstart

```bash
curl https://api.sference.com/v1/decisions \
  -H "Authorization: Bearer $SFERENCE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "Cloudflare/clef",
    "state": "Please refund the duplicate charge.",
    "questions": {
      "urgent": { "type": "noul", "instructions": "Does this require immediate action?" }
    }
  }'
```

## Positioning

Clef is not a chat model. It reads one input (a support ticket, an email, a review, a transaction note) together with every question you need answered, and returns probabilities instead of generated text. Use it at the front of a pipeline where you would otherwise burn chat tokens on classification: routing, triage, prioritization, moderation, extraction of discrete fields. It answers up to 16 questions about the same input in one call, so a full routing-plus-urgency-plus-sentiment pass costs one request, not three.

In the sference portfolio it complements the chat models rather than competing with them: Clef does the cheap, deterministic front-of-pipeline work, and the chat models do the generation behind it. If you find yourself prompting a chat model to output only a single word or a JSON enum, that workload belongs here.

Upstream, Clef accepts image input as well. Serving images on sference's decisions endpoint is in the works — until it ships, the API accepts text-only inputs, and the catalog capability flag will flip when vision goes live.

## Request limits

| Limit | Value |
| --- | --- |
| Context | 65,536 tokens (state + encoded questions/schema) |
| Questions | ≤ 16 per request |
| Options | ≤ 256 per question |
| Request body | 1 MiB (HTTP) / 2 MiB (gRPC) |
| Latency target | 2 s admission; 20 s hard deadline; 21 s API timeout |
| Prefix cache | Off |
| Billing | Successful responses only; usage aggregated in 5 s flushes |

## Question types

- **choice** — pick one criterion → label + probabilities per option
- **score** — ordinal criteria → expected 0-based index
- **noul** — yes/no → P(true) only
- **confidence** — the largest per-question probability, returned alongside every answer

## Example request

```json
{
  "model": "Cloudflare/clef",
  "state": "Please refund the duplicate charge.",
  "questions": {
    "route":  { "type": "choice",
                "criteria": { "billing": "Payments and refunds",
                              "support": "Technical help" } },
    "urgent": { "type": "noul",
                "instructions": "Does this require immediate action?" }
  }
}
```

## Serving stack

- Weights: `simonlehmann/clef-NVFP4` @ 817ac58a
- Reference: `Cloudflare/clef` (Transformers BF16) @ 2f3de3dd
- Precision: backbone NVFP4/FP8 mixed · head + option embeddings BF16 · softmax FP32

## Resources

- Weights: [huggingface.co/simonlehmann/clef-NVFP4](https://huggingface.co/simonlehmann/clef-NVFP4)
- Reference model: [huggingface.co/Cloudflare/clef](https://huggingface.co/Cloudflare/clef)
- Decisions guide: [sference.com/docs/guides/decisions](https://sference.com/docs/guides/decisions)
- API reference: [sference.com/docs/api-reference](https://sference.com/docs/api-reference)

## FAQ

### When should I use Clef instead of a chat model?

Whenever the output you need is a label, a score, or a yes/no answer rather than generated prose. Clef is cheaper (input tokens only), faster (no autoregressive decode), and returns calibrated probabilities per option, which chat models do not. If your prompt asks a chat model to reply with a single word, that is a decisions workload.

### How is Clef billed?

On input tokens only: $0.24 per 1M tokens, and only for successful responses. There is no output rate because decisions models generate no text; usage is aggregated in 5-second flushes.

### Can Clef see images?

Upstream Clef accepts image input. On sference, serving images on the decisions endpoint is in the works; until it ships the endpoint accepts text-only inputs, and the model card's vision capability will update automatically when it goes live.

### What happens on the hard deadline?

Requests admitted under the 2-second latency target are given a 20-second hard deadline; if a response is not ready by then the request fails with a timeout and is not billed. Retry with the same input — decisions are stateless.
