Clef 27B

Cloudflare/clef

A 27B decisions model for realtime classification, routing, and scoring: up to 16 questions answered about one input in a single call, billed on input tokens only.

Input
$0.24 / 1M
Context
64K
Type
Decisions
Maker
Cloudflare
Released
2026-09-30

Positioning

Clef is not a chat model. It reads one input (a support ticket, an email, a review, a transaction note) together with every question you need answered, and returns probabilities instead of generated text. Use it at the front of a pipeline where you would otherwise burn chat tokens on classification: routing, triage, prioritization, moderation, extraction of discrete fields. It answers up to 16 questions about the same input in one call, so a full routing-plus-urgency-plus-sentiment pass costs one request, not three.

In the sference portfolio it complements the chat models rather than competing with them: Clef does the cheap, deterministic front-of-pipeline work, and the chat models do the generation behind it. If you find yourself prompting a chat model to output only a single word or a JSON enum, that workload belongs here.

Upstream, Clef accepts image input as well. Serving images on sference's decisions endpoint is in the works — until it ships, the API accepts text-only inputs, and the catalog capability flag will flip when vision goes live.

Compared to the portfolio

ModelContextInputOutputIndexVision
Kimi K31M$3.00$15.0044—
GLM 5.31M$1.20$4.2045—
DeepSeek V4.1 Flash1M$0.50$1.5039✓
GLM 5.3 Flash1M$0.20$0.6042✓
DeepSeek V4 Flash (0731)1M$0.28$0.56——
Clef 27Bdecisions64K$0.24———

Serverless rates per 1M tokens at realtime priority, confirmed per model when your account is provisioned. Dedicated deployments are billed per GPU-hour on monthly commitments.

Use it

POST /v1/decisions
Realtime only; no batch, no streaming
curl https://api.sference.com/v1/decisions \
  -H "Authorization: Bearer $SFERENCE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "Cloudflare/clef",
    "state": "Please refund the duplicate charge.",
    "questions": {
      "urgent": { "type": "noul", "instructions": "Does this require immediate action?" }
    }
  }'

Decisions guide · API reference

Request limits

LimitValue
Context65,536 tokens (state + encoded questions/schema)
Questions≤ 16 per request
Options≤ 256 per question
Request body1 MiB (HTTP) / 2 MiB (gRPC)
Latency target2 s admission; 20 s hard deadline; 21 s API timeout
Prefix cacheOff
BillingSuccessful responses only; usage aggregated in 5 s flushes

Question types

  • choice — pick one criterion → label + probabilities per option
  • score — ordinal criteria → expected 0-based index
  • noul — yes/no → P(true) only
  • confidence — the largest per-question probability, returned alongside every answer

Example request

{
  "model": "Cloudflare/clef",
  "state": "Please refund the duplicate charge.",
  "questions": {
    "route":  { "type": "choice",
                "criteria": { "billing": "Payments and refunds",
                              "support": "Technical help" } },
    "urgent": { "type": "noul",
                "instructions": "Does this require immediate action?" }
  }
}

Serving stack

  • Weights: simonlehmann/clef-NVFP4 @ 817ac58a
  • Reference: Cloudflare/clef (Transformers BF16) @ 2f3de3dd
  • Precision: backbone NVFP4/FP8 mixed · head + option embeddings BF16 · softmax FP32

Resources

FAQ

When should I use Clef instead of a chat model?

Whenever the output you need is a label, a score, or a yes/no answer rather than generated prose. Clef is cheaper (input tokens only), faster (no autoregressive decode), and returns calibrated probabilities per option, which chat models do not. If your prompt asks a chat model to reply with a single word, that is a decisions workload.

How is Clef billed?

On input tokens only: $0.24 per 1M tokens, and only for successful responses. There is no output rate because decisions models generate no text; usage is aggregated in 5-second flushes.

Can Clef see images?

Upstream Clef accepts image input. On sference, serving images on the decisions endpoint is in the works; until it ships the endpoint accepts text-only inputs, and the model card's vision capability will update automatically when it goes live.

What happens on the hard deadline?

Requests admitted under the 2-second latency target are given a 20-second hard deadline; if a response is not ready by then the request fails with a timeout and is not billed. Retry with the same input — decisions are stateless.

Generated from the live catalog (GET /v1/models), so pricing, context, and capabilities always match the API. Machine-readable version: https://sference.com/models/Cloudflare/clef.md