Kimi K3

moonshotai/Kimi-K3

Moonshot AI's frontier reasoning model: 1M context, thinking, and tool calling, served with the lowest average latency of any provider measured on the Requesty gateway.

thinkingtool calling
Input
$3.00 / 1M
Output
$15.00 / 1M
Cached
$0.45 / 1M
Context
1M
Type
Chat
Maker
Moonshot
Released
2026-07-27

Positioning

The largest reasoning model in the sference catalog and the default pick for hard planning turns: long-horizon agentic workloads, planning-heavy coding sessions, and anything where the model needs to think before it answers. Kimi K3 trades speed for depth — for short interactive turns, the GLM 5.3 family is cheaper and faster; when the question is hard, Kimi K3 is the one that reasons its way through.

Thinking is supported on the OpenAI-shaped endpoints with thinking: {"type": "enabled"}.

Compared to the portfolio

Intelligence vs price
Chat models on sference
Index
44/100
Blended
$6.00/1M
↖ Smarter for less
35
40
50
$0.25$1.00
44$6.00
Kimi K3
↑ IndexBlended $ / 1M · log →

Artificial Analysis Intelligence Index v4.3 (0–100), measured across all providers by Artificial Analysis, not sference specifically. Blended price weights input and output 3:1. Hover a point for the reasoning-effort variant scored. Source

ModelContextInputOutputIndexVision
Kimi K31M$3.00$15.0044—
GLM 5.31M$1.20$4.2045—
DeepSeek V4.1 Flash1M$0.50$1.5039✓
GLM 5.3 Flash1M$0.20$0.6042✓
DeepSeek V4 Flash (0731)1M$0.28$0.56——
Clef 27Bdecisions64K$0.24———

Serverless rates per 1M tokens at realtime priority, confirmed per model when your account is provisioned. Dedicated deployments are billed per GPU-hour on monthly commitments.

Benchmarks

Average latency — Requesty live gateway comparison, mid-August 2026426 mslower is better
source
Output throughput — Requesty live gateway comparison, mid-August 202676 tok/s
source

Use it

POST /v1/chat/completions
Realtime, flex (service_tier: "flex")
POST /v1/responses
Realtime, flex, and the 24h async window (background: true)
POST /v1/messages
Anthropic Messages shape, realtime only
curl https://api.sference.com/v1/chat/completions \
  -H "Authorization: Bearer $SFERENCE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "moonshotai/Kimi-K3",
    "messages": [{ "role": "user", "content": "Say hello in one sentence." }]
  }'

OpenAI-compatible; most clients switch with a base-URL change. Quickstart · Python SDK · Claude Code

Generated from the live catalog (GET /v1/models), so pricing, context, and capabilities always match the API. Machine-readable version: https://sference.com/models/moonshotai/Kimi-K3.md