# Kimi K3

> Moonshot AI's frontier reasoning model: 1M context, thinking, and tool calling, served with the lowest average latency of any provider measured on the Requesty gateway.

Source: https://sference.com/models/moonshotai/Kimi-K3

Catalog id: `moonshotai/Kimi-K3` · Maker: Moonshot · Modality: text_generation · Released: 2026-07-27

## Pricing

USD per 1M tokens, realtime priority: $3.00 in / $15.00 out / $0.45 cached input, per 1M tokens.

## Specs

- Context window: 1M tokens
- Capabilities: thinking, tool calling
- Endpoints: `POST /v1/chat/completions` and `POST /v1/responses` (realtime, flex via `service_tier: "flex"`, and the 24h async window), plus `POST /v1/messages` (Anthropic-shaped, realtime only)

## Portfolio comparison

Every model in the live catalog (this one in bold):

| Model | Modality | Context | Input | Output | Vision |
| --- | --- | --- | --- | --- | --- |
| **`moonshotai/Kimi-K3`** | text | 1M | $3.00 | $15.00 | — |
| `zai-org/GLM-5.3` | text | 1M | $1.20 | $4.20 | — |
| `deepseek-ai/DeepSeek-V4.1-Flash` | text | 1M | $0.50 | $1.50 | ✓ |
| `zai-org/GLM-5.3-Flash` | text | 1M | $0.20 | $0.60 | ✓ |
| `deepseek-ai/DeepSeek-V4-Flash-0731` | text | 1M | $0.28 | $0.56 | — |
| `Cloudflare/clef` | decisions | 64K | $0.24 | — | — |

## Quickstart

```bash
curl https://api.sference.com/v1/chat/completions \
  -H "Authorization: Bearer $SFERENCE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "moonshotai/Kimi-K3",
    "messages": [{ "role": "user", "content": "Say hello in one sentence." }]
  }'
```

## Intelligence

Artificial Analysis Intelligence Index v4.3 (max effort): **44 / 100**. Source: https://artificialanalysis.ai/evaluations/artificial-analysis-intelligence-index. Measured across all providers, not sference specifically.

## Benchmarks

- Average latency — Requesty live gateway comparison, mid-August 2026: **426 ms** (lower is better). Source: https://sference.com/#engine
- Output throughput — Requesty live gateway comparison, mid-August 2026: **76 tok/s**. Source: https://sference.com/#engine

## Positioning

The largest reasoning model in the sference catalog and the default pick for hard planning turns: long-horizon agentic workloads, planning-heavy coding sessions, and anything where the model needs to think before it answers. Kimi K3 trades speed for depth — for short interactive turns, the GLM 5.3 family is cheaper and faster; when the question is hard, Kimi K3 is the one that reasons its way through.

Thinking is supported on the OpenAI-shaped endpoints with `thinking: {"type": "enabled"}`.
