# GLM 5.3 Flash

> The fast, cheap tier of Z.AI's GLM 5.3 family: 1M context with vision and thinking, tuned for high-volume interactive traffic where cost per token dominates.

Source: https://sference.com/models/zai-org/GLM-5.3-Flash

Catalog id: `zai-org/GLM-5.3-Flash` · Maker: Zhipu · Modality: text_generation · Released: 2026-08-25

## Pricing

USD per 1M tokens, realtime priority: $0.20 in / $0.60 out / $0.07 cached input, per 1M tokens.

## Specs

- Context window: 1M tokens
- Capabilities: vision, thinking, tool calling
- Endpoints: `POST /v1/chat/completions` and `POST /v1/responses` (realtime, flex via `service_tier: "flex"`, and the 24h async window), plus `POST /v1/messages` (Anthropic-shaped, realtime only)

## Portfolio comparison

Every model in the live catalog (this one in bold):

| Model | Modality | Context | Input | Output | Vision |
| --- | --- | --- | --- | --- | --- |
| `moonshotai/Kimi-K3` | text | 1M | $3.00 | $15.00 | — |
| `zai-org/GLM-5.3` | text | 1M | $1.20 | $4.20 | — |
| `deepseek-ai/DeepSeek-V4.1-Flash` | text | 1M | $0.50 | $1.50 | ✓ |
| **`zai-org/GLM-5.3-Flash`** | text | 1M | $0.20 | $0.60 | ✓ |
| `deepseek-ai/DeepSeek-V4-Flash-0731` | text | 1M | $0.28 | $0.56 | — |
| `Cloudflare/clef` | decisions | 64K | $0.24 | — | — |

## Quickstart

```bash
curl https://api.sference.com/v1/chat/completions \
  -H "Authorization: Bearer $SFERENCE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "zai-org/GLM-5.3-Flash",
    "messages": [{ "role": "user", "content": "Say hello in one sentence." }]
  }'
```

## Intelligence

Artificial Analysis Intelligence Index v4.3: **42 / 100**. Source: https://artificialanalysis.ai/evaluations/artificial-analysis-intelligence-index. Measured across all providers, not sference specifically.
