GLM 5.3 Flash
zai-org/GLM-5.3-FlashThe fast, cheap tier of Z.AI's GLM 5.3 family: 1M context with vision and thinking, tuned for high-volume interactive traffic where cost per token dominates.
- Input
- $0.20 / 1M
- Output
- $0.60 / 1M
- Cached
- $0.07 / 1M
- Context
- 1M
- Type
- Chat
- Maker
- Zhipu
- Released
- 2026-08-25
Compared to the portfolio
Artificial Analysis Intelligence Index v4.3 (0–100), measured across all providers by Artificial Analysis, not sference specifically. Blended price weights input and output 3:1. Hover a point for the reasoning-effort variant scored. Source
| Model | Context | Input | Output | Index | Vision |
|---|---|---|---|---|---|
| Kimi K3 | 1M | $3.00 | $15.00 | 44 | — |
| GLM 5.3 | 1M | $1.20 | $4.20 | 45 | — |
| DeepSeek V4.1 Flash | 1M | $0.50 | $1.50 | 39 | ✓ |
| GLM 5.3 Flash | 1M | $0.20 | $0.60 | 42 | ✓ |
| DeepSeek V4 Flash (0731) | 1M | $0.28 | $0.56 | — | — |
| Clef 27Bdecisions | 64K | $0.24 | — | — | — |
Serverless rates per 1M tokens at realtime priority, confirmed per model when your account is provisioned. Dedicated deployments are billed per GPU-hour on monthly commitments.
Use it
POST /v1/chat/completions- Realtime, flex (service_tier: "flex")
POST /v1/responses- Realtime, flex, and the 24h async window (background: true)
POST /v1/messages- Anthropic Messages shape, realtime only
curl https://api.sference.com/v1/chat/completions \
-H "Authorization: Bearer $SFERENCE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "zai-org/GLM-5.3-Flash",
"messages": [{ "role": "user", "content": "Say hello in one sentence." }]
}'OpenAI-compatible; most clients switch with a base-URL change. Quickstart · Python SDK · Claude Code
Generated from the live catalog (GET /v1/models), so pricing, context, and capabilities always match the API. Machine-readable version: https://sference.com/models/zai-org/GLM-5.3-Flash.md