GLM 5.3

zai-org/GLM-5.3

Z.AI's current-generation frontier model: 1M context, thinking, tool calling, and the strongest coding scores in the Z.AI line-up, per its maker and independent benchmarking.

thinkingtool calling
Input
$1.20 / 1M
Output
$4.20 / 1M
Cached
$0.26 / 1M
Context
1M
Type
Chat
Maker
Zhipu
Released
2026-08-25

Compared to the portfolio

Intelligence vs price
Chat models on sference
Index
45/100
Blended
$1.95/1M
↖ Smarter for less
35
40
50
$0.25$1.00$4.00
45$1.95
GLM 5.3
↑ IndexBlended $ / 1M · log →

Artificial Analysis Intelligence Index v4.3 (0–100), measured across all providers by Artificial Analysis, not sference specifically. Blended price weights input and output 3:1. Hover a point for the reasoning-effort variant scored. Source

ModelContextInputOutputIndexVision
Kimi K31M$3.00$15.0044—
GLM 5.31M$1.20$4.2045—
DeepSeek V4.1 Flash1M$0.50$1.5039✓
GLM 5.3 Flash1M$0.20$0.6042✓
DeepSeek V4 Flash (0731)1M$0.28$0.56——
Clef 27Bdecisions64K$0.24———

Serverless rates per 1M tokens at realtime priority, confirmed per model when your account is provisioned. Dedicated deployments are billed per GPU-hour on monthly commitments.

Use it

POST /v1/chat/completions
Realtime, flex (service_tier: "flex")
POST /v1/responses
Realtime, flex, and the 24h async window (background: true)
POST /v1/messages
Anthropic Messages shape, realtime only
curl https://api.sference.com/v1/chat/completions \
  -H "Authorization: Bearer $SFERENCE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "zai-org/GLM-5.3",
    "messages": [{ "role": "user", "content": "Say hello in one sentence." }]
  }'

OpenAI-compatible; most clients switch with a base-URL change. Quickstart · Python SDK · Claude Code

Generated from the live catalog (GET /v1/models), so pricing, context, and capabilities always match the API. Machine-readable version: https://sference.com/models/zai-org/GLM-5.3.md