Model cards
Every model we serve, with pricing, capabilities, and a quickstart — generated from the live catalog, so it never drifts.
Artificial Analysis Intelligence Index v4.3 (0–100), measured across all providers by Artificial Analysis, not sference specifically. Blended price weights input and output 3:1. Source
Chat · 5
zai-org/GLM-5.3Z.AI's current-generation frontier model: 1M context, thinking, tool calling, and the strongest coding scores in the Z.AI line-up, per its maker and independent benchmarking.
moonshotai/Kimi-K3Moonshot AI's frontier reasoning model: 1M context, thinking, and tool calling, served with the lowest average latency of any provider measured on the Requesty gateway.
zai-org/GLM-5.3-FlashThe fast, cheap tier of Z.AI's GLM 5.3 family: 1M context with vision and thinking, tuned for high-volume interactive traffic where cost per token dominates.
deepseek-ai/DeepSeek-V4.1-FlashDeepSeek's fast tier with vision input: 1M context and the best price-to-throughput ratio in the sference catalog for vision-capable workloads.
deepseek-ai/DeepSeek-V4-Flash-0731Decisions · 1
Decisions models answer questions about an input — classify, route, score — and return probabilities instead of generated text. Billed on input tokens only.