Models / GLM-5.3-Flash
GLM-5.3-Flash
GLM-5.3-Flash family · Provider: zhipu-ai · Status: active · Released: 2026-08-26
GLM-5.3-Flash is a GLM-5.3-Flash model developed by zhipu-ai, a Chinese AI company headquartered in Beijing Zhipu Huazhang Technology Co., Ltd. (北京智谱华章科技股份有限公司), Beijing, China, released 2026-08-26 with a 1,048,576-token context window and open weights.
| Key fact | Value |
|---|---|
| Model ID | glm-5.3-flash |
| Version | Flash |
| Architecture | 320B total / 18B active; first open-source frontier model combining sparse + linear attention; mHC hyper-connections; 30T-token multimodal pre-training corpus |
| Context window | 1,048,576 tokens |
| Max output | 131,072 tokens |
| Open weights | Yes |
| License | Apache-2.0 (per GitHub repo metadata; README has no separate weights-license section - verify per-model HF cards before reuse) |
| Self-hosting | Yes |
| API available | Yes |
| API pricing | $0.15 input / $0.5 output per 1M tokens (USD) · provider pricing page |
| Regions | international, china |
| Cloud providers | Z.ai, BigModel |
Capabilities
| Capability | Supported |
|---|---|
| Reasoning | Yes |
| Coding | Yes |
| Math | Unknown |
| Chinese | Unknown |
| English | Unknown |
| Multilingual | Unknown |
| Vision | Yes |
| Audio | Unknown |
| Video | Yes |
| Tool calling | Unknown |
| Function calling | Unknown |
| Structured output | Unknown |
| Agent capability | Yes |
| RAG | Unknown |
| Computer use | Yes |
Benchmark results
| Benchmark | Version | Score | Metric | Date | Source type | Source |
|---|---|---|---|---|---|---|
| Artificial Analysis Intelligence Index v4.1.1 | — | 57 | index score | 2026-08-26 | vendor_reported | link |
| DeepSWE v1.1 | — | 63.4 | accuracy | 2026-08-26 | vendor_reported | link |
| AutomationBench | — | 48.8 | accuracy | 2026-08-26 | vendor_reported | link |
Benchmark scores are single data points, not universal rankings. Vendor-reported scores are labeled as such.
Known limitations
- FlashX tier not yet available on the GLM Coding Plan (pay-as-you-go only)
- Reasoning always enabled; cannot be disabled
- Z.ai Code Bench is a private in-house benchmark
- Benchmarks vendor-reported; not independently verified
GLM-5.3-Flash (2026-08-26) is Zhipu’s multimodal coding model with open weights (320B total / 18B active, FP8): 1M context, 128K max output, input modalities of video, image, text and file. It combines sparse and linear attention and is positioned for visual coding loops (observe - code - test), computer use (BUA/CUA), browser/GUI agents, office workflows, video understanding, 3D and CAD tasks. The FlashX tier serves at up to 200 tokens/s.
International pricing: Flash $0.15 input / $0.50 output per 1M tokens (cached $0.03); FlashX $0.37 / $1.25 (cached $0.075), as of 2026-09-20. Zhipu states all Flash traffic is served on Chinese AI chips.
Provider
Pricing
Benchmarks with results for this model
Agents built on this model
Related models (same family)
API
Model family timeline
| Model | Released | Status |
|---|---|---|
| GLM-5.3-FlashX | — | active |
| GLM-5.3-Flash (this page) | 2026-08-26 | active |
Data interpretation
| Field | Value | Evidence type |
|---|---|---|
| Context window | 1,048,576 tokens | Official |
| Architecture | 320B total / 18B active; first open-source frontier model combining sparse + linear attention; mHC hyper-connections; 30T-token multimodal pre-training corpus | Official |
| Open weights | Yes | Official |
| License | Apache-2.0 (per GitHub repo metadata; README has no separate weights-license section - verify per-model HF cards before reuse) | Official |
| API pricing | $0.15 / $0.5 per 1M tokens (USD) | Official |
| Artificial Analysis Intelligence Index v4.1.1 | 57 index score | Vendor-reported |
| DeepSWE v1.1 | 63.4 accuracy | Vendor-reported |
| AutomationBench | 48.8 accuracy | Vendor-reported |
| Family position | 2 models in the GLM-5.3-Flash family | China AI Hub analysis |
Evidence types: Official = vendor documentation, pricing pages or model cards. Vendor-reported = benchmark scores published by the vendor. China AI Hub analysis = derived from the database itself. See the sourcing policy.
Sources
Confidence and source hierarchy per the sourcing policy. Facts change; check the source before relying on this page.
What is the context window of GLM-5.3-Flash?
GLM-5.3-Flash has a 1,048,576-token context window and a maximum output of 131,072 tokens.
Is GLM-5.3-Flash open weight?
Yes — GLM-5.3-Flash weights are openly available under the Apache-2.0 (per GitHub repo metadata; README has no separate weights-license section - verify per-model HF cards before reuse) license.
How much does GLM-5.3-Flash cost through the API?
$0.15 per 1M input tokens and $0.5 per 1M output tokens (USD).
Where does China AI Hub get its GLM-5.3-Flash data?
From 3 sources (official pages first), last verified 2026-09-20. Benchmark scores are labeled by source type; see the sourcing policy for details.