GLM 5.3 Flash
GLM-5.3-Flash is Z.ai's MIT-licensed, open-weight multimodal MoE for coding, long-horizon agents, computer use, and professional work. It has 320B total parameters with 18B activated per token, hybrid linear and sparse attention, native image input, a 1M-token context window, and official support for SGLang, vLLM, and TokenSpeed. Its total weight footprint makes it a server-class local model despite its efficient active-parameter count.
Deployment and license note
GLM 5.3 Flash is a server-class local model. Its 18B active parameters reduce per-token compute, but the complete 320B weights still require roughly 384GB for the official FP8 deployment class or substantially more for BF16, before context and KV-cache overhead. Use current SGLang, vLLM, KTransformers, or TokenSpeed releases and plan for multi-GPU serving.
Task Fit
Tool use, repo work, terminal workflows, and coding benchmarks.
Code generation, debugging, refactoring, and benchmark signal.
General writing, Q&A, and assistant use.
Document QA benefits from long context and instruction following.
Image or visual understanding, not necessarily image generation.
Not marked for image generation in the current library.
Not marked for video generation in the current library.
Not marked for voice in the current library.
Source Confidence
Variants and Quant Artifacts
Choose the artifact first; hardware fit follows from RAM, VRAM, format, and runtime.
| Quant | Format | Quality | Min RAM | Reco RAM | Runtime | Action |
|---|---|---|---|---|---|---|
| FP8 | safetensors | balanced | 384GB | 512GB | transformers, vllm, sglang, ktransformers, tokenspeed | Plan with this |
| BF16 | safetensors | high | 768GB | 1TB | transformers, vllm, sglang, ktransformers, tokenspeed | Plan with this |
Recommended Hardware
No compatible local hardware found in the current hardware pool.
Benchmarks
Source and Review
Execution evidence
Run GLM 5.3 Flash with a documented recipe
Recipes connect hardware, a model artifact, tools, settings, verification, and a reportable result.
No verified recipe is linked to this record yet.
Compatibility estimates remain available in the planner. A recipe appears here only after its exact stack and verification protocol are documented.
Similar Models
Z.ai's GLM 5.2 FP8. Auto-imported from Hugging Face seed; review memory and benchmark details before using for final ranking.
Z.ai's GLM 5.1 FP8. Auto-imported from Hugging Face seed; review memory and benchmark details before using for final ranking.
Z.ai's GLM 4.7 Flash. Auto-imported from Hugging Face seed; review memory and benchmark details before using for final ranking.