Model Detail Human reviewed

GLM 5.3 Flash

GLM-5.3-Flash is Z.ai's MIT-licensed, open-weight multimodal MoE for coding, long-horizon agents, computer use, and professional work. It has 320B total parameters with 18B activated per token, hybrid linear and sparse attention, native image input, a 1M-token context window, and official support for SGLang, vLLM, and TokenSpeed. Its total weight footprint makes it a server-class local model despite its efficient active-parameter count.

Z.aiGLMmit2026-08-26

Deployment and license note

GLM 5.3 Flash is a server-class local model. Its 18B active parameters reduce per-token compute, but the complete 320B weights still require roughly 384GB for the official FP8 deployment class or substantially more for BF16, before context and KV-cache overhead. Use current SGLang, vLLM, KTransformers, or TokenSpeed releases and plan for multi-GPU serving.

Parameters
320B
18B active
Context window
1M
Long context
Architecture
hybrid linear sparse attention moe
moe
Quality score
94
Planner signal

Task Fit

Code AgentSupported

Tool use, repo work, terminal workflows, and coding benchmarks.

CodeSupported

Code generation, debugging, refactoring, and benchmark signal.

ChatSupported

General writing, Q&A, and assistant use.

RAGSupported

Document QA benefits from long context and instruction following.

VisionSupported

Image or visual understanding, not necessarily image generation.

Image GenerationNot a fit

Not marked for image generation in the current library.

Video GenerationNot a fit

Not marked for video generation in the current library.

VoiceNot a fit

Not marked for voice in the current library.

Source Confidence

Overallhigh · 100/100
ParametersReviewed / seeded
Task fitReviewed / seeded
MemorySeeded artifact
LicenseSource / seed
BenchmarksAvailable
Hardware fitCalculated

Variants and Quant Artifacts

Choose the artifact first; hardware fit follows from RAM, VRAM, format, and runtime.

2 artifacts
QuantFormatQualityMin RAMReco RAMRuntimeAction
FP8safetensorsbalanced384GB512GBtransformers, vllm, sglang, ktransformers, tokenspeed Plan with this
BF16safetensorshigh768GB1TBtransformers, vllm, sglang, ktransformers, tokenspeed Plan with this

Recommended Hardware

No compatible local hardware found in the current hardware pool.

Benchmarks

Terminal Bench 2.184.3%
Toolathlon Verified78.4%
Deep SWE 1.163.4%
Officeqa Pro62.4%
Artificial Analysis Intelligence V4.1.157%
Automation Bench 1.0.648.8%

Source and Review

Hugging Facezai-org/GLM-5.3-Flash
OllamaNot mapped
VerificationHuman reviewed
Artifact sourceofficial-glm-5.3-flash-fp8-weights
Default variantGLM 5.3 Flash Coder
Tool callingSupported

Execution evidence

Run GLM 5.3 Flash with a documented recipe

Recipes connect hardware, a model artifact, tools, settings, verification, and a reportable result.

Browse all recipes →

No verified recipe is linked to this record yet.

Compatibility estimates remain available in the planner. A recipe appears here only after its exact stack and verification protocol are documented.

Similar Models

Continue planning your local AI setup