A newer version is available: GLM-5.3View →

GLM-5

MIT

Zhipu AI · 744B (40B active) · Mixture of Experts

MoE with 256 experts, 40B active — frontier-class agentic coding Check if your GPU or Mac can run GLM-5 locally — 415.7 GB min, 692.9 GB recommended.

2026-02128K context

Mixture of Experts

Total experts: 256
Active experts: 8
Active params: 40.0B

Quantization Options

QuantBitsVRAMQualityStatus
Q2_K2238.7 GBlow
Q3_K_M3334 GBmoderate
Q4_K_M4381.6 GBgood
Q5_K_M5476.9 GBgood
Q6_K6572.1 GBexcellent
Q8_08762.7 GBexcellent
F16161524.9 GBlossless

Can I run GLM-5 locally?

Can I run GLM-5 locally?
GLM-5 needs about 415.7 GB of memory at a minimum and 692.9 GB recommended. Open this page to grade it against your GPU or Mac, then run it with runai, Ollama or LM Studio.
How much VRAM does GLM-5 need?
At Q4_K_M, GLM-5 uses about 381.6 GB of VRAM. Higher quants need more memory; lower quants fit tighter cards with a quality tradeoff.