Skip to content
View original post on X: Daniel Han· 36/100AI score36/100

GLM-5.3-Flash quantizes to 4-bit with 93% accuracy retained

AISummary

GLM-5.3-Flash (ox-alpha) can be quantized to 4-bit while retaining 93% accuracy, according to Daniel Han. The post says the 4-bit model runs on a 256GB Mac or two DGX Sparks, and 5-bit may also work. Unsloth separately says 3-bit GGUF runs on 128GB RAM and that the model rivals Claude Opus 4.8 on DeepSWE, coding, and agentic benchmarks.

Post on XView on X
@danielhanchen

GLM-5.3-Flash (ox-alpha) can be quantized down to 4-bit and retain 93% accuracy!

The 4-bit model is essentially like a local Claude 4.7 Opus.

It runs perfectly on a 256GB Mac or two DGX Sparks. 5-bit may even work.

Unsloth AI@UnslothAI
GLM-5.3-Flash can now be run locally! ✨ Run 3-bit on 128GB RAM via Unsloth GGUF. GLM-5.3-Flash (ox-alpha) rivals Claude Opus 4.8 on DeepSWE, coding & agentic benchmarks. Guide: https://unsloth.ai/docs/models/glm-5.3-flash GGUF: https://huggingface.co/unsloth/GLM-5.3-Flash-GGUF
View quoted post on X

Source: Daniel Han · x.comPublished · added here