Community EXL3 4-bit quantization shrinks MiniCPM5-2B to 1.61 GB
Original title⚡ Making MiniCPM5-2B even more lightweight for local inference!
AISummary
A community-built 4.0 bpw EXL3 quantization of MiniCPM5-2B reduces the quantized model weights to 1.61 GB for local inference. The author reports roughly 68–70 tokens/s on an NVIDIA Tesla T4, and the model runs with ExLlamaV3 and TabbyAPI.
Source: OpenBMB · x.comPublished · added here