Introducing GLM-4.6-REAP-252B-A32B, a memory-efficient compressed variant of GLM-4.6 that maintains near-identical performance while being 30% lighter.

The original model available here: https://huggingface.co/cerebras/GLM-4.6-REAP-252B-A32B

Note: this is a MXFP4-MOE accurate downstream low-bit quantization version. It has been tested with the latest version of llama.cpp.

bitcoin:

bc1q6nvh39fcmy0de0ezepnn2z0rn4dme9yjal77ah
Downloads last month
20
GGUF
Model size
252B params
Architecture
glm4moe
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for andyjack/GLM-4.6-REAP-252B-A32B-MXFP4_MOE_GGUF

Base model

zai-org/GLM-4.6
Quantized
(39)
this model

Collection including andyjack/GLM-4.6-REAP-252B-A32B-MXFP4_MOE_GGUF