A code-specialized MoE carved out of Qwen3.6-35B-A3B by pure expert pruning β no fine-tuning, no distillation. I profiled all 256 experts on balanced corpora plus targeted code benchmarks (LiveCodeBench +
MultiPL-E), built a competence map with the code classes up-weighted 1.5Γ, and dropped the 72 weakest experts per layer (256β184, ~35Bβ27B). Router, attention, norms, the MTP head and the vision tower are all
preserved; active params stay at A3B and routing is baked to top-10 (revert to top-8 anytime).
π Benchmarks (Q6_K, temp 0.6):
β’ MultiPL-E 0.840
β’ HumanEval 0.970
β’ LiveCodeBench 0.688
β’ GSM8K 0.970 Β· ARC-C 0.944 Β· AIME 0.733
β’ GPQA-Diamond 0.773 Β· MATH-500 0.620 Β· IFEval 0.730
β’ Average 0.808
27B footprint, A3B speed, coding that punches well above its size β and the preserved MTP head gives you speculative decoding out of the box (text + vision).
π Model: ManniX-ITA/Qwen3.6-27B-A3B-Coder
π¦ GGUF (+MTP): ManniX-ITA/Qwen3.6-27B-A3B-Coder-MTP-GGUF
π¦ Ollama: https://ollama.com/mannix/qwen3.6-27b-a3b-coder