News
Enhancing Quantization Efficiency in Mixture-of-Experts Large Language Models
Abstract
Models based on the Mixture-of-Experts (MoE) architecture demand substantial storage and computational resources, making their compression a pressing challenge. Post-training weight quantization is among the most effective means of reducing model size; however, existing methods such as GPTQ exhibit degraded performance when directly applied to MoE models due to imbalanced expert activation and the specific nature of sparse routing. To address these limitations, this work introduces three modifications. First, a calibration dataset with elevated expert activation entropy is constructed by sampling data from diverse domains. Second, the Hessian matrix is weighted by router coefficients to better capture the contribution of individual experts. Third, an expert sensitivity metric is proposed, enabling 2-bit quantization for the least important experts while retaining 3-bit precision for the remainder. Experiments on the Qwen3-30B-A3B and Mixtral-8x7b model demonstrate that the proposed method outperforms existing approaches in the INT3 regime, and the sensitivity metric further reduces the average bit-width to 2.8 with minimal degradation in text generation quality. These findings confirm the effectiveness of the proposed modifications and pave the way toward more compact deployment of MoE models.
Keywords
Edition
Proceedings of the Institute for System Programming, vol. 38, issue 6, part 1, 2026, pp. 123-138
ISSN 2220-6426 (Online), ISSN 2079-8156 (Print).
DOI: 10.15514/ISPRAS-2026-38(6)-8
For citation
Full text of the paper in pdf (in Russian)
Back to the contents of the volume