---
title: "Quantizing Large Language Models for Code Generation: A Differentiated Replication"
canonical_url: "https://www.modelscope.cn/papers/125103"
md_url: "https://www.modelscope.cn/papers/125103.md"
arxiv_id: 2503.07103
published: 2025-03-10
last_updated: 2025-03-10
authors:
  - "Alessandro Giagnorio"
  - "Antonio Mastropaolo"
  - "Saima Afrin"
  - "Massimiliano Di Penta"
  - "Gabriele Bavota"
model_name: "AQLM (Additive Quantization with Learned Multi-Codebooks)"
model_developer: "SEART Q软件研究所, 威廉玛丽学院, 桑尼奥大学"
domain:
  - "自然语言处理"
  - "软件工程"
  - "深度学习"
type:
  - "自然语言处理"
  - "软件工程"
  - "深度学习"
  - "Software Engineering (cs.SE)"
arxiv_url: "https://arxiv.org/abs/2503.07103"
pdf_url: "https://arxiv.org/pdf/2503.07103.pdf"
---

# Quantizing Large Language Models for Code Generation: A Differentiated Replication

> Large Language Models (LLMs) have shown an impressive capability in code generation and, specifically, to automatically implement requirements described in natural language. The LLM effectiveness generally increases with its size: The higher the number of…

「Quantizing Large Language Models for Code Generation: A Differentiated Replication」是 ModelScope 魔搭社区收录的论文，arXiv 2503.07103，作者为 Alessandro Giagnorio, Antonio Mastropaolo, Saima Afrin et al.，发表于 2025-03-10，属于 自然语言处理、软件工程、深度学习 领域。

- **ArXiv**: 2503.07103
- **Published**: 2025-03-10
- **Authors**: Alessandro Giagnorio, Antonio Mastropaolo, Saima Afrin, Massimiliano Di Penta, Gabriele Bavota
- **Model**: AQLM (Additive Quantization with Learned Multi-Codebooks)
- **Developer**: SEART Q软件研究所, 威廉玛丽学院, 桑尼奥大学
- **Domain**: 自然语言处理, 软件工程, 深度学习
- **ArXiv URL**: https://arxiv.org/abs/2503.07103
- **PDF**: https://arxiv.org/pdf/2503.07103.pdf

Source: https://www.modelscope.cn/papers/125103

---

> 突破极限：新型量化技术使大型语言模型在代码生成中实现极致压缩与高效性能

## 摘要

本文探讨了大型语言模型（LLMs）在代码生成任务中的量化技术。随着LLMs规模的扩大，其内存和碳足迹问题日益显著，限制了实际部署。Wei等人之前的研究表明，通过量化可以减少LLMs的内存占用而不明显降低性能。在此基础上，本文提出了一种差异化的复制研究，使用更近期、更大规模的代码相关LLMs（如CodeLlama和DeepSeek Coder），并应用最新的量化技术，包括极端量化至2位精度。研究结果表明，在4位精度下，量化后的模型平均内存占用减少了70%，且性能没有显著下降。此外，当进一步压缩到3位或2位时，使用特定于代码的校准数据集可以有效减轻性能损失。这项工作不仅验证了先前的研究成果，还展示了最新进展使得极端量化更具吸引力，为未来高效代码生成LLMs奠定了基础。

## Abstract

Large Language Models (LLMs) have shown an impressive capability in code generation and, specifically, to automatically implement requirements described in natural language. The LLM effectiveness generally increases with its size: The higher the number of LLM's trainable parameters the better its ability to implement code. However, when it comes to deploying LLM-based code generators, larger LLMs pose significant challenges related to their memory (and, consequently, carbon) footprint. A previous work by Wei et al. proposed to leverage quantization techniques to reduce the memory footprint of LLM-based code generators without substantially degrading their effectiveness. In short, they studied LLMs featuring up to 16B parameters, quantizing their precision from floating point 32 bits down to int 8 bits and showing their limited impact on code generation performance. Given the fast pace at which LLM capabilities and quantization techniques are evolving, in this work we present a differentiated replication of the work by Wei et al. in which we consider (i) on the one side, more recent and larger code-related LLMs, of up to 34B parameters; (ii) the latest advancements in model quantization techniques, which allow pushing the compression to the extreme quantization level of 2 bits per model parameter and; (iii) different types of calibration datasets to guide the quantization process, including code-specific ones. Our empirical evaluation reveals that the new frontier for LLM quantization is 4-bit precision, resulting in an average memory footprint reduction of 70% compared to the original model without observing any significant decrease in performance. Additionally, when the quantization becomes even more extreme (3 and 2 bits), a code-specific calibration dataset helps to limit the loss of performance.
