← Open Source
AI-Efficiency

Awesome-Model-Quantization

A curated collection of papers, benchmarks, surveys, and tools for model quantization, covering low-bit networks, LLMs, multimodal and generative models, vector and lattice quantization, and efficient deployment.

ListsPaper collections
Open on GitHub
Momentum
+0stars in 24 hours0.0%
2.45k
Stars
241
Forks
+1
This week
34
Contributors
Created 2018-10-18 · Updated 2026-10-05 · #8480 today
Top developers
README

Awesome Model Quantization Awesome

Awesome Model Quantization is a curated, continuously updated collection of papers, benchmarks, surveys, and open-source implementations on neural network and model quantization. It spans binary and ternary networks, post-training quantization, quantization-aware training, vector and lattice quantization, low-bit LLMs, multimodal and generative models, KV-cache quantization, low-precision training, and hardware-efficient deployment. The project was initiated by Haotong Qin. Thanks to all contributors for helping grow and maintain this collection.

Quick Navigation

Research Landscape

Model quantization can be organized along five dimensions:

  • Optimization: post-training quantization (PTQ), quantization-aware training (QAT), quantized fine-tuning, data-free methods, and low-precision training.
  • Representation: scalar, vector/codebook, lattice, binary-coded, binary/ternary, and mixed-precision quantization.
  • Error reduction: rotations, outlier smoothing, residual reconstruction, error compensation, and sensitivity-aware methods.
  • Quantized tensors: weights, activations, KV caches, training states, gradients, and communication.
  • Models and deployment: vision, language, multimodal, generative, state space, and graph models, alongside edge and hardware systems.

Methods often combine several dimensions, such as PTQ with rotations and vector codebooks.

🔎 Explore the taxonomy and method connections · Click to expand

Optimization paradigm

Post-Training Quantization (PTQ) converts a pretrained model, often with calibration: GPTQ, SmoothQuant, AWQ, OmniQuant, QuaRot, SpinQuant, FlatQuant, BiLLM. Quantization-Aware Training (QAT) models quantization during optimization: PACT, LSQ, IR-Net. Quantized Fine-Tuning / Parameter-Efficient Fine-Tuning (PEFT) adapts low-bit models: QLoRA, QA-LoRA, LoftQ, IR-QLoRA, L4Q. Data-Free / Zero-Shot Quantization avoids original training data, using model statistics or synthetic samples: ZeroQ, Qimera. Low-Precision Training also reduces precision in training computation or stored states: INT8/FP8 training, 8-bit Optimizers.

Representation / coding structure

Scalar quantization codes individual values; non-uniform, logarithmic, and floating-point quantization change the available levels (AdaLog, LLM-FP4). Vector quantization jointly codes tuples; codebook quantization stores reusable representatives; product / grouped vector quantization partitions vectors into groups (GPTVQ, VPTQ, EPQuant). Lattice quantization uses structured geometric codebooks (QuIP#, NestQuant, grouped lattice vector quantizers). Binary-coded quantization combines binary bases (AnyBCQ); binary / ternary quantization constrains values to two / three levels (IR-Net, BiBERT, PT²-LLM). Mixed precision allocates different bit widths or formats across tensors or groups (HAWQ, SliM-LLM).

Transformation / error handling

Rotation / orthogonal transforms redistribute coordinates (QuaRot, SpinQuant); outlier smoothing / redistribution balances quantization difficulty (SmoothQuant, AWQ). Residual / low-rank reconstruction models remaining errors or outliers (LQER, SVDQuant); error compensation corrects quantization effects (GPTQ, First-Order Error Matters). Saliency-aware / Hessian-aware quantization uses importance or curvature to guide precision, reconstruction, or rounding (HAWQ, GPTQ, BiLLM). These techniques can accompany scalar or structured coding.

Quantized object

Weights (GPTQ, AWQ); activations (PACT); weight + activation (SmoothQuant, BiBERT); KV cache (KIVI, KVQuant, ZipCache, PM-KVQ); training states / optimizer states (ActNN, 8-bit Optimizers); gradients / communication (DoReFa-Net, SDP4Bit). Weight bit width does not imply the same activation, accumulator, or cache precision.

Model family / deployment

CNNs / classical vision (XNOR-Net, BRECQ); Vision Transformers (PTQ4ViT); Large Language Models (GPTQ, QLoRA); multimodal / VLM / VLA (Q-VLM, MQuant, AutoQVLA); diffusion / generative models (PTQ4DM, Q-Diffusion, PTQD, ViDiT-Q, SVDQuant, BinaryDM, Q-VDiT, S²Q-VDiT, QuantSparse); Mamba / state space models (Quamba2, SSDi8); graph / point cloud models (EPQuant, BiPointNet); edge / embedded / hardware-oriented systems (HAQ, FINN, LUT-GEMM).

For the structured-coding lineage, QuIP introduces incoherence processing for low-bit LLM quantization; QuIP# connects this direction to lattice codebooks, while QTIP uses trellis coding. GPTVQ, VPTQ, and NestQuant explore vector or lattice representations. TurboQuant and RaBitQ are also retained for their vector-quantization methodology; RaBitQ targets approximate nearest-neighbor search rather than LLM weight quantization.

Bit width alone does not specify storage overhead, arithmetic precision, or deployment speed.

Representative Works

Selected works grouped by technical approach, with authors and short method summaries. Each paper has one primary home here; the yearly collection includes the wider literature.

Reading the links: Scholar opens a title search on Google Scholar, where citation counts can be viewed. Star badges show the linked GitHub repository’s stars, which may cover several papers. These are discovery aids, not rankings.

Classical Quantization and QAT

From binary weights and learned codebooks to integer arithmetic, learned quantizers and reconstruction-based PTQ.

  • BinaryConnect: Training Deep Neural Networks with binary weights during propagations

    Matthieu Courbariaux, Yoshua Bengio, Jean-Pierre David

    NeurIPS 2015 · Neural Networks QAT Binary Weights · Paper · Code · Scholar GitHub stars

    Trains neural networks with binary weights during forward and backward propagation.

  • XNOR-Net: ImageNet Classification Using Binary Convolutional Neural Networks

    Mohammad Rastegari, Vicente Ordonez, Joseph Redmon, Ali Farhadi

    ECCV 2016 · CNN Binary Weight + Activation · Paper · Code · Scholar GitHub stars

    Approximates convolutions with binary weights and inputs for efficient CNN inference.

  • Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding

    Song Han, Huizi Mao, William J. Dally

    ICLR 2016 · CNN Weight Sharing Codebook · Paper · Scholar

    Combines pruning, trained weight sharing and Huffman coding, connecting learned quantization to compressed model storage.

  • PACT: Parameterized Clipping Activation for Quantized Neural Networks

    Jungwook Choi, Zhuo Wang, Swagath Venkataramani, Pierce I-Jen Chuang, Vijayalakshmi Srinivasan, Kailash Gopalakrishnan

    ICLR 2018 · CNN QAT Activations · Paper · Scholar

    Learns activation clipping thresholds to support low-bit network training.

  • Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference

    Benoit Jacob, Skirmantas Kligys, Bo Chen, Menglong Zhu, Matthew Tang, Andrew Howard, Hartwig Adam, Dmitry Kalenichenko

    CVPR 2018 · QAT INT8 Integer-Only Inference · Paper · Scholar

    Co-designs quantization-aware training and integer arithmetic for mobile inference, including scale and zero-point handling.

  • HAWQ: Hessian AWare Quantization of Neural Networks With Mixed-Precision

    Zhen Dong, Zhewei Yao, Amir Gholami, Michael W. Mahoney, Kurt Keutzer

    ICCV 2019 · Mixed Precision Hessian-Aware · Paper · Scholar

    Uses Hessian information to guide mixed-precision neural network quantization.

  • Learned Step Size Quantization

    Steven K. Esser, Jeffrey L. McKinstry, Deepika Bablani, Rathinakumar Appuswamy, Dharmendra S. Modha

    ICLR 2020 · QAT Low-Bit · Paper · Scholar

    Learns quantizer step sizes alongside network parameters.

  • Up or Down? Adaptive Rounding for Post-Training Quantization

    Markus Nagel, Rana Ali Amjad, Mart van Baalen, Christos Louizos, Tijmen Blankevoort

    ICML 2020 · PTQ Rounding · Paper · Scholar

    Optimizes rounding decisions when converting pretrained weights to low precision.

  • HAWQ-V2: Hessian Aware trace-Weighted Quantization of Neural Networks

    Zhen Dong, Zhewei Yao, Daiyaan Arfeen, Amir Gholami, Michael Mahoney, Kurt Keutzer

    NeurIPS 2020 · Mixed Precision Hessian-Aware · Paper · Scholar

    Develops trace-weighted Hessian sensitivity for mixed-precision allocation.

  • BRECQ: Pushing the Limit of Post-Training Quantization by Block Reconstruction

    Yuhang Li, Ruihao Gong, Xu Tan, Yang Yang, Peng Hu, Qi Zhang, Fengwei Yu, Wei Wang, Shi Gu

    ICLR 2021 · CNN PTQ Reconstruction · Paper · Code · Scholar GitHub stars

    Uses block reconstruction to reduce post-training quantization error.

  • QDrop: Randomly Dropping Quantization for Extremely Low-bit Post-Training Quantization

    Xiuying Wei, Ruihao Gong, Yuhang Li, Xianglong Liu, Fengwei Yu

    ICLR 2022 · PTQ Activations Reconstruction · Paper · Code · Scholar GitHub stars

    Randomly bypasses activation quantization during reconstruction to improve low-bit generalization beyond the calibration data.

Data-Free and Zero-Shot Quantization

These methods replace access to the original dataset with model statistics or generated samples; they may still require calibration or optimization.

  • Data-Free Quantization Through Weight Equalization and Bias Correction

    Markus Nagel, Mart van Baalen, Tijmen Blankevoort, Max Welling

    ICCV 2019 · CNN PTQ Data-Free · Paper · Scholar

    Equalizes channel ranges and corrects quantization-induced bias using model parameters and statistics.

  • ZeroQ: A Novel Zero Shot Quantization Framework

    Yaohui Cai, Zhewei Yao, Zhen Dong, Amir Gholami, Michael W. Mahoney, Kurt Keutzer

    CVPR 2020 · CNN Data-Free Mixed Precision · Paper · Code · Scholar GitHub stars

    Synthesizes calibration inputs from batch-normalization statistics to quantize without the original training dataset.

  • Diversifying Sample Generation for Accurate Data-Free Quantization

    Xiangguo Zhang, Haotong Qin, Yifu Ding, Ruihao Gong, Qinghua Yan, Renshuai Tao, Yuhang Li, Fengwei Yu, Xianglong Liu

    CVPR 2021 · Oral · CNN Data-Free Synthetic Data PTQ · Paper · Scholar

    Relaxes batch-normalization statistic matching and varies layer-wise emphasis to diversify synthetic calibration samples for data-free quantization.

  • Diverse Sample Generation: Pushing the Limit of Generative Data-Free Quantization

    Haotong Qin, Yifu Ding, Xiangguo Zhang, Jiakai Wang, Xianglong Liu, Jiwen Lu

    IEEE TPAMI 2023 · CNN Data-Free PTQ + QAT Sample Diversity · Paper · Code · Scholar GitHub stars

    Extends the CVPR 2021 DSG method with theoretical analysis and inter-sample decorrelation, improving synthetic-data generation for both PTQ and QAT.

  • LLM-QAT: Data-Free Quantization Aware Training for Large Language Models

    Zechun Liu, Barlas Oguz, Changsheng Zhao, Ernie Chang, Pierre Stock, Yashar Mehdad, Yangyang Shi, Raghuraman Krishnamoorthi, Vikas Chandra

    ACL Findings 2024 · LLM QAT Data-Free KV Cache · Paper · Scholar

    Uses the pretrained model’s generated text for distillation-based QAT of weights, activations and KV caches.

Transformer and LLM Quantization

Weight-only compression, weight–activation quantization and QAT address different deployment needs. Sparse outlier handling and rotations offer complementary ways to control error.

  • LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale

    Tim Dettmers, Mike Lewis, Younes Belkada, Luke Zettlemoyer

    NeurIPS 2022 · Transformer INT8 Mixed Precision · Paper · Code · Scholar GitHub stars

    Enables 8-bit matrix multiplication at transformer scale while handling outlier features in higher precision.

  • GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers

    Elias Frantar, Saleh Ashkboos, Torsten Hoefler, Dan Alistarh

    ICLR 2023 · LLM PTQ Weights · Paper · Code · Scholar GitHub stars

    Uses approximate second-order information and error compensation for low-bit weight quantization.

  • SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models

    Guangxuan Xiao, Ji Lin, Mickael Seznec, Hao Wu, Julien Demouth, Song Han

    ICML 2023 · LLM PTQ Weight + Activation · Paper · Code · Scholar GitHub stars

    Redistributes activation outlier difficulty into weights to enable low-precision matrix multiplication.

  • AWQ: Activation-aware Weight Quantization for On-Device LLM Compression and Acceleration

    Ji Lin, Jiaming Tang, Haotian Tang, Shang Yang, Wei-Ming Chen, Wei-Chen Wang, Guangxuan Xiao, Xingyu Dang, Chuang Gan, Song Han

    MLSys 2024 · LLM PTQ Weights Saliency-Aware · Paper · Code · Scholar GitHub stars

    Uses activation information to guide weight quantization for on-device compression and acceleration.

  • OmniQuant: Omnidirectionally Calibrated Quantization for Large Language Models

    Wenqi Shao, Mengzhao Chen, Zhaoyang Zhang, Peng Xu, Lirui Zhao, Zhiqian Li, Kaipeng Zhang, Peng Gao, Yu Qiao, Ping Luo

    ICLR 2024 · LLM PTQ Calibration · Paper · Code · Scholar GitHub stars

    Optimizes clipping and equivalent transformations to calibrate low-bit LLMs.

  • QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs

    Saleh Ashkboos, Amirkeivan Mohtashami, Maximilian L. Croci, Bo Li, Pashmina Cameron, Martin Jaggi, Dan Alistarh, Torsten Hoefler, James Hensman

    NeurIPS 2024 · LLM PTQ 4-Bit Rotation · Paper · Code · Scholar GitHub stars

    Uses rotations to suppress outliers and enable 4-bit inference.

  • SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression

    Tim Dettmers, Ruslan Svirschevski, Vage Egiazarian, Denis Kuznedelev, Elias Frantar, Saleh Ashkboos, Alexander Borzunov, Torsten Hoefler, Dan Alistarh

    ICLR 2024 · LLM PTQ Weights Sparse Outliers · Paper · Code · Scholar GitHub stars

    Separates sensitive outlier weights into a sparse higher-precision component while quantizing the remaining weights.

  • SqueezeLLM: Dense-and-Sparse Quantization

    Sehoon Kim, Coleman Hooper, Amir Gholami, Zhen Dong, Xiuyu Li, Sheng Shen, Michael W. Mahoney, Kurt Keutzer

    ICML 2024 · LLM PTQ Non-uniform Sparse Outliers · Paper · Code · Scholar GitHub stars

    Combines sensitivity-weighted non-uniform scalar quantization with a sparse component for outlier weights.

  • SpinQuant: LLM Quantization with Learned Rotations

    Zechun Liu, Changsheng Zhao, Igor Fedorov, Bilge Soran, Dhruv Choudhary, Raghuraman Krishnamoorthi, Vikas Chandra, Yuandong Tian, Tijmen Blankevoort

    ICLR 2025 · LLM PTQ Learned Rotation · Paper · Code · Scholar GitHub stars

    Learns rotations to make LLM representations more amenable to quantization.

  • FlatQuant: Flatness Matters for LLM Quantization

    Yuxuan Sun, Ruikang Liu, Haoli Bai, Han Bao, Kang Zhao, Yuening Li, Jiaxin Hu, Xianzhi Yu, Lu Hou, Chun Yuan, Xin Jiang, Wulong Liu, Jun Yao

    ICML 2025 · LLM PTQ Transformation · Paper · Code · Scholar GitHub stars

    Targets distribution flatness to improve LLM quantization.

  • EfficientQAT: Efficient Quantization-Aware Training for Large Language Models

    Mengzhao Chen, Wenqi Shao, Peng Xu, Jiahao Wang, Peng Gao, Kaipeng Zhang, Ping Luo

    ACL 2025 · LLM QAT Low-Bit · Paper · Code · Scholar GitHub stars

    Trains block parameters first, then quantization parameters end to end, to reduce the cost of LLM QAT.

Quantized Fine-Tuning

QLoRA and related methods adapt low-bit bases with low-rank updates; PV-Tuning also optimizes discrete compressed representations.

  • QLoRA: Efficient Finetuning of Quantized LLMs

    Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke Zettlemoyer

    NeurIPS 2023 · LLM PEFT 4-Bit · Paper · Code · Scholar GitHub stars

    Fine-tunes low-rank adapters through a frozen 4-bit quantized base model.

  • QA-LoRA: Quantization-Aware Low-Rank Adaptation of Large Language Models

    Yuhui Xu, Lingxi Xie, Xiaotao Gu, Xin Chen, Heng Chang, Hengheng Zhang, Zhengsu Chen, Xiaopeng Zhang, Qi Tian

    ICLR 2024 · LLM PEFT Quantization-Aware · Paper · Code · Scholar GitHub stars

    Combines quantization-aware optimization with low-rank adaptation.

  • LoftQ: LoRA-Fine-Tuning-aware Quantization for Large Language Models

    Yixiao Li, Yifan Yu, Chen Liang, Pengcheng He, Nikos Karampatziakis, Weizhu Chen, Tuo Zhao

    ICLR 2024 · LLM PEFT Low-Bit · Paper · Code · Scholar GitHub stars

    Aligns quantization with LoRA initialization to reduce the error encountered during adaptation.

  • Accurate LoRA-Finetuning Quantization of LLMs via Information Retention

    Haotong Qin, Xudong Ma, Xingyu Zheng, Xiaoyang Li, Yang Zhang, Shouda Liu, Jie Luo, Xianglong Liu, Michele Magno

    ICML 2024 · LLM PEFT Information-Aware · Paper · Code · Scholar GitHub stars

    Uses information retention to improve low-bit quantization and LoRA adaptation.

  • PV-Tuning: Beyond Straight-Through Estimation for Extreme LLM Compression

    Vladimir Malinovskii, Denis Mazur, Ivan Ilin, Denis Kuznedelev, Konstantin Burlachenko, Kai Yi, Dan Alistarh, Peter Richtarik

    NeurIPS 2024 · LLM Quantized Fine-Tuning Discrete Optimization · Paper · Code · Scholar GitHub stars

    Alternates continuous and discrete optimization to fine-tune extremely compressed models, including additive-codebook representations.

  • L4Q: Parameter Efficient Quantization-Aware Fine-Tuning on Large Language Models

    Hyesung Jeon, Yulhwa Kim, Jae-Joon Kim

    ACL 2025 · LLM PEFT QAT · Paper · Scholar

    Combines parameter-efficient fine-tuning with quantization-aware training.

Extreme Low-Bit, Binary and Ternary

Binary CNNs and transformers, post-training binarization, and native ternary pretraining have different training costs and arithmetic requirements.

  • Bi-Real Net: Enhancing the Performance of 1-bit CNNs With Improved Representational Capability and Advanced Training Algorithm

    Zechun Liu, Baoyuan Wu, Wenhan Luo, Xin Yang, Wei Liu, Kwang-Ting Cheng

    ECCV 2018 · CNN Binary QAT · Paper · Code · Scholar GitHub stars

    Connects real-valued intermediate activations through shortcuts to improve information flow in 1-bit CNNs.

  • Forward and Backward Information Retention for Accurate Binary Neural Networks

    Haotong Qin, Ruihao Gong, Xianglong Liu, Mingzhu Shen, Ziran Wei, Fengwei Yu, Jingkuan Song

    CVPR 2020 · CNN QAT Binary 1-Bit · Paper · Code · Scholar GitHub stars

    Retains information in both forward activations and backward gradients when training binary neural networks.

  • ReActNet: Towards Precise Binary Neural Network with Generalized Activation Functions

    Zechun Liu, Zhiqiang Shen, Marios Savvides, Kwang-Ting Cheng

    ECCV 2020 · CNN Binary QAT · Paper · Code · Scholar GitHub stars

    Learns activation shifts and reshaping functions to reduce the accuracy gap between binary and real-valued networks.

  • BiBERT: Accurate Fully Binarized BERT

    Haotong Qin, Yifu Ding, Mingyuan Zhang, Qinghua Yan, Aishan Liu, Qingqing Dang, Ziwei Liu, Xianglong Liu

    ICLR 2022 · Transformer NLP Binary Weight + Activation · Paper · Code · Scholar GitHub stars

    Targets fully binarized BERT, extending binary networks to transformer language models.

  • BiLLM: Pushing the Limit of Post-Training Quantization for LLMs

    Wei Huang, Yangdong Liu, Haotong Qin, Ying Li, Shiming Zhang, Xianglong Liu, Michele Magno, Xiaojuan Qi

    ICML 2024 · LLM PTQ Binary Extreme Low-Bit · Paper · Code · Scholar GitHub stars

    Uses saliency-aware binarization to push pretrained LLM weights into the extreme low-bit regime.

  • DB-LLM: Accurate Dual-Binarization for Efficient LLMs

    Hong Chen, Chengtao Lv, Liang Ding, Haotong Qin, Xiabin Zhou, Yifu Ding, Xuebo Liu, Min Zhang, Jinyang Guo, Xianglong Liu, Dacheng Tao

    ACL Findings 2024 · LLM Dual Binarization Extreme Low-Bit · Paper · Scholar

    Uses dual binarization to compress LLMs while retaining accuracy.

  • The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits

    Shuming Ma, Hongyu Wang, Lingxiao Ma, Lei Wang, Wenhui Wang, Shaohan Huang, Li Dong, Ruiping Wang, Jilong Xue, Furu Wei

    arXiv 2024 · LLM QAT Ternary Weights 8-Bit Activations · Paper · Scholar

    Extends BitNet’s quantization-aware pretraining to ternary weights; this is a training recipe, distinct from post-training binarization.

  • ARB-LLM: Alternating Refined Binarizations for Large Language Models

    Zhiteng Li, Xianglong Yan, Tianao Zhang, Haotong Qin, Dong Xie, Jiang Tian, Zhongchao Shi, Linghe Kong, Yulun Zhang, Xiaokang Yang

    ICLR 2025 · LLM Binary Extreme Low-Bit · Paper · Code · Scholar GitHub stars

    Refines alternating binarizations for low-bit LLM representation.

  • PTQ1.61: Push the Real Limit of Extremely Low-Bit Post-Training Quantization Methods for Large Language Models

    Jiaqi Zhao, Miao Zhang, Ming Wang, Yuzhang Shang, Kaihao Zhang, Weili Guan, Yaowei Wang, Min Zhang

    ACL 2025 · LLM PTQ Extreme Low-Bit · Paper · Code · Scholar GitHub stars

    Explores extremely low-bit post-training quantization for LLMs.

  • PT²-LLM: Post-Training Ternarization for Large Language Models

    Xianglong Yan, Chengzhu Bao, Zhiteng Li, Tianao Zhang, Kaicheng Yang, Haotong Qin, Ruobing Xie, Xingwu Sun, Yulun Zhang

    ICLR 2026 · LLM PTQ Ternary · Paper · Code · Scholar GitHub stars

    Converts pretrained large language models to ternary representations.

Vector, Lattice and Codebook Quantization

From CNN product quantization to LLM additive, lattice and trellis codes. QuIP provides the incoherence-processing precursor; RaBitQ contributes vector-search methodology.

  • Compressing Deep Convolutional Networks using Vector Quantization

    Yunchao Gong, Liu Liu, Ming Yang, Lubomir Bourdev

    arXiv 2014 · CNN Vector Quantization Product Quantization · Paper · Scholar

    Studies clustering and product quantization of CNN parameters as early approaches to reducing model storage.

  • QuIP: 2-Bit Quantization of Large Language Models With Guarantees

    Jerry Chee, Yaohui Cai, Volodymyr Kuleshov, Christopher De Sa

    NeurIPS 2023 · LLM PTQ 2-Bit Incoherence · Paper · Code · Scholar GitHub stars

    Uses incoherence processing for low-bit quantization with guarantees, forming a precursor to the QuIP# lattice-codebook lineage.

  • QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

    Albert Tseng, Jerry Chee, Qingyao Sun, Volodymyr Kuleshov, Christopher De Sa

    ICML 2024 · LLM Lattice Codebook Hadamard · Paper · Code · Scholar GitHub stars

    Combines Hadamard incoherence processing with lattice codebooks for LLM quantization.

  • QTIP: Quantization with Trellises and Incoherence Processing

    Albert Tseng, Qingyao Sun, David Hou, Christopher De Sa

    NeurIPS 2024 · LLM Trellis Coding Incoherence · Paper · Code · Scholar GitHub stars

    Combines trellis-based quantization with incoherence processing for compact LLM representation.

  • GPTVQ: The Blessing of Dimensionality for LLM Quantization

    Mart van Baalen, Andrey Kuzmin, Ivan Koryakovskiy, Markus Nagel, Peter Couperus, Cedric Bastoul, Eric Mahurin, Tijmen Blankevoort, Paul Whatmough

    arXiv 2024 · LLM Vector Quantization Weights · Paper · Code · Scholar GitHub stars

    Exploits joint quantization of multiple weight coordinates rather than coding each weight independently.

  • VPTQ: Extreme Low-bit Vector Post-Training Quantization for Large Language Models

    Yifei Liu, Jicheng Wen, Yang Wang, Shengyu Ye, Li Lyna Zhang, Ting Cao, Cheng Li, Mao Yang

    EMNLP 2024 · LLM PTQ Vector Quantization Extreme Low-Bit · Paper · Code · Scholar GitHub stars

    Uses vector post-training quantization for extremely low-bit LLM compression.

  • RaBitQ: Quantizing High-Dimensional Vectors with a Theoretical Error Bound for Approximate Nearest Neighbor Search

    Jianyang Gao, Cheng Long

    SIGMOD 2024 · Vector Quantization Binary Codes Vector Search · Paper · Code · Scholar GitHub stars

    Quantizes high-dimensional vectors with a theoretical error bound for approximate nearest-neighbor search.

  • Extreme Compression of Large Language Models via Additive Quantization

    Vage Egiazarian, Andrei Panferov, Denis Kuznedelev, Elias Frantar, Artem Babenko, Dan Alistarh

    ICML 2024 · LLM PTQ Additive Codebooks 2–3 Bit · Paper · Code · Scholar GitHub stars

    Represents weight vectors as sums of learned codewords and jointly optimizes codebooks within transformer blocks.

  • NestQuant: nested lattice quantization for matrix products and LLMs

    Semyon Savkin, Eitan Porat, Or Ordentlich, Yury Polyanskiy

    ICML 2025 · LLM Lattice Matrix Products · Paper · Scholar

    Uses nested lattice quantization for matrix products and LLMs.

  • Learning Grouped Lattice Vector Quantizers for Low-Bit Large Language Models

    Xi Zhang, Xiaolin Wu, Jiamang Wang, Weisi Lin

    NeurIPS 2025 · LLM Grouped Vector Quantization Lattice · Paper · Scholar

    Learns grouped lattice vector quantizers for low-bit LLM representation.

  • AnyBCQ: Hardware Efficient Flexible Binary-Coded Quantization for Multi-Precision LLMs

    Gunho Park, Jeongin Bae, Beomseok Kwon, Byeongwook Kim, Se Jung Kwon, Dongsoo Lee

    ICLR 2026 · LLM Binary-Coded Mixed Precision Hardware · Paper · Code · Scholar GitHub stars

    Develops flexible binary-coded quantization for hardware-efficient multi-precision LLMs.

  • TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate

    Amir Zandieh, Majid Daliri, Majid Hadian, Vahab Mirrokni

    ICLR 2026 · Vector Quantization Online Distortion · Paper · Scholar

    Studies online vector quantization with near-optimal distortion rate.

KV Cache Quantization

These methods compress inference-time key and value tensors; their bit widths are separate from model weight precision.

  • KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache

    Zirui Liu, Jiayi Yuan, Hongye Jin, Shaochen Zhong, Zhaozhuo Xu, Vladimir Braverman, Beidi Chen, Xia Hu

    ICML 2024 · LLM KV Cache 2-Bit · Paper · Code · Scholar GitHub stars

    Uses asymmetric, tuning-free 2-bit quantization to compress key and value caches.

  • KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

    Coleman Hooper, Sehoon Kim, Hiva Mohammadzadeh, Michael Mahoney, Sophia Shao, Kurt Keutzer, Amir Gholami

    NeurIPS 2024 · LLM KV Cache Long Context · Paper · Code · Scholar GitHub stars

    Targets long-context inference by reducing the memory occupied by the KV cache.

  • ZipCache: Accurate and Efficient KV Cache Quantization with Salient Token Identification

    Yefei He, Luoming Zhang, Weijia Wu, Jing Liu, Hong Zhou, Bohan Zhuang

    NeurIPS 2024 · LLM KV Cache Salient Tokens · Paper · Code · Scholar GitHub stars

    Uses salient-token identification to guide accurate and efficient cache quantization.

  • PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs

    Tengxuan Liu, Shiyao Li, Jiayi Yang, Tianchen Zhao, Feng Zhou, Xiaohui Song, Guohao Dai, Shengen Yan, Huazhong Yang, Yu Wang

    ICLR 2026 · LLM KV Cache Mixed Precision · Paper · Code · Scholar GitHub stars

    Progressively quantizes KV caches with mixed precision for long chain-of-thought inference.

Diffusion and Generative Model Quantization

Early diffusion PTQ addresses denoising-step sensitivity; later work extends to diffusion transformers, low-rank outlier handling and video generation.

  • Post-training Quantization on Diffusion Models

    Yuzhang Shang, Zhihang Yuan, Bin Xie, Bingzhe Wu, Yan Yan

    CVPR 2023 · Diffusion PTQ · Paper · Code · Scholar GitHub stars

    Adapts post-training quantization to diffusion model inference.

  • Q-diffusion: Quantizing Diffusion Models

    Xiuyu Li, Yijiang Liu, Long Lian, Huanrui Yang, Zhen Dong, Daniel Kang, Shanghang Zhang, Kurt Keutzer

    ICCV 2023 · Diffusion PTQ · Paper · Code · Scholar GitHub stars

    Quantizes diffusion models to reduce the cost of iterative generation.

  • PTQD: Accurate Post-Training Quantization for Diffusion Models

    Yefei He, Luping Liu, Jing Liu, Weijia Wu, Hong Zhou, Bohan Zhuang

    NeurIPS 2023 · Diffusion PTQ Error Handling · Paper · Code · Scholar GitHub stars

    Targets accurate diffusion generation through post-training quantization error handling.

  • ViDiT-Q: Efficient and Accurate Quantization of Diffusion Transformers for Image and Video Generation

    Tianchen Zhao, Tongcheng Fang, Haofeng Huang, Rui Wan, Widyadewi Soedarmadji, Enshu Liu, Shiyao Li, Zinan Lin, Guohao Dai, Shengen Yan, Huazhong Yang, Xuefei Ning, Yu Wang

    ICLR 2025 · Diffusion Transformer Image + Video Low-Bit · Paper · Code · Scholar GitHub stars

    Quantizes diffusion transformers for both image and video generation.

  • SVDQuant: Absorbing Outliers by Low-Rank Component for 4-Bit Diffusion Models

    Muyang Li, Yujun Lin, Zhekai Zhang, Tianle Cai, Xiuyu Li, Junxian Guo, Enze Xie, Chenlin Meng, Jun-Yan Zhu, Song Han

    ICLR 2025 · Diffusion 4-Bit Low-Rank · Paper · Code · Scholar GitHub stars

    Absorbs outliers into a low-rank component to support 4-bit diffusion models.

  • BinaryDM: Accurate Weight Binarization for Efficient Diffusion Models

    Xingyu Zheng, Xianglong Liu, Haotong Qin, Xudong Ma, Mingyuan Zhang, Haojie Hao, Jiakai Wang, Zixiang Zhao, Jinyang Guo, Michele Magno

    ICLR 2025 · Diffusion Binary Weights · Paper · Code · Scholar GitHub stars

    Binarizes diffusion model weights for efficient generation.

  • Q-VDiT: Towards Accurate Quantization and Distillation of Video-Generation Diffusion Transformers

    Weilun Feng, Chuanguang Yang, Haotong Qin, Xiangqi Li, Yu Wang, Zhulin An, Libo Huang, Boyu Diao, Zixiang Zhao, Yongjun Xu, Michele Magno

    ICML 2025 · Video Diffusion Quantization Distillation · Paper · Code · Scholar GitHub stars

    Combines quantization and distillation for video-generation diffusion transformers.

  • S²Q-VDiT: Accurate Quantized Video Diffusion Transformer with Salient Data and Sparse Token Distillation

    Weilun Feng, Haotong Qin, Chuanguang Yang, Xiangqi Li, Han Yang, Yuqi Li, Zhulin An, Libo Huang, Michele Magno, Yongjun Xu

    NeurIPS 2025 · Video Diffusion Quantization Distillation · Paper · Code · Scholar GitHub stars

    Uses salient data and sparse-token distillation to improve quantized video diffusion transformers.

  • QuantSparse: Comprehensively Compressing Video Diffusion Transformer with Model Quantization and Attention Sparsification

    Weilun Feng, Chuanguang Yang, Haotong Qin, Mingqiang Wu, Yuqi Li, Xiangqi Li, Zhulin An, Libo Huang, Yulun Zhang, Michele Magno, Yongjun Xu

    ICLR 2026 · Video Diffusion Quantization Attention Sparsity · Paper · Code · Scholar GitHub stars

    Combines model quantization and attention sparsification to compress video diffusion transformers.

Multimodal and State Space Models

Vision-language models and selective state space models introduce quantization sensitivities beyond those of language-only transformers.

  • Q-VLM: Post-training Quantization for Large Vision-Language Models

    Changyuan Wang, Ziwei Wang, Xiuwei Xu, Yansong Tang, Jie Zhou, Jiwen Lu

    NeurIPS 2024 · VLM PTQ Cross-Layer Dependency · Paper · Code · Scholar GitHub stars

    Uses cross-layer dependencies to guide block partitioning and quantization of vision-language models.

  • Quamba2: A Robust and Scalable Post-training Quantization Framework for Selective State Space Models

    Hung-Yueh Chiang, Chi-Chih Chang, Natalia Frumkin, Kai-Chiang Wu, Mohamed S. Abdelfattah, Diana Marculescu

    ICML 2025 · Mamba State Space Models PTQ W4A8 / W8A8 · Paper · Code · Scholar GitHub stars

    Uses channel clustering and state-group quantization to accommodate the sensitivity of Mamba’s selective state-space computations.

Vision, Edge and Hardware

Vision methods and deployment systems connect quantizer design to integer kernels, memory movement and hardware costs.

  • FINN: A Framework for Fast, Scalable Binarized Neural Network Inference

    Yaman Umuroglu, Nicholas J. Fraser, Giulio Gambardella, Michaela Blott, Philip Leong, Magnus Jahre, Kees Vissers

    FPGA 2017 · Binary Networks FPGA Inference · Paper · Code · Scholar GitHub stars

    Provides a framework for fast, scalable binarized neural network inference on FPGA hardware.

  • HAQ: Hardware-Aware Automated Quantization with Mixed Precision

    Kuan Wang, Zhijian Liu, Yujun Lin, Ji Lin, Song Han

    CVPR 2019 · CNN Mixed Precision Hardware-Aware · Paper · Code · Scholar GitHub stars

    Automates mixed-precision quantization with hardware deployment costs in view.

  • BiPointNet: Binary Neural Network for Point Clouds

    Haotong Qin, Zhongang Cai, Mingyuan Zhang, Yifu Ding, Haiyu Zhao, Shuai Yi, Xianglong Liu, Hao Su

    ICLR 2021 · Point Clouds Binary QAT 1-Bit · Paper · Code · Scholar GitHub stars

    Uses entropy-maximizing aggregation and layer-wise scale recovery to address feature homogenization and scale distortion in binary point-cloud networks.

  • PTQ4ViT: Post-Training Quantization for Vision Transformers with Twin Uniform Quantization

    Zhihang Yuan, Chenhao Xue, Yiqi Chen, Qiang Wu, Guangyu Sun

    ECCV 2022 · Vision Transformer PTQ · Paper · Code · Scholar GitHub stars

    Uses twin uniform quantization to support post-training compression of vision transformers.

  • QuantSR: Accurate Low-bit Quantization for Efficient Image Super-Resolution

    Haotong Qin, Yulun Zhang, Yifu Ding, Yifan Liu, Xianglong Liu, Martin Danelljan, Fisher Yu

    NeurIPS 2023 · Super-Resolution QAT 2–4 Bit · Paper · Code · Scholar GitHub stars

    Combines a redistribution-driven learnable quantizer with a depth-dynamic architecture for accurate low-bit image super-resolution.

  • LUT-GEMM: Quantized Matrix Multiplication based on LUTs for Efficient Inference in Large-Scale Generative Language Models

    Gunho Park, Baeseong Park, Minsub Kim, Sungjae Lee, Jeonghoon Kim, Beomseok Kwon, Se Jung Kwon, Byeongwook Kim, Youngjoo Lee, Dongsoo Lee

    ICLR 2024 · LLM Quantized Matrix Multiplication Lookup Tables · Paper · Scholar

    Uses lookup tables for efficient quantized matrix multiplication in generative language models.

  • Quant-LLM: Accelerating the Serving of Large Language Models via FP6-Centric Algorithm-System Co-Design on Modern GPUs

    Haojun Xia, Zhen Zheng, Xiaoxia Wu, Shiyang Chen, Zhewei Yao, Stephen Youn, Arash Bakhtiari, Michael Wyatt, Donglin Zhuang, Zhongzhu Zhou, Olatunji Ruwase, Yuxiong He, Shuaiwen Leon Song

    USENIX ATC 2024 · LLM FP6 GPU Kernels · Paper · Code · Scholar GitHub stars

    Uses TC-FPx kernels to support non-power-of-two weight formats efficiently on GPUs; the codebase is also known as FP6-LLM.

  • QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving

    Yujun Lin, Haotian Tang, Shang Yang, Zhekai Zhang, Guangxuan Xiao, Chuang Gan, Song Han

    MLSys 2025 · LLM W4A8KV4 GPU Serving · Paper · Code · Scholar GitHub stars

    Co-designs progressive quantization, attention and GPU kernels to turn reduced precision into serving throughput.

Floating-Point and Microscaling Formats

Low-bit floating-point and shared-scale formats complement integer quantization. Format design and model calibration are separate choices.

  • FP8 Formats for Deep Learning

    Paulius Micikevicius, Dusan Stosic, Neil Burgess, Marius Cornea, Pradeep Dubey, Richard Grisenthwaite, Sangwon Ha, Alexander Heinecke, Patrick Judd, John Kamalu, Naveen Mellempudi, Stuart Oberman, Mohammad Shoeybi, Michael Siu, Hao Wu

    arXiv 2022 · FP8 E4M3 / E5M2 Training + Inference · Paper · Scholar

    Defines complementary FP8 encodings and evaluates their use in neural network training and inference.

  • Microscaling Data Formats for Deep Learning

    Bita Darvish Rouhani, Ritchie Zhao, Ankit More, Mathew Hall, Alireza Khodamoradi, Summer Deng, Dhruv Choudhary, Marius Cornea, Eric Dellinger, Kristof Denolf, Stosic Dusan, Venmugil Elango, Maximilian Golub, Alexander Heinecke, Phil James-Roxby, Dharmesh Jani, Gaurav Kolhe, Martin Langhammer, Ada Li, Levi Melnick, Maral Mesmakhosroshahi, Andres Rodriguez, Michael Schulte, Rasoul Shafipour, Lei Shao, Michael Siu, Pradeep Dubey, Paulius Micikevicius, Maxim Naumov, Colin Verrilli, Ralph Wittig, Doug Burger, Eric Chung

    arXiv 2023 · MX Formats Block Scaling Training + Inference · Paper · Code · Scholar GitHub stars

    Combines shared block scales with narrow element formats to balance numerical range and hardware efficiency.

  • LLM-FP4: 4-Bit Floating-Point Quantized Transformers

    Shih-yang Liu, Zechun Liu, Xijie Huang, Pingcheng Dong, Kwang-Ting Cheng

    EMNLP 2023 · LLM PTQ FP4 · Paper · Code · Scholar GitHub stars

    Searches exponent configurations and quantization parameters to handle weight and activation range differences in 4-bit floating point.

Low-Precision Training and States

Quantization can reduce saved activations, optimizer states, gradient communication or training arithmetic; each targets a different part of the training cost.

  • QSGD: Communication-Efficient SGD via Gradient Quantization and Encoding

    Dan Alistarh, Demjan Grubic, Jerry Li, Ryota Tomioka, Milan Vojnovic

    NeurIPS 2017 · Training Gradients Communication · Paper · Scholar

    Uses randomized gradient quantization with convergence guarantees to trade communication bandwidth against estimator variance.

  • ActNN: Reducing Training Memory Footprint via 2-Bit Activation Compressed Training

    Jianfei Chen, Lianmin Zheng, Zhewei Yao, Dequan Wang, Ion Stoica, Michael Mahoney, Joseph Gonzalez

    ICML 2021 · Training Activations 2-Bit · Paper · Code · Scholar GitHub stars

    Compresses saved activations to reduce the memory footprint of neural network training.

  • 8-bit Optimizers via Block-wise Quantization

    Tim Dettmers, Mike Lewis, Sam Shleifer, Luke Zettlemoyer

    ICLR 2022 · Training Optimizer States 8-Bit · Paper · Code · Scholar GitHub stars

    Uses block-wise quantization to reduce optimizer-state memory.

  • SDP4Bit: Toward 4-bit Communication Quantization in Sharded Data Parallelism for LLM Training

    Jinda Jia, Cong Xie, Hanlin Lu, Daoce Wang, Hao Feng, Chengming Zhang, Baixi Sun, Haibin Lin, Zhi Zhang, Xin Liu, Dingwen Tao

    NeurIPS 2024 · LLM Training Communication 4-Bit · Paper · Code · Scholar GitHub stars

    Targets 4-bit communication quantization in sharded data-parallel LLM training.

  • Optimizing Large Language Model Training Using FP4 Quantization

    Ruizhe Wang, Yeyun Gong, Xiao Liu, Guoshuai Zhao, Ziyue Yang, Baining Guo, Zhengjun Zha, Peng Cheng

    ICML 2025 · LLM Training FP4 Gradient Estimation · Paper · Scholar

    Combines differentiable quantization estimation with outlier handling to stabilize FP4 LLM training.

Benchmarks

Choose by evaluation scope: deployment reproducibility, binary networks, LLM capabilities or robustness. Expand a resource below for authors, figures and citation details.

Resource What it covers
MQBench: Towards Reproducible and Deployable Model Quantization Benchmark
NeurIPS 2021 Datasets and Benchmarks
Code · Scholar GitHub stars QAT + deployment
Compares quantization algorithms under reproducible settings and hardware backend constraints.
BiBench: Benchmarking and Analyzing Network Binarization
ICML 2023
Code · Scholar GitHub stars Binary networks
Compares binarization methods across tasks, architectures and deployment settings.
Evaluating Quantized Large Language Models
ICML 2024
Code · Scholar GitHub stars Weights, activations + KV cache
Evaluates 11 model families on basic NLP, emergent abilities, trustworthiness, dialogue and long-context tasks.
LLMC: Benchmarking Large Language Model Quantization with a Versatile Compression Toolkit
EMNLP 2024 Industry Track
Code · Scholar GitHub stars LLM toolkit
Compares calibration data, method pipelines and quantization configurations; the toolkit is now LightCompress.
An empirical study of LLaMA3 quantization: from LLMs to MLLMs
Visual Intelligence 2024
Code · Scholar GitHub stars LLMs + multimodal
Examines low-bit behavior across LLaMA3 language and multimodal models.
An Empirical Study of Qwen3 Quantization
Visual Intelligence 2026
Code · Scholar GitHub stars Dense + MoE LLMs
Studies quantization across Qwen3 model sizes, architectures and reasoning settings.
RobustMQ: Benchmarking Robustness of Quantized Models
Visual Intelligence 2023
Scholar Model robustness
Tests quantized models beyond clean accuracy, including robustness under input perturbations.

MQBench · Authors and BibTeX

Yuhang Li, Mingzhu Shen, Jian Ma, Yan Ren, Mingxin Zhao, Qi Zhang, Ruihao Gong, Fengwei Yu, Junjie Yan

@inproceedings{li2021mqbench,
  title={MQBench: Towards Reproducible and Deployable Model Quantization Benchmark},
  author={Li, Yuhang and Shen, Mingzhu and Ma, Jian and Ren, Yan and Zhao, Mingxin and Zhang, Qi and Gong, Ruihao and Yu, Fengwei and Yan, Junjie},
  booktitle={NeurIPS Datasets and Benchmarks},
  year={2021}
}

BiBench · Authors, overview and BibTeX

Haotong Qin, Mingyuan Zhang, Yifu Ding, Aoyu Li, Zhongang Cai, Ziwei Liu, Fisher Yu, Xianglong Liu

BiBench: benchmarking binary neural networks

@inproceedings{qin2023bibench,
  title={BiBench: Benchmarking and Analyzing Network Binarization},
  author={Qin, Haotong and Zhang, Mingyuan and Ding, Yifu and Li, Aoyu and Cai, Zhongang and Liu, Ziwei and Yu, Fisher and Liu, Xianglong},
  booktitle={International Conference on Machine Learning (ICML)},
  year={2023}
}

QLLM-Eval · Authors and BibTeX

Shiyao Li, Xuefei Ning, Luning Wang, Tengxuan Liu, Xiangsheng Shi, Shengen Yan, Guohao Dai, Huazhong Yang, Yu Wang

@inproceedings{li2024evaluating,
  title={Evaluating Quantized Large Language Models},
  author={Li, Shiyao and Ning, Xuefei and Wang, Luning and Liu, Tengxuan and Shi, Xiangsheng and Yan, Shengen and Dai, Guohao and Yang, Huazhong and Wang, Yu},
  booktitle={International Conference on Machine Learning},
  year={2024},
  url={https://proceedings.mlr.press/v235/li24bb.html}
}

LLMC · Authors, overview and BibTeX

Ruihao Gong, Yang Yong, Shiqiao Gu, Yushi Huang, Chengtao Lv, Yunchen Zhang, Dacheng Tao, Xianglong Liu

LLMC quantization benchmark and toolkit

@inproceedings{gong2024llmc,
  title={Llmc: Benchmarking large language model quantization with a versatile compression toolkit},
  author={Gong, Ruihao and Yong, Yang and Gu, Shiqiao and Huang, Yushi and Lv, Chengtao and Zhang, Yunchen and Tao, Dacheng and Liu, Xianglong},
  booktitle={Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: Industry Track},
  pages={132--152},
  year={2024}
}

LLaMA3 study · Authors, overview and BibTeX

Wei Huang, Xingyu Zheng, Xudong Ma, Haotong Qin, Chengtao Lv, Hong Chen, Jie Luo, Xiaojuan Qi, Xianglong Liu, Michele Magno

LLaMA3 Quantization Benchmark

@article{huang2024empirical,
  title={An empirical study of llama3 quantization: From llms to mllms},
  author={Huang, Wei and Zheng, Xingyu and Ma, Xudong and Qin, Haotong and Lv, Chengtao and Chen, Hong and Luo, Jie and Qi, Xiaojuan and Liu, Xianglong and Magno, Michele},
  journal={Visual Intelligence},
  volume={2},
  number={1},
  pages={36},
  year={2024},
  publisher={Springer}
}

Qwen3 study · Authors, overview and BibTeX

Xingyu Zheng, Yuye Li, Haoran Chu, Yue Feng, Xudong Ma, Zining Wang, Jie Luo, Jinyang Guo, Haotong Qin, Michele Magno, Xianglong Liu

Preprint

Qwen3 quantization empirical study

@article{zheng2026empirical,
  title={An empirical study of Qwen3 quantization},
  author={Zheng, Xingyu and Li, Yuye and Chu, Haoran and Feng, Yue and Ma, Xudong and Wang, Zining and Luo, Jie and Guo, Jinyang and Qin, Haotong and Magno, Michele and Liu, Xianglong},
  journal={Visual Intelligence},
  volume={4},
  pages={11},
  year={2026},
  doi={10.1007/s44267-026-00114-4}
}

RobustMQ · Authors, overview and BibTeX

Yisong Xiao, Aishan Liu, Tianyuan Zhang, Haotong Qin, Jinyang Guo, Xianglong Liu

RobustMQ: robustness of quantized models

@article{xiao2023robustmq,
  title={Robustmq: benchmarking robustness of quantized models},
  author={Xiao, Yisong and Liu, Aishan and Zhang, Tianyuan and Qin, Haotong and Guo, Jinyang and Liu, Xianglong},
  journal={Visual Intelligence},
  volume={1},
  number={1},
  pages={30},
  year={2023},
  publisher={Springer}
}

Survey Papers

Start with the white paper for practical PTQ/QAT, then choose a survey for broader context or a specific model family. Figures and citation details are available below.

Resource What it covers
A White Paper on Neural Network Quantization
arXiv 2021
Scholar Practical PTQ + QAT
Explains quantizer design, common failure modes and practical post-training and quantization-aware training workflows.
A Survey of Quantization Methods for Efficient Neural Network Inference
arXiv 2021
Scholar Foundations + taxonomy
Reviews quantization design choices, mixed precision and the trade-offs between model accuracy and efficient inference.
Binary Neural Networks: A Survey
Pattern Recognition 2020
Scholar Binary networks
Surveys binary network representations, training methods and applications.
A Survey of Low-bit Large Language Models: Basics, Systems, and Algorithms
Neural Networks 2025
Scholar LLM algorithms + systems
Connects low-bit LLM algorithms with numerical formats and inference systems.
Low-bit Model Quantization for Deep Neural Networks: A Survey
arXiv 2025
Scholar Broad low-bit methods
Maps low-bit quantization methods across neural network architectures and applications.

Quantization white paper · Authors and BibTeX

Markus Nagel, Marios Fournarakis, Rana Ali Amjad, Yelysei Bondarenko, Mart van Baalen, Tijmen Blankevoort

@article{nagel2021white,
  title={A White Paper on Neural Network Quantization},
  author={Nagel, Markus and Fournarakis, Marios and Amjad, Rana Ali and Bondarenko, Yelysei and van Baalen, Mart and Blankevoort, Tijmen},
  journal={arXiv preprint arXiv:2106.08295},
  year={2021}
}

Quantization methods survey · Authors and BibTeX

Amir Gholami, Sehoon Kim, Zhen Dong, Zhewei Yao, Michael W. Mahoney, Kurt Keutzer

Chapter

@article{gholami2021survey,
  title={A Survey of Quantization Methods for Efficient Neural Network Inference},
  author={Gholami, Amir and Kim, Sehoon and Dong, Zhen and Yao, Zhewei and Mahoney, Michael W. and Keutzer, Kurt},
  journal={arXiv preprint arXiv:2103.13630},
  year={2021}
}

Binary networks survey · Authors, overview and BibTeX

Haotong Qin, Ruihao Gong, Xianglong Liu, Xiao Bai, Jingkuan Song, Nicu Sebe

Blog

Binary Neural Networks survey overview

@article{Qin:pr20_bnn_survey,
    title = "Binary neural networks: A survey",
    author = "Haotong Qin and Ruihao Gong and Xianglong Liu and Xiao Bai and Jingkuan Song and Nicu Sebe",
    journal = "Pattern Recognition",
    volume = "105",
    pages = "107281",
    year = "2020"
}

Low-bit LLM survey · Authors, overview and BibTeX

Ruihao Gong, Yifu Ding, Zining Wang, Chengtao Lv, Xingyu Zheng, Jinyang Du, Yang Yong, Shiqiao Gu, Haotong Qin, Jinyang Guo, Dahua Lin, Michele Magno, Xianglong Liu

A Survey of Low-bit Large Language Models

@article{gong2025survey,
  title={A survey of low-bit large language models: Basics, systems, and algorithms},
  author={Gong, Ruihao and Ding, Yifu and Wang, Zining and Lv, Chengtao and Zheng, Xingyu and Du, Jinyang and Yong, Yang and Gu, Shiqiao and Qin, Haotong and Guo, Jinyang and Lin, Dahua and Magno, Michele and Liu, Xianglong},
  journal={Neural Networks},
  pages={107856},
  year={2025}
}

Low-bit model survey · Authors, overview and BibTeX

Kai Liu, Qian Zheng, Kaiwen Tao, Zhiteng Li, Haotong Qin, Wenbo Li, Yong Guo, Xianglong Liu, Linghe Kong, Guihai Chen, Yulun Zhang, Xiaokang Yang

Low-bit model quantization survey overview

@article{liu2025low,
  title={Low-bit Model Quantization for Deep Neural Networks: A Survey},
  author={Liu, Kai and Zheng, Qian and Tao, Kaiwen and Li, Zhiteng and Qin, Haotong and Li, Wenbo and Guo, Yong and Liu, Xianglong and Kong, Linghe and Chen, Guihai and Zhang, Yulun and Yang, Xiaokang},
  journal={arXiv preprint arXiv:2505.05530},
  year={2025}
}

Papers by Year

All paper titles and links are kept in this README. Published work is grouped by venue year where verified; otherwise the recorded preprint year is used. Representative works, benchmarks and surveys also appear here for chronological browsing. Within each year, entries are grouped by conference or journal, with preprints at the end.

2026

  • [AAAI] First-Order Error Matters: Accurate Compensation for Quantized Large Language Models [code] GitHub stars
  • [AAAI] TR-DQ: Time-Rotation Diffusion Quantization
  • [ICLR] PT²-LLM: Post-Training Ternarization for Large Language Models [code] GitHub stars
  • [ICLR] Quant-dLLM: Post-Training Extreme Low-Bit Quantization for Diffusion Large Language Models
  • [ICLR] DVD-Quant: Data-free Video Diffusion Transformers Quantization
  • [ICLR] Q&C: When Quantization Meets Cache in Efficient Generation
  • [ICLR] Quantized Visual Geometry Grounded Transformer
  • [ICLR] Post-Training Quantization for Video Matting
  • [ICLR] QVGen: Pushing the Limit of Quantized Video Generative Models
  • [ICLR] QuantSparse: Comprehensively Compressing Video Diffusion Transformer with Model Quantization and Attention Sparsification [code] GitHub stars
  • [ICLR] TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate
  • [ICLR] Optimal Brain Restoration for Joint Quantization and Sparsification of LLMs [code] GitHub stars
  • [ICLR] AnyBCQ: Hardware Efficient Flexible Binary-Coded Quantization for Multi-Precision LLMs [code] GitHub stars
  • [ICLR] Tequila: Deadzone-free Ternary Quantization for Large Language Models
  • [ICLR] LogART: Pushing the Limit of Efficient Logarithmic Post-Training Quantization [code] GitHub stars
  • [ICLR] ParoQuant: Pairwise Rotation Quantization for Efficient Reasoning LLM Inference [code] GitHub stars
  • [ICLR] Improving Block-Wise LLM Quantization by 4-bit Generalized Normal Float Formats
  • [ICLR] Channel-Aware Mixed-Precision Quantization for Efficient Long-Context Inference
  • [ICLR] CodeQuant: Unified Clustering and Quantization for Enhanced Outlier Smoothing in Low-Precision Mixture-of-Experts
  • [ICLR] QeRL: Beyond Efficiency - Quantization-enhanced Reinforcement Learning for LLMs [code] GitHub stars
  • [ICLR] AutoQVLA: Not All Channels Are Equal in Vision-Language-Action Model's Quantization
  • [ICLR] Achieving low-bit Muon through subspace preservation and grid quantization
  • [ICLR] Shift-and-Sum Quantization for Visual Autoregressive Models
  • [ICLR] Inlier-Centric Post-Training Quantization for Object Detection Models
  • [ICLR] Efficient Quantization of Mixture-of-Experts with Theoretical Generalization Guarantees
  • [ICLR] BBQ: Boosting Quantization Entropy with Bell Box Quantization
  • [ICLR] Improving Block-Wise LLM Quantization by 4-bit Block-Wise Optimal Float (BOF4): Analysis and Variations [code] GitHub stars
  • [ICLR] Learning under Quantization for High-Dimensional Linear Regression
  • [ICLR] On-the-Fly Adaptation to Quantization: Configuration-Aware LoRA for Efficient Fine-Tuning of Quantized LLMs
  • [ICLR] Bridging the Gap Between Promise and Performance for FP4 Quantization [code] GitHub stars
  • [ICLR] KBVQ-MoE: KLT-guided SVD with Bias-Corrected Vector Quantization for MoE Large Language Models [code] GitHub stars
  • [ICLR] UniQL: Unified Quantization and Low-rank Compression for Adaptive Edge LLMs [code] GitHub stars
  • [ICLR] The Lattice Geometry of Neural Network Quantization: A Short Equivalence Proof of GPTQ and Babai's algorithm
  • [ICLR] DPQuant: Efficient and Private Model Training via Dynamic Quantization Scheduling
  • [ICLR] Towards Quantization-Aware Training for Ultra-Low-Bit Reasoning LLMs
  • [ICLR] A Convergence Analysis of Adaptive Optimizers under Floating-point Quantization
  • [ICLR] Training Dynamics Impact Post-Training Quantization Robustness [code] GitHub stars
  • [ICLR] SSDi8: Accurate and Efficient 8-bit Quantization for State Space Duality
  • [ICLR] The Geometry of LLM Quantization: GPTQ as Babai's Nearest Plane Algorithm
  • [ICLR] PTQ4ARVG: Post-Training Quantization for AutoRegressive Visual Generation Models [code] GitHub stars
  • [ICLR] QWHA: Quantization-Aware Walsh-Hadamard Adaptation for Parameter-Efficient Fine-Tuning on Large Language Models [code] GitHub stars
  • [ICLR] Gradient-Aligned Calibration for Post-Training Quantization of Diffusion Models
  • [ICLR] SERQ: Saliency-Aware Low-Rank Error Reconstruction for LLM Quantization
  • [ICLR] Compute-Optimal Quantization-Aware Training
  • [ICLR] PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs [code] GitHub stars
  • [ICLR] Beyond Outliers: A Study of Optimizers Under Quantization
  • [ICLR] Qronos: Correcting the Past by Shaping the Future... in Post-Training Quantization
  • [ICLR] MicroMix: Efficient Mixed-Precision Quantization with Microscaling Formats for Large Language Models [code] GitHub stars
  • [ICLR] TurboBoA: Faster and Exact Attention-aware Quantization without Backpropagation
  • [ICLR] Beyond Uniformity: Sample and Frequency Meta Weighting for Post-Training Quantization of Diffusion Models
  • [ICLR] Rethinking Residual Errors in Compensation-based LLM Quantization
  • [ICLR] SPR²Q: Static Priority-based Rectifier Routing Quantization for Image Super-Resolution [code] GitHub stars
  • [ICLR] STaMP: Sequence Transformation and Mixed Precision for Low-Precision Activation Quantization
  • [CVPR Findings] Q-MambaIR: Accurate Quantized Mamba for Efficient Image Restoration
  • [ICCAD] Interface-Aware KV Cache Quantization for Dense On-Chip NVM in Long-Context LLM Decoding [Scholar]
  • [EMNLP] All for 1-Bit: Towards Genuine 1-Bit Post-Training Quantization for LLMs [Scholar]
  • [EMNLP] Colla-Q: Toward Collaborative Experts in MoE Quantization via Minimax Precision Balancing [code] GitHub stars [Scholar]
  • [NeurIPS] SPHQuant: Efficient extreme low bit weight quantization for Vision-Language Models [official repository; code pending] GitHub stars [Scholar]
  • [NeurIPS] D²Quant: Accurate Low-bit Post-Training Weight Quantization for LLMs [code] GitHub stars [Scholar]
  • [NeurIPS] OSCAR: Offline Spectral Covariance-Aware Rotation for 2-bit KV Cache Quantization [code] GitHub stars [Scholar]
  • [NeurIPS] HESTIA: A Hessian-Guided Differentiable Quantization-Aware Training Framework for Extremely Low-Bit LLMs [code] GitHub stars [Scholar]
  • [NeurIPS] KroQuant: Kronecker-Structured Block Transforms for Efficient Post-Training Quantization of Diffusion Transformers [Scholar]
  • [NeurIPS] AlphaQ: Calibration-Free Bit Allocation for Mixture-of-Experts Quantization [code] GitHub stars [Scholar]
  • [NeurIPS] MoBiQuant: Mixture-of-Bits Quantization for Token-Adaptive Any-Precision LLM [Scholar]
  • [NeurIPS] Normalized Architectures are Natively 4-Bit [code] GitHub stars [Scholar]
  • [Visual Intelligence] An Empirical Study of Qwen3 Quantization [code] GitHub stars [arXiv]
  • [arXiv] EdgeRazor: A Lightweight Framework for Large Language Models via Mixed-Precision Quantization-Aware Distillation [code] GitHub stars
  • [arXiv] Sherry: Hardware-Efficient 1.25-Bit Ternary Quantization via Fine-grained Sparsification [code] GitHub stars
  • [arXiv] QuantLRM: Quantization of Large Reasoning Models via Fine-Tuning Signals
  • [arXiv] SliderQuant: Accurate Post-Training Quantization for LLMs
  • [arXiv] What Makes Low-Bit Quantization-Aware Training Work for Reasoning LLMs? A Systematic Study
  • [arXiv] OneComp: One-Line Revolution for Generative AI Model Compression [Code] GitHub stars

2025

  • [AAAI] MPQ-DM: Mixed Precision Quantization for Extremely Low Bit Diffusion Models
  • [AAAI] JAQ: Joint Efficient Architecture Design and Low-Bit Quantization
  • [AAAI] OAC: Output-adaptive Calibration for Accurate Post-Training Quantization of LLMs
  • [AAAI] Optimizing Quantized Diffusion Models via Distillation with Decay Timestep-Aware Loss
  • [AAAI] Quantifiable Quantization Sensitivity of Diffusion Models
  • [AAAI] TCAQ-DM: Timestep-Channel Adaptive Quantization for Diffusion Models
  • [AAAI] Thinking in Granularity: Dynamic Quantization for Image Super-Resolution by Intriguing Multi-Granularity Clues [code] GitHub stars
  • [AAAI] D2-DPM: Dual Denoising for Quantized Diffusion Probabilistic Models [code] GitHub stars
  • [ICLR] ARB-LLM: Alternating Refined Binarizations for Large Language Models [code] GitHub stars
  • [ICLR] BinaryDM: Accurate Weight Binarization for Efficient Diffusion Models [code] GitHub stars
  • [ICLR] CBQ: Cross-Block Quantization for Large Language Models
  • [ICLR] DGQ: Distribution-Aware Group Quantization for Text-to-Image Diffusion Models
  • [ICLR] LeanQuant: Accurate and Scalable Large Language Model Quantization with Loss-error-aware Grid
  • [ICLR] OSTQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting [code] GitHub stars
  • [ICLR] QERA: an Analytical Framework for Quantization Error Reconstruction [code] GitHub stars
  • [ICLR] SpinQuant: LLM Quantization with Learned Rotations [code] GitHub stars
  • [ICLR] SVDQuant: Absorbing Outliers by Low-Rank Component for 4-Bit Diffusion Models [code] GitHub stars
  • [ICLR] ViDiT-Q: Efficient and Accurate Quantization of Diffusion Transformers for Image and Video Generation [code] GitHub stars
  • [ICLR] SynQ: Accurate Zero-shot Quantization by Synthesis-aware Fine-tuning [code] GitHub stars
  • [MLSys] QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving [code] GitHub stars
  • [CVPR] PassionSR: Post-Training Quantization with Adaptive Scale in One-Step Diffusion based Image Super-Resolution [code] GitHub stars
  • [CVPR] Quantization without Tears
  • [CVPR] APHQ-ViT: Post-Training Quantization with Average Perturbation Hessian Based Reconstruction for Vision Transformer [code] GitHub stars
  • [SIGMOD] Practical and Asymptotically Optimal Quantization of High-Dimensional Vectors in Euclidean Space for Approximate Nearest Neighbor Search [code] GitHub stars
  • [ICML] Q-VDiT: Towards Accurate Quantization and Distillation of Video-Generation Diffusion Transformers [code] GitHub stars
  • [ICML] SliM-LLM: Salience-Driven Mixed-Precision Quantization for Large Language Models [code] GitHub stars
  • [ICML] FlatQuant: Flatness Matters for LLM Quantization [code] GitHub stars
  • [ICML] RoSTE: An Efficient Quantization-Aware Supervised Fine-Tuning Approach for Large Language Models [code] GitHub stars
  • [ICML] GANQ: GPU-Adaptive Non-Uniform Quantization for Large Language Models
  • [ICML] Modulated Diffusion: Accelerating Generative Modeling with Modulated Quantization [code] GitHub stars
  • [ICML] GuidedQuant: Large Language Model Quantization via Exploiting End Loss Guidance [code] GitHub stars
  • [ICML] ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals [code] GitHub stars
  • [ICML] MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design [code] GitHub stars
  • [ICML] Learning from Loss Landscape: Generalizable Mixed-Precision Quantization via Adaptive Sharpness-Aware Gradient Aligning
  • [ICML] PARQ: Piecewise-Affine Regularized Quantization [code] GitHub stars
  • [ICML] Quamba2: A Robust and Scalable Post-training Quantization Framework for Selective State Space Models [code] GitHub stars
  • [ICML] LRA-QViT: Integrating Low-Rank Approximation and Quantization for Robust and Efficient Vision Transformers
  • [ICML] BoA: Attention-aware Post-training Quantization without Backpropagation
  • [ICML] MoEQuant: Enhancing Quantization for Mixture-of-Experts Large Language Models via Expert-Balanced Sampling and Affinity Guidance [code] GitHub stars
  • [ICML] NestQuant: nested lattice quantization for matrix products and LLMs
  • [ICML] Q-resafe: Assessing Safety Risks and Quantization-aware Safety Patching for Quantized Large Language Models [code] GitHub stars
  • [ICML] SLiM: One-shot Quantization and Sparsity with Low-rank Approximation for LLM Weight Compression [code] GitHub stars
  • [ICML] QT-DoG: Quantization-Aware Training for Domain Generalization [code] GitHub stars
  • [ICML] Matryoshka Quantization
  • [ICML] Merge-Friendly Post-Training Quantization for Multi-Target Domain Adaptation [code] GitHub stars
  • [ICML] Layer-wise Quantization for Quantized Optimistic Dual Averaging
  • [ICML] Outlier-Aware Post-Training Quantization for Discrete Graph Diffusion Models
  • [ICML] BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference
  • [ICML] GPTAQ: Efficient Finetuning-Free Quantization with Asymmetric Calibration [code] GitHub stars
  • [ICML] Optimizing Large Language Model Training Using FP4 Quantization
  • [ICML] SKIM: Any-bit Quantization Pushing The Limits of Post-Training Quantization
  • [ICML] SageAttention2: Efficient Attention with Thorough Outlier Smoothing and Per-thread INT4 Quantization [code] GitHub stars
  • [ACL] EfficientQAT: Efficient Quantization-Aware Training for Large Language Models [code] GitHub stars
  • [ACL] L4Q: Parameter Efficient Quantization-Aware Fine-Tuning on Large Language Models
  • [ACL] MoQAE: Mixed-Precision Quantization for Long-Context LLM Inference via Mixture of Quantization-Aware Experts
  • [ACL] Outlier-Safe Pre-Training for Robust 4-Bit Quantization of Large Language Models
  • [ACL] PTQ1.61: Push the Real Limit of Extremely Low-Bit Post-Training Quantization Methods for Large Language Models [code] GitHub stars
  • [ACL] Unifying Uniform and Binary-coding Quantization for Accurate Compression of Large Language Models
  • [ACL] “Give Me BF16 or Give Me Death”? Accuracy-Performance Trade-Offs in LLM Quantization
  • [ACL Findings] Achieving Binary Weight and Activation for LLMs using Post-Training Quantization
  • [ICCV] Scheduling Weight Transitions for Quantization-Aware Training [code] GitHub stars
  • [ICCV] Task-Specific Zero-shot Quantization-Aware Training for Object Detection [code] GitHub stars
  • [ICCV] OuroMamba: A Data-Free Quantization Framework for Vision Mamba
  • [ICCV] FedWSQ: Efficient Federated Learning with Weight Standardization and Distribution-Aware Non-Uniform Quantization [code] GitHub stars
  • [ICCV] Semantic Alignment and Reinforcement for Data-Free Quantization of Vision Transformers [code] GitHub stars
  • [ICCV] QuantCache: Adaptive Importance-Guided Quantization with Hierarchical Latent and Layer Caching for Video Generation [code] GitHub stars
  • [ICCV] MixA-Q: Revisiting Activation Sparsity for Vision Transformers from a Mixed-Precision Quantization Perspective
  • [ICCV] DMQ: Dissecting Outliers of Diffusion Models for Post-Training Quantization [code] GitHub stars
  • [ICCV] AHCPTQ: Accurate and Hardware-Compatible Post-Training Quantization for Segment Anything Model
  • [ICCV] MSQ: Memory-Efficient Bit Sparsification Quantization
  • [ICCV] QuEST: Low-bit Diffusion Model Quantization via Efficient Selective Finetuning [code] GitHub stars
  • [ACM MM] DilateQuant: Accurate and Efficient Quantization-Aware Training for Diffusion Models via Weight Dilation
  • [ACM MM] Learning Binarized Representations with Pseudo-positive Distillation
  • [ACM MM] MQuant: Unleashing the Inference Potential of Multimodal Large Language Models with Post-Training Quantization
  • [ACM MM] Pushing the Limit of Binarized Neural Network for Image Super Resolution with Smooth Information Transmission
  • [ACM MM] Quantization Meets OOD: Generalizable Quantization-aware Training from a Flatness Perspective
  • [EMNLP] AMQ: Enabling AutoML for Mixed-precision Weight-Only Quantization of Large Language Models
  • [EMNLP] Does quantization affect models' performance on long-input and long-output tasks?
  • [EMNLP Findings] KurTail: Kurtosis-based LLM Quantization
  • [NeurIPS] S²Q-VDiT: Accurate Quantized Video Diffusion Transformer with Salient Data and Sparse Token Distillation [code] GitHub stars
  • [NeurIPS] DartQuant: Efficient Rotational Distribution Calibration for LLM Quantization [code] GitHub stars
  • [NeurIPS] A Double Norm