Awesome LLM Compression
[](https://awesome.re)



Awesome LLM compression research papers and tools to accelerate LLM training and inference.
Papers in each category are grouped by year in collapsible blocks, newest first — click a year to expand it. The current year is expanded by default.
Contents
Papers
Survey
2026 · 6 papers
-
Dynamic Model Routing and Cascading for Efficient LLM Inference: A Survey
Arxiv 2026 [Paper] -
A Survey of On-Policy Distillation for Large Language Models
Arxiv 2026 [Paper] -
Unifying Data, Memory, and Compute Efficiency in LLM training: A Survey
TAI 2026 [Paper] -
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving
Arxiv 2026 [Paper] -
Towards Efficient Large Language Model Serving: A Survey on System-Aware KV Cache Optimization
ACL Findings 2026 [Paper] -
Accelerating Masked Diffusion Large Language Models: A Survey of Efficient Inference Techniques
IJCAI-ECAI 2026 (Survey Track) [Paper]
2025 · 14 papers
-
Compressed but Compromised? A Study of Jailbreaking in Compressed LLMs
NeurIPS Lock-LLM Workshop 2025 [Paper] [Blog] -
Scaling Down, Serving Fast: Compressing and Deploying Efficient LLMs for Recommendation Systems
EMNLP 2025 [Paper] -
Key, Value, Compress: A Systematic Exploration of KV Cache Compression Techniques
CICC 2025 [Paper] -
Are We There Yet? A Measurement Study of Efficiency for LLM Applications on Mobile Devices
Arxiv 2025 [Paper] -
Knowledge Distillation and Dataset Distillation of Large Language Models: Emerging Trends, Challenges, and Future Directions
Arxiv 2025 [Paper] -
An Empirical Study on Prompt Compression for Large Language Models
Building Trust Workshop @ ICLR 2025 2025 [Paper] [PCToolkit] -
Low-Precision Training of Large Language Models: Methods, Challenges, and Opportunities
Arxiv 2025 [Paper] [GitHub Page] -
Optimizing LLMs for Resource-Constrained Environments: A Survey of Model Compression Techniques
COMPSAC 2025 [Paper] -
A Survey on Inference Engines for Large Language Models: Perspectives on Optimization and Efficiency
Arxiv 2025 [Paper] [GitHub Page] -
EfficientLLM: Efficiency in Large Language Models
Arxiv 2025 [Paper] [Homepage] [Huggingface Page] -
KV Cache Compression for Inference Efficiency in LLMs: A Review
Arxiv 2025 [Paper] -
A Survey on Parallel Text Generation: From Parallel Decoding to Diffusion Language Models
Arxiv 2025 [Paper] [Code] -
A Survey on Collaborating Small and Large Language Models for Performance, Cost-effectiveness, Cloud-edge Privacy, and Trustworthiness
Arxiv 2025 [Paper] [Code] -
A Survey on Efficient Large Language Model Training: From Data-centric Perspectives
ACL 2025 [Paper] [Code]
2024 · 17 papers
-
Understanding LLMs: A Comprehensive Overview from Training to Inference
Arxiv 2024 [Paper] -
Faster and Lighter LLMs: A Survey on Current Challenges and Way Forward
IJCAI 2024 (Survey Track) [Paper] [GitHub Page] -
A Survey of Resource-efficient LLM and Multimodal Foundation Models
Arxiv 2024 [Paper] -
A Survey on Hardware Accelerators for Large Language Models
Arxiv 2024 [Paper] -
A Comprehensive Survey of Compression Algorithms for Language Models
Arxiv 2024 [Paper] -
A Survey on Transformer Compression
Arxiv 2024 [Paper] -
Model Compression and Efficient Inference for Large Language Models: A Survey
Arxiv 2024 [Paper] -
LLM Inference Unveiled: Survey and Roofline Model Insights
Arxiv 2024 [Paper] -
A Survey on Knowledge Distillation of Large Language Models
Arxiv 2024 [Paper] [GitHub Page] -
Efficient Prompting Methods for Large Language Models: A Survey
Arxiv 2024 [Paper] -
Survey on Knowledge Distillation for Large Language Models: Methods, Evaluation, and Application
Arxiv 2024 [Paper] -
On-Device Language Models: A Comprehensive Review
Arxiv 2024 [Paper] [Download On-device LLMs] -
A Survey of Low-bit Large Language Models: Basics, Systems, and Algorithms
Arxiv 2024 [Paper] -
Contextual Compression in Retrieval-Augmented Generation for Large Language Models: A Survey
Arxiv 2024 [Paper] -
Prompt Compression for Large Language Models: A Survey
Arxiv 2024 [Paper] -
A Comprehensive Study on Quantization Techniques for Large Language Models
Arxiv 2024 [Paper] -
A Survey on Large Language Model Acceleration based on KV Cache Management
TMLR 2025 [Paper]
2023 · 5 papers
-
A Survey on Model Compression for Large Language Models
TACL [Paper] -
The Cost of Compression: Investigating the Impact of Compression on Parametric Knowledge in Language Models
EMNLP 2023 [Paper] [Code] -
The Efficiency Spectrum of Large Language Models: An Algorithmic Survey
Arxiv 2023 [Paper] -
Efficient Large Language Models: A Survey
TMLR [Paper] [GitHub Page] -
Towards Efficient Generative Large Language Model Serving: A Survey from Algorithms to Systems
ICML 2024 Tutorial [Paper] [Tutorial]
Quantization
2026 · 56 papers
-
Can Large Language Models Still Explain Themselves? Investigating the Impact of Quantization on Self-Explanations
EMNLP Findings 2026 [Paper] -
QSLM: A Performance- and Memory-aware Quantization Framework with Tiered Search Strategy for Spike-driven Language Models
DATE 2026 [Paper] -
HAS-VQ: Hessian-Adaptive Sparse Vector Quantization for High-Fidelity LLM Compression
Arxiv 2026 [Paper] [Code] -
ARCQuant: Boosting NVFP4 Quantization with Augmented Residual Channels for LLMs
ACL 2026 [Paper] -
Sherry: Hardware-Efficient 1.25-Bit Ternary Quantization via Fine-grained Sparsification
Arxiv 2026 [Paper] [Code] -
Calibrating Beyond English: Language Diversity for Better Quantized Multilingual LLM
EACL 2026 [Paper] -
M2XFP: A Metadata-Augmented Microscaling Data Format for Efficient Low-bit Quantization
ASPLOS 2026 [Paper] -
Quartet II: Accurate LLM Pre-Training in NVFP4 by Improved Unbiased Gradient Estimation
Arxiv 2026 [Paper] [Code] -
Two-Stage Grid Optimization for Group-wise Quantization of LLMs
ICASSP 2026 [Paper] -
Quantized Evolution Strategies: High-precision Fine-tuning of Quantized LLMs at Low-precision Cost
Arxiv 2026 [Paper] [Code] -
QVLA: Not All Channels Are Equal in Vision-Language-Action Model's Quantization
ICLR 2026 [Paper] -
TurboBoA: Faster and Exact Attention-aware Quantization without Backpropagation
ICLR 2026 [Paper] -
RaBiT: Residual-Aware Binarization Training for Accurate and Efficient LLMs
ICML 2026 [Paper] -
NanoQuant: Efficient Sub-1-Bit Quantization of Large Language Models
ICML 2026 [Paper] -
On the Importance of a Multi-Scale Calibration for Quantization
ICASSP 2026 [Paper] -
QTALE: Quantization-Robust Token-Adaptive Layer Execution for LLMs
ICML 2026 [Paper] -
QuRL: Efficient Reinforcement Learning with Quantized Rollout
ICLR 2026 [Paper] -
SPQ: An Ensemble Technique for Large Language Model Compression
LREC 2026 [Paper] [Code] -
Quant Experts: Token-aware Adaptive Error Reconstruction with Mixture of Experts for Large Vision-Language Models Quantization
CVPR 2026 [Paper] -
MASQuant: Modality-Aware Smoothing Quantization for Multimodal Large Language Models
CVPR 2026 [Paper] -
SliderQuant: Accurate Post-Training Quantization for LLMs
ICLR 2026 [Paper] [Code] -
OneComp: One-Line Revolution for Generative AI Model Compression
Arxiv 2026 [Paper] [Code] -
Fast NF4 Dequantization Kernels for Large Language Model Inference
ASPLOS 2026 Workshop [Paper] -
RUQuant: Towards Refining Uniform Quantization for Large Language Models
KDD 2026 [Paper] -
SEPTQ: A Simple and Effective Post-Training Quantization Paradigm for Large Language Models
KDD 2025 [Paper] -
ReSpinQuant: Efficient Layer-Wise LLM Quantization via Subspace Residual Rotation Approximation
ICML 2026 [Paper] -
Robust Ultra Low-Bit Post-Training Quantization via Stable Diagonal Curvature Estimate
MLSys 2026 [Paper] -
AQPIM: Breaking the PIM Capacity Wall for LLMs with In-Memory Activation Quantization
HPCA 2026 [Paper] -
From Signal Degradation to Computation Collapse: Uncovering the Two Failure Modes of LLM Quantization
ACL Findings 2026 [Paper] -
Statistically-Lossless Quantization of Large Language Models
Arxiv 2026 [Paper] [Code] -
EdgeRazor: A Lightweight Framework for Large Language Models via Mixed-Precision Quantization-Aware Distillation
Arxiv 2026 [Paper] [Code] [Model] [Playground] -
OSAQ: Outlier Self-Absorption for Accurate Low-bit LLM Quantization
ICML 2026 [Paper] -
Normalized Architectures are Natively 4-Bit
Arxiv 2026 [Paper] [Code] -
RQ-MoE: Residual Quantization via Mixture of Experts for Efficient Input-Dependent Vector Compression
ICML 2026 [Paper] -
XFP: Quality-Targeted Adaptive Codebook Quantization with Sparse Outlier Separation for LLM Inference
Arxiv 2026 [Paper] [Code] -
GEMQ: Global Expert-Level Mixed-Precision Quantization for MoE LLMs
ICML 2026 [Paper] [Code] -
EVA: Accelerating LLM Decoding via an Efficient Vector Quantization Architecture
ISCA 2026 [Paper] -
Influence-Inspired Spectral Rotations for Extreme Low-Bit LLM Quantization
Arxiv 2026 [Paper] [Code] -
InfoQuant: Shaping Activation Distributions for Low-Bit LLM Quantization
Arxiv 2026 [Paper] [Code] -
LFQ: Logit-aware Final-block Quantization for Boosting the Generation Quality of Low-Bit Quantized LLMs
ICML 2026 [Paper] -
LiftQuant: Continuous Bit-Width LLM via Dimensional Lifting and Projection
ICML 2026 [Paper] -
UniSVQ: 2-bit Unified Scalar-Vector Quantization
ICML 2026 [Paper] -
LC-QAT: Data-Efficient 2-Bit QAT for LLMs via Linear-Constrained Vector Quantization
ICML 2026 [Paper] -
TWLA: Achieving Ternary Weights and Low-Bit Activations for LLMs via Post-Training Quantization
ICML 2026 [Paper] -
CAT-Q: Cost-efficient and Accurate Ternary Quantization for LLMs
ICML 2026 [Paper] [Code] -
Quantize the Target, Quantize the Drafter: Efficient Inference with Qwen3.5-4B
Arxiv 2026 [Paper] [Code] -
KronQ: LLM Quantization via Kronecker-Factored Hessian
COLM 2026 [Paper] -
Rethinking Small VLM Quantization: From Component-Wise Analysis to Hardware-Aware Edge Deployment
ICML 2026 Workshop [Paper] -
Reliability Scaling Laws for Quantized Large Language Models
TMLR 2026 [Paper] -
PolyQ: Codesigning End-to-End Quantization Framework for Scalable Edge CPU LLM Inference
ICCAD 2026 [Paper] -
ExaGEMM: Exploration Framework for CPU-Driven ML Inference via Associative In-Register Computing for Low-Bit GEMM
ICCAD 2026 [Paper] -
GyRot: Leveraging Hidden Synergy between Rotation and Fine-grained Group Quantization for Low-bit LLM Inference
HPCA 2026 [Paper] -
Attend to Your Own Thoughts: Breaking the Barrier for Post-Training Quantization of Reasoning LLMs through the Lens of 1.58-Bit Quantization
Arxiv 2026 [Paper] [Code] -
ARCHead: Activation-Metric Residual Correction for Large Language Model Output Heads
Arxiv 2026 [Paper] [Code] -
Compress and Forget: bitsandbytes Quantization Amplifies Proactive Interference in LLMs
Arxiv 2026 [Paper] [Code] -
Jacobian-guided Noise Injection for Quantization Robustness in Large Language Models
ICML 2026 [Paper]
2025 · 118 papers
-
HALO: Hadamard-Assisted Lower-Precision Optimization for LLMs
Arxiv 2025 [Paper] -
RaZeR: Pushing the Limits of NVFP4 Quantization with Redundant Zero Remapping
Arxiv 2025 [Paper] -
FlexQuant: Elastic Quantization Framework for Locally Hosted LLM on Edge Devices
Arxiv 2025 [Paper] -
Rethinking Post-Training Quantization: Introducing a Statistical Pre-Calibration Approach
Arxiv 2025 [Paper] -
Qrazor: Reliable and effortless 4-bit llm quantization by significant data razoring
Arxiv 2025 [Paper] -
OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting
ICLR 2025 [Paper] [Code] -
SwiftPrune: Hessian-Free Weight Pruning for Large Language Models
Arxiv 2025 [Paper] -
Progressive Binarization with Semi-Structured Pruning for LLMs
Arxiv 2025 [Paper] [Code] -
Physics-Inspired Binary Neural Networks: Interpretable Compression with Theoretical Guarantees
Arxiv 2025 [Paper] -
QuEST: Stable Training of LLMs with 1-Bit Weights and Activations
Arxiv 2025 [Paper] -
ParetoQ: Scaling Laws in Extremely Low-bit LLM Quantization
NeurIPS 2025 [Paper] -
Systematic Outliers in Large Language Models
ICLR 2025 [Paper] [Code] -
Can Post-Training Quantization Benefit from an Additional QLoRA Integration?
NAACL 2025 [Paper] -
1bit-Merging: Dynamic Quantized Merging for Large Language Models
Arxiv 2025 [Paper] -
Towards Efficient Pre-training: Exploring FP4 Precision in Large Language Models
Arxiv 2025 [Paper] -
Continual Quantization-Aware Pre-Training: When to transition from 16-bit to 1.58-bit pre-training for BitNet language models?
Arxiv 2025 [Paper] -
QuZO: Quantized Zeroth-Order Fine-Tuning for Large Language Models
Arxiv 2025 [Paper] -
Benchmarking Post-Training Quantization in LLMs: Comprehensive Taxonomy, Unified Evaluation, and Comparative Analysis
Arxiv 2025 [Paper] -
Compression Scaling Laws:Unifying Sparsity and Quantization
Arxiv 2025 [Paper] -
M-ANT: Efficient Low-bit Group Quantization for LLMs via Mathematically Adaptive Numerical Type
Arxiv 2025 [Paper] -
Identifying Sensitive Weights via Post-quantization Integral
Arxiv 2025 [Paper] -
RSQ: Learning from Important Tokens Leads to Better Quantized LLMs
Arxiv 2025 [Paper] [Code] -
VQ-LLM: High-performance Code Generation for Vector Quantization Augmented LLM Inference
HPCA 2025 [Paper] -
Universality of Layer-Level Entropy-Weighted Quantization Beyond Model Architecture and Size
Arxiv 2025 [Paper] -
Towards Superior Quantization Accuracy: A Layer-sensitive Approach
Arxiv 2025 [Paper] -
MergeQuant: Accurate 4-bit Static Quantization of Large Language Models by Channel-wise Calibration
Arxiv 2025 [Paper] -
ClusComp: A Simple Paradigm for Model Compression and Efficient Finetuning
Arxiv 2025 [Paper] -
DynaMo: Runtime Switchable Quantization for MoE with Cross-Dataset Adaptation
DATE 2026 [Paper] -
Cocktail: Chunk-Adaptive Mixed-Precision Quantization for Long-Context LLM Inference
DATE 2025 [Paper] -
GPTQv2: Efficient Finetuning-Free Quantization for Asymmetric Calibration
ICML 2025 [Paper] [Code] -
Task-Circuit Quantization: Leveraging Knowledge Localization and Interpretability for Compression
COLM 2025 [Paper] [Code] -
Quantization Error Propagation: Revisiting Layer-Wise Post-Training Quantization
NeurIPS 2025 [Paper] [Code] -
RaanA: A Fast, Flexible, and Data-Efficient Post-Training Quantization Algorithm
Arxiv 2025 [Paper] [Code] -
Achieving binary weight and activation for LLMs using Post-Training Quantization
Arxiv 2025 [Paper] -
DL-QAT: Weight-Decomposed Low-Rank Quantization-Aware Training for Large Language Models
EMNLP 2024 [Paper] -
Gradual Binary Search and Dimension Expansion : A general method for activation quantization in LLMs
Arxiv 2025 [Paper] -
FGMP: Fine-Grained Mixed-Precision Weight and Activation Quantization for Hardware-Accelerated LLM Inference
Arxiv 2025 [Paper] -
BitNet v2: Native 4-bit Activations with Hadamard Transformation for 1-bit LLMs
Arxiv 2025 [Paper] -
FineQ: Software-Hardware Co-Design for Low-Bit Fine-Grained Mixed-Precision Quantization of LLMs
DATE 2025 [Paper] -
Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models
Arxiv 2025 [Paper] -
ICQuant: Index Coding enables Low-bit LLM Quantization
Arxiv 2025 [Paper] -
Radio: Rate-Distortion Optimization for Large Language Model Compression
ICML 2025 [Paper] -
Balancing Fidelity and Plasticity: Aligning Mixed-Precision Fine-Tuning with Linguistic Hierarchies
Arxiv 2025 [Paper] -
MoEQuant: Enhancing Quantization for Mixture-of-Experts Large Language Models via Expert-Balanced Sampling and Affinity Guidance
ICML 2025 [Paper] -
Grouped Sequency-arranged Rotation: Optimizing Rotation Transformation for Quantization for Free
Arxiv 2025 [Paper] -
Improving Block-Wise LLM Quantization by 4-bit Block-Wise Optimal Float (BOF4): Analysis and Variations
Arxiv 2025 [Paper] -
GuidedQuant: Large Language Model Quantization via Exploiting End Loss Guidance
ICML 2025 [Paper] [Code] -
QuantX: A Framework for Hardware-Aware Quantization of Generative AI Workloads
Arxiv 2025 [Paper] -
An Extra RMSNorm is All You Need for Fine Tuning to 1.58 Bits
Arxiv 2025 [Paper] -
ITERA-LLM: Boosting Sub-8-Bit Large Language Model Inference via Iterative Tensor Decomposition
Arxiv 2025 [Paper] -
Fine-tuning Quantized Neural Networks with Zeroth-order Optimization
ICLR 2026 [Paper] -
Scaling Law for Quantization-Aware Training
Arxiv 2025 [Paper] -
Quaff: Quantized Parameter-Efficient Fine-Tuning under Outlier Spatial Stability Hypothesis
ACL 2025 [Paper] [Code] -
Is (Selective) Round-To-Nearest Quantization All You Need?
Arxiv 2025 [Paper] -
NeUQI: Near-Optimal Uniform Quantization Parameter Initialization
ICML 2026 [Paper] -
LoTA-QAF: Lossless Ternary Adaptation for Quantization-Aware Fine-Tuning
Arxiv 2025 [Paper] [Code] -
FP4 All the Way: Fully Quantized Training of LLMs
Arxiv 2025 [Paper] [Code] -
FireQ: Fast INT4-FP8 Kernel and RoPE-aware Quantization for LLM Inference Acceleration
Arxiv 2025 [Paper] -
Rethinking the Outlier Distribution in Large Language Models: An In-depth Study
Arxiv 2025 [Paper] -
Assigning Distinct Roles to Quantized and Low-Rank Matrices Toward Optimal Weight Decomposition
ACL Findings 2025 [Paper] -
Unifying Uniform and Binary-coding Quantization for Accurate Compression of Large Language Models
ACL 2025 [Paper] -
FPTQuant: Function-Preserving Transforms for LLM Quantization
ICML 2026 [Paper] -
BAQ: Efficient Bit Allocation Quantization for Large Language Models
Arxiv 2025 [Paper] [Code] -
MoQAE: Mixed-Precision Quantization for Long-Context LLM Inference via Mixture of Quantization-Aware Experts
ACL 2025 [Paper] -
Unifying Block-wise PTQ and Distillation-based QAT for Progressive Quantization toward 2-bit Instruction-Tuned LLMs
Arxiv 2025 [Paper] -
Boost Post-Training Quantization via Null Space Optimization for Large Language Models
Arxiv 2025 [Paper] [Code] -
FlexQuant: A Flexible and Efficient Dynamic Precision Switching Framework for LLM Quantization
Arxiv 2025 [Paper] -
BTC-LLM: Efficient Sub-1-Bit LLM Quantization via Learnable Transformation and Binary Codebook
Arxiv 2025 [Paper] -
ROSAQ: Rotation-based Saliency-Aware Weight Quantization for Efficiently Compressing Large Language Models
Arxiv 2025 [Paper] -
LittleBit: Ultra Low-Bit Quantization via Latent Factorization
NeurIPS 2025 [Paper] -
BASE-Q: Bias and Asymmetric Scaling Enhanced Rotational Quantization for Large Language Models
Arxiv 2025 [Paper] [Code] -
UltraSketchLLM: Sub-1-Bit LLM Compression via Sketch and Hardware-Friendly Operators
DAC 2026 [Paper] -
DBellQuant: Breaking the Bell with Double-Bell Transformation for LLMs Post Training Binarization
Arxiv 2025 [Paper] -
any4: Learned 4-bit Numeric Representation for LLMs
ICML 2025 [Paper] -
CCQ: Convolutional Code for Extreme Low-bit Quantization in LLMs
Arxiv 2025 [Paper] -
First-Order Error Matters: Accurate Compensation for Quantized Large Language Models
AAAI 2026 [Paper] -
PoTPTQ: A Two-step Power-of-Two Post-training for LLMs
ECAI 2025 [Paper] -
EAC-MoE: Expert-Selection Aware Compressor for Mixture-of-Experts Large Language Models
ACL 2025 [Paper] -
MicroMix: Efficient Mixed-Precision Quantization with Microscaling Formats for Large Language Models
Arxiv 2025 [Paper] [Code] -
VLMQ: Token Saliency-Driven Post-Training Quantization for Vision-language Models
Arxiv 2025 [Paper] -
FlashCommunication V2: Bit Splitting and Spike Reserving for Any Bit Communication
Arxiv 2025 [Paper] -
FlexQ: Efficient Post-training INT6 Quantization for LLM Serving via Algorithm-System Co-Design
Arxiv 2025 [Paper] [Code] -
Provable Post-Training Quantization: Theoretical Analysis of OPTQ and Qronos
Arxiv 2025 [Paper] -
iFairy: the First 2-bit Complex LLM with All Parameters in ${\pm1, \pm i}$
Arxiv 2025 [Paper] -
Pushing the Envelope of LLM Inference on AI-PC and Intel GPUs
Arxiv 2025 [Paper] -
Rethinking 1-bit Optimization Leveraging Pre-trained Large Language Models
Arxiv 2025 [Paper] -
Efficient Edge LLMs Deployment via HessianAware Quantization and CPU GPU Collaborative
Arxiv 2025 [Paper] -
Profiling Large Language Model Inference on Apple Silicon: A Quantization Perspective
Arxiv 2025 [Paper] -
LLM Compression: How Far Can We Go in Balancing Size and Performance?
RANLP 2025 [Paper] -
DLLMQuant: Quantizing Diffusion-based Large Language Models
Arxiv 2025 [Paper] -
Quantization Meets dLLMs: A Systematic Study of Post-training Quantization for Diffusion LLMs
Machine Intelligence Research 2025 [Paper] -
Systematic Characterization of LLM Quantization: A Performance, Energy, and Quality Perspective
Arxiv 2025 [Paper] -
Interpreting the Effects of Quantization on LLMs
AACL 2025 [Paper] -
Task-Stratified Knowledge Scaling Laws for Post-Training Quantized Large Language Models
ACL Findings 2026 [Paper] -
APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration
TCAD 2025 [Paper] -
Quantized but Deceptive? A Multi-Dimensional Truthfulness Evaluation of Quantized LLMs
EMNLP 2025 [Paper] -
The Uneven Impact of Post-Training Quantization in Machine Translation
Arxiv 2025 [Paper] -
BitROM: Weight Reload-Free CiROM Architecture Towards Billion-Parameter 1.58-bit LLM Inference
ASP-DAC 2026 [Paper] -
AMQ: Enabling AutoML for Mixed-precision Weight-Only Quantization of Large Language Models
EMNLP 2025 [Paper] -
Fair-GPTQ: Bias-Aware Quantization for Large Language Models
Arxiv 2025 [Paper] -
QWHA: Quantization-Aware Walsh-Hadamard Adaptation for Parameter-Efficient Fine-Tuning on Large Language Models
ICLR 2026 [Paper] -
Q-Palette: Fractional-Bit Quantizers Toward Optimal Bit Allocation for Efficient LLM Deployment
NeurIPS 2025 [Paper] -
AnyBCQ: Hardware Efficient Flexible Binary-Coded Quantization for Multi-Precision LLMs
ICLR 2026 [Paper] -
QeRL: Beyond Efficiency -- Quantization-enhanced Reinforcement Learning for LLMs
Arxiv 2025 [Paper] [Code] -
F-BFQ: Flexible Block Floating-Point Quantization Accelerator for LLMs
ISCA 2025 Workshop [Paper] -
Learning Grouped Lattice Vector Quantizers for Low-Bit LLM Compression
NeurIPS 2025 [Paper] -
A Convergence Analysis of Adaptive Optimizers under Floating-point Quantization
ICLR 2026 [Paper] -
FALQON: Accelerating LoRA Fine-tuning with Low-Bit Floating-Point Arithmetic
NeurIPS 2025 [Paper] -
TetraJet-v2: Accurate NVFP4 Training for Large Language Models with Oscillation Suppression and Outlier Control
ICML 2026 [Paper] -
DartQuant: Efficient Rotational Distribution Calibration for LLM Quantization
NeurIPS 2025 [Paper] -
You Had One Job: Per-Task Quantization Using LLMs' Hidden Representations
ICML 2026 Workshop [Paper] -
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats
ISCA 2026 [Paper] -
ParoQuant: Pairwise Rotation Quantization for Efficient Reasoning LLM Inference
ICLR 2026 [Paper] [Code] -
SpecQuant: Spectral Decomposition and Adaptive Truncation for Ultra-Low-Bit LLMs Quantization
AAAI 2026 [Paper] -
T-SAR: A Full-Stack Co-design for CPU-Only Ternary LLM Inference via In-Place SIMD ALU Reorganization
DATE 2026 [Paper] -
Enhancing Trustworthiness with Mixed Precision: Benchmarks, Opportunities, and Challenges
ASP-DAC 2026 [Paper] -
SignRoundV2: Toward Closing the Performance Gap in Extremely Low-Bit Post-Training Quantization for LLMs
Arxiv 2025 [Paper] [Code] -
CodeGEMM: A Codebook-Centric Approach to Efficient GEMM in Quantized LLMs
NeurIPS 2025 [Paper]
2024 · 162 papers
-
FlightLLM: Efficient Large Language Model Inference with a Complete Mapping Flow on FPGA
FPGA 2024 [Paper] -
Extreme Compression of Large Language Models via Additive Quantization
ICML 2024 [Paper] [Code] -
Quantized Side Tuning: Fast and Memory-Efficient Tuning of Quantized Large Language Models
Arxiv 2024 [Paper] -
Inferflow: an Efficient and Highly Configurable Inference Engine for Large Language Models
Arxiv 2024 [Paper] -
FP6-LLM: Efficiently Serving Large Language Models Through FP6-Centric Algorithm-System Co-Design
USENIX ATC 2024 [Paper] -
Can Large Language Models Understand Context?
EACL Findings 2024 [Paper] -
Squat: Quant Small Language Models on the Edge
ICCAD 2025 [Paper] [Code] -
LQER: Low-Rank Quantization Error Reconstruction for LLMs
ICML 2024 [Paper] -
BiLLM: Pushing the Limit of Post-Training Quantization for LLMs
Arxiv 2024 [Paper] [Code] -
QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks
ICML 2024 [Paper] [Code] -
L4Q: Parameter Efficient Quantization-Aware Training on Large Language Models via LoRA-wise LSQ
Arxiv 2024 [Paper] -
TP-Aware Dequantization
Arxiv 2024 [Paper] -
ApiQ: Finetuning of 2-Bit Quantized Large Language Model
EMNLP 2024 [Paper] -
Accurate LoRA-Finetuning Quantization of LLMs via Information Retention
Arxiv 2024 [Paper] [Code] -
BitDelta: Your Fine-Tune May Only Be Worth One Bit
NeurIPS 2024 [Paper] [Code] -
QDyLoRA: Quantized Dynamic Low-Rank Adaptation for Efficient Large Language Model Tuning
EMNLP 2024 Industry Track [Paper] -
Any-Precision LLM: Low-Cost Deployment of Multiple, Different-Sized LLMs
ICML 2024 [Paper] -
BitDistiller: Unleashing the Potential of Sub-4-Bit LLMs via Self-Distillation
ACL 2024 [Paper] [Code] -
OneBit: Towards Extremely Low-bit Large Language Models
NeurIPS 2024 [Paper] -
DB-LLM: Accurate Dual-Binarization for Efficient LLMs
ACL Findings 2024 [Paper] -
WKVQuant: Quantizing Weight and Key/Value Cache for Large Language Models Gains More
Arxiv 2024 [Paper] -
GPTVQ: The Blessing of Dimensionality for LLM Quantization
Arxiv 2024 [Paper] [Code] -
APTQ: Attention-aware Post-Training Mixed-Precision Quantization for Large Language Models
DAC 2024 [Paper] -
A Comprehensive Evaluation of Quantization Strategies for Large Language Models
ACL Findings 2024 [Paper] -
Evaluating Quantized Large Language Models
Arxiv 2024 [Paper] -
FlattenQuant: Breaking Through the Inference Compute-bound for Large Language Models with Per-tensor Quantization
Arxiv 2024 [Paper] -
LLM-PQ: Serving LLM on Heterogeneous Clusters with Phase-Aware Partition and Adaptive Quantization
Arxiv 2024 [Paper] -
IntactKV: Improving Large Languagze Model Quantization by Keeping Pivot Tokens Intact
ACL Findings 2024 [Paper] [Code] -
On the Compressibility of Quantized Large Language Models
Arxiv 2024 [Paper] -
EasyQuant: An Efficient Data-free Quantization Algorithm for LLMs
Arxiv 2024 [Paper] -
What Makes Quantization for Large Language Models Hard? An Empirical Study from the Lens of Perturbation
Arxiv 2024 [Paper] -
AffineQuant: Affine Transformation Quantization for Large Language Models
ICLR 2024 [Paper] [Code] -
Oh! We Freeze: Improving Quantized Knowledge Distillation via Signal Propagation Analysis for Large Language Models
ICLR Practical ML for Low Resource Settings Workshop 2024 [Paper] -
Accurate Block Quantization in LLMs with Outliers
Arxiv 2024 [Paper] -
QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs
Arxiv 2024 [Paper] [Code] -
Minimize Quantization Output Error with Bias Compensation
Arxiv 2024 [Paper] [Code] -
Cherry on Top: Parameter Heterogeneity and Quantization in Large Language Models
Arxiv 2024 [Paper] -
Fine-Tuning, Quantization, and LLMs: Navigating Unintended Outcomes
Arxiv 2024 [Paper] -
Quantization of Large Language Models with an Overdetermined Basis
Arxiv 2024 [Paper] -
An empirical study of LLaMA3 quantization: from LLMs to MLLMs
Arxiv 2024 [Paper] [Code] -
How to Parameterize Asymmetric Quantization Ranges for Quantization-Aware Training
Arxiv 2024 [Paper] -
Mitigating the Impact of Outlier Channels for Language Model Quantization with Activation Regularization
Arxiv 2024 [Paper] [Code] -
When Quantization Affects Confidence of Large Language Models?
NAACL 2024 [Paper] -
QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving
Arxiv 2024 [Paper] [Code] -
Learning from Students: Applying t-Distributions to Explore Accurate and Efficient Formats for LLMs
ICML 2024 [Paper] -
LLMC: Benchmarking Large Language Model Quantization with a Versatile Compression Toolkit
EMNLP 2024 [Paper] [Code] -
SKVQ: Sliding-window Key and Value Cache Quantization for Large Language Models
Arxiv 2024 [Paper] -
Post Training Quantization of Large Language Models with Microscaling Formats
Arxiv 2024 [Paper] -
Edge Intelligence Optimization for Large Language Model Inference with Batching and Quantization
Arxiv 2024 [Paper] -
SliM-LLM: Salience-Driven Mixed-Precision Quantization for Large Language Models
Arxiv 2024 [Paper] [Code] -
OAC: Output-adaptive Calibration for Accurate Post-training Quantization
Arxiv 2024 [Paper] -
PV-Tuning: Beyond Straight-Through Estimation for Extreme LLM Compression
Arxiv 2024 [Paper] -
SpinQuant -- LLM quantization with learned rotations
ICLR 2025 [Paper] -
Compressing Large Language Models using Low Rank and Low Precision Decomposition
NeurIPS 2024 [Paper] [Code] -
Athena: Efficient Block-Wise Post-Training Quantization for Large Language Models Using Second-Order Matrix Derivative Information
Arxiv 2024 [Paper] -
Exploiting LLM Quantization
Arxiv 2024 [Paper] -
One QuantLLM for ALL: Fine-tuning Quantized LLMs Once for Efficient Deployments
Arxiv 2024 [Paper] -
LCQ: Low-Rank Codebook based Quantization for Large Language Models
Arxiv 2024 [Paper] -
LoQT: Low Rank Adapters for Quantized Training
Arxiv 2024 [Paper] [Code] -
CLAQ: Pushing the Limits of Low-Bit Post-Training Quantization for LLMs
Arxiv 2024 [Paper] -
I-LLM: Efficient Integer-Only Inference for Fully-Quantized Low-Bit Large Language Models
Arxiv 2024 [Paper] -
Outliers and Calibration Sets have Diminishing Effect on Quantization of Modern LLMs
Arxiv 2024 [Paper] -
DuQuant: Distributing Outliers via Dual Transformation Makes Stronger Quantized LLMs
NeurIPS 2024 [Paper] [Code] -
ShiftAddLLM: Accelerating Pretrained LLMs via Post-Training Multiplication-Less Reparameterization
NeurIPS 2024 [Paper] [Code] -
Low-Rank Quantization-Aware Training for LLMs
Arxiv 2024 [Paper] -
TernaryLLM: Ternarized Large Language Model
Arxiv 2024 [Paper] -
Examining Post-Training Quantization for Mixture-of-Experts: A Benchmark
Arxiv 2024 [Paper] [Code] -
Delta-CoMe: Training-Free Delta-Compression with Mixed-Precision for Large Language Models
NeurIPS 2024 [Paper] -
QQQ: Quality Quattuor-Bit Quantization for Large Language Models
Arxiv 2024 [Paper] [Code] -
QTIP: Quantization with Trellises and Incoherence Processing
NeurIPS 2024 [Paper] [Code] -
Prefixing Attention Sinks can Mitigate Activation Outliers for Large Language Model Quantization
EMNLP 2024 [Paper] -
Mixture of Scales: Memory-Efficient Token-Adaptive Binarization for Large Language Models
NeurIPS 2024 [Paper] -
Tender: Accelerating Large Language Models via Tensor Decomposition and Runtime Requantization
ISCA 2024 [Paper] -
SDQ: Sparse Decomposed Quantization for LLM Inference
Arxiv 2024 [Paper] -
Attention-aware Post-training Quantization without Backpropagation
ICML 2025 [Paper] -
EDGE-LLM: Enabling Efficient Large Language Model Adaptation on Edge Devices via Layerwise Unified Compression and Adaptive Layer Tuning and Voting
Arxiv 2024 [Paper] [Code] -
Compensate Quantization Errors: Make Weights Hierarchical to Compensate Each Other
Arxiv 2024 [Paper] -
Layer-Wise Quantization: A Pragmatic and Effective Method for Quantizing LLMs Beyond Integer Bit-Levels
Arxiv 2024 [Paper] [Code] -
CDQuant: Greedy Coordinate Descent for Accurate LLM Quantization
Arxiv 2024 [Paper] -
OutlierTune: Efficient Channel-Wise Quantization for Large Language Models
Arxiv 2024 [Paper] -
T-MAC: CPU Renaissance via Table Lookup for Low-Bit LLM Deployment on Edge
EuroSys 2025 [Paper] [Code] -
GPTQT: Quantize Large Language Models Twice to Push the Efficiency
ICORIS 2024 [Paper] -
Improving Conversational Abilities of Quantized Large Language Models via Direct Preference Alignment
ACL 2024 [Paper] -
How Does Quantization Affect Multilingual LLMs?
EMNLP Findings 2024 [Paper] -
RoLoRA: Fine-tuning Rotated Outlier-free LLMs for Effective Weight-Activation Quantization
EMNLP Findings 2024 [Paper] [Code] -
Q-GaLore: Quantized GaLore with INT4 Projection and Layer-Adaptive Low-Rank Gradients
Arxiv 2024 [Paper] [Code] -
FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision
Arxiv 2024 [Paper] [Code] -
Accuracy is Not All You Need
Arxiv 2024 [Paper] -
BitNet b1.58 Reloaded: State-of-the-art Performance Also on Smaller Networks
Arxiv 2024 [Paper] -
LeanQuant: Accurate Large Language Model Quantization with Loss-Error-Aware Grid
ICLR 2025 [Paper] -
Fast Matrix Multiplications for Lookup Table-Quantized LLMs
EMNLP Findings 2024 [Paper] [Code] -
EfficientQAT: Efficient Quantization-Aware Training for Large Language Models
ACL 2025 [Paper] [Code] -
LRQ: Optimizing Post-Training Quantization for Large Language Models by Learning Low-Rank Weight-Scaling Matrices
NAACL 2025 [Paper] -
Exploring Quantization for Efficient Pre-Training of Transformer Language Models
EMNLP Findings 2024 [Paper] [Code] -
Spectra: Surprising Effectiveness of Pretraining Ternary Language Models at Scale
Arxiv 2024 [Paper] [Code] -
Mamba-PTQ: Outlier Channels in Recurrent Large Language Models
Efficient Systems for Foundation Models Workshop @ ICML 2024 [Paper] -
Compensate Quantization Errors+: Quantized Models Are Inquisitive Learners
Arxiv 2024 [Paper] -
Accurate and Efficient Fine-Tuning of Quantized Large Language Models Through Optimal Balance
Arxiv 2024 [Paper] [Code] -
STBLLM: Breaking the 1-Bit Barrier with Structured Binary LLMs
Arxiv 2024 [Paper] -
Advancing Multimodal Large Language Models with Quantization-Aware Scale Learning for Efficient Adaptation
ACM MM 2024 [Paper] -
ABQ-LLM: Arbitrary-Bit Quantized Inference Acceleration for Large Language Models
AAAI 2025 [Paper] -
MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models
Arxiv 2024 [Paper] [Code (Marlin)] [Code (Sparse Marlin)] -
Matmul or No Matmal in the Era of 1-bit LLMs
Arxiv 2024 [Paper] -
MobileQuant: Mobile-friendly Quantization for On-device Language Models
EMNLP Findings 2024 [Paper] [Code] -
GIFT-SW: Gaussian noise Injected Fine-Tuning of Salient Weights for LLMs
Arxiv 2024 [Paper] [Code] -
Foundations of Large Language Model Compression -- Part 1: Weight Quantization
Arxiv 2024 [Paper] -
OPAL: Outlier-Preserved Microscaling Quantization A ccelerator for Generative Large Language Models
DAC 2024 [Paper] -
VPTQ: Extreme Low-bit Vector Post-Training Quantization for Large Language Models
EMNLP 2024 [Paper] [Code] -
Scaling FP8 training to trillion-token LLMs
Arxiv 2024 [Paper] -
Accumulator-Aware Post-Training Quantization
Arxiv 2024 [Paper] -
Efficient Arbitrary Precision Acceleration for Large Language Models on GPU Tensor Cores
ASP-DAC 2025 [Paper] -
Rotated Runtime Smooth: Training-Free Activation Smoother for accurate INT4 inference
Arxiv 2024 [Paper] [Code] -
EXAQ: Exponent Aware Quantization For LLMs Acceleration
Arxiv 2024 [Paper] -
ARB-LLM: Alternating Refined Binarizations for Large Language Models
Arxiv 2024 [Paper] [Code] -
PrefixQuant: Eliminating Outliers by Prefixed Tokens for Large Language Models Quantization
Arxiv 2024 [Paper] [Code] -
Sketch to Adapt: Fine-Tunable Sketches for Efficient LLM Adaptation
ICML 2025 [Paper] -
Scaling Laws For Mixed Quantization
Arxiv 2024 [Paper] -
Q-VLM: Post-training Quantization for Large Vision-Language Models
NeurIPS 2024 [Paper] [Code] -
CrossQuant: A Post-Training Quantization Method with Smaller Quantization Kernel for Precise Large Language Model Compression
Arxiv 2024 [Paper] -
FlatQuant: Flatness Matters for LLM Quantization
ICML 2025 [Paper] [Code] -
DeltaDQ: Ultra-High Delta Compression for Fine-Tuned LLMs via Group-wise Dropout and Separate Quantization
Arxiv 2024 [Paper] -
QEFT: Quantization for Efficient Fine-Tuning of LLMs
EMNLP Findings 2024 [Paper] [Code] -
Continuous Approximations for Improving Quantization Aware Training of LLMs
Arxiv 2024 [Paper] -
DAQ: Density-Aware Post-Training Weight-Only Quantization For LLMs
Arxiv 2024 [Paper] -
COMET: Towards Partical W4A4KV4 LLMs Serving
Arxiv 2024 [Paper] -
Scaling laws for post-training quantized large language models
Arxiv 2024 [Paper] -
Channel-Wise Mixed-Precision Quantization for Large Language Models
Arxiv 2024 [Paper] -
Understanding the difficulty of low-precision post-training quantization of large language models
Arxiv 2024 [Paper] -
QuAILoRA: Quantization-Aware Initialization for LoRA
NeurIPS Workshop on Efficient Natural Language and Speech Processing (ENLSP-IV) 2024 [Paper] -
SDP4Bit: Toward 4-bit Communication Quantization in Sharded Data Parallelism for LLM Training
NeurIPS 2024 [Paper] -
Pyramid Vector Quantization for LLMs
Arxiv 2024 [Paper] -
TesseraQ: Ultra Low-Bit LLM Post-Training Quantization with Block Reconstruction
Arxiv 2024 [Paper] [Code] -
COAT: Compressing Optimizer states and Activation for Memory-Efficient FP8 Training
ICLR 2025 [Paper] [Code] -
GWQ: Gradient-Aware Weight Quantization for Large Language Models
Arxiv 2024 [Paper] -
"Give Me BF16 or Give Me Death"? Accuracy-Performance Trade-Offs in LLM Quantization
ACL 2025 [Paper] -
Interactions Across Blocks in Post-Training Quantization of Large Language Models
Arxiv 2024 [Paper] -
BitNet a4.8: 4-bit Activations for 1-bit LLMs
Arxiv 2024 [Paper] -
The Super Weight in Large Language Models
Arxiv 2024 [Paper] [Code] -
ASER: Activation Smoothing and Error Reconstruction for Large Language Model Quantization
AAAI 2025 [Paper] -
Towards Low-bit Communication for Tensor Parallel LLM Inference
Arxiv 2024 [Paper] -
AMXFP4: Taming Activation Outliers with Asymmetric Microscaling Floating-Point for 4-bit LLM Inference
Arxiv 2024 [Paper] [Code] -
Scaling Laws for Precision
Arxiv 2024 [Paper] -
BitMoD: Bit-serial Mixture-of-Datatype LLM Acceleration
HPCA 2025 [Paper] [Code] -
SageAttention2: Efficient Attention with Thorough Outlier Smoothing and Per-thread INT4 Quantization
ICML 2025 [Paper] [Code] -
AutoMixQ: Self-Adjusting Quantization for High Performance Memory-Efficient Fine-Tuning
Arxiv 2024 [Paper] -
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format
HPCA 2025 [Paper] -
MixPE: Quantization and Hardware Co-design for Efficient LLM Inference
Arxiv 2024 [Paper] -
Pushing the Limits of Large Language Model Quantization via the Linearity Theorem
Arxiv 2024 [Paper] -
Low-Bit Quantization Favors Undertrained LLMs: Scaling Laws for Quantized LLMs with 100T Training Tokens
Arxiv 2024 [Paper] [Models] -
DFRot: Achieving Outlier-Free and Massive Activation-Free for Rotated LLMs with Refined Rotation
COLM 2025 [Paper] [Code] -
RILQ: Rank-Insensitive LoRA-based Quantization Error Compensation for Boosting 2-bit Large Language Model Accuracy
AAAI 2025 [Paper] -
CPTQuant -- A Novel Mixed Precision Post-Training Quantization Techniques for Large Language Models
Arxiv 2024 [Paper] -
SKIM: Any-bit Quantization Pushing The Limits of Post-Training Quantization
Arxiv 2024 [Paper] -
Direct Quantized Training of Language Models with Stochastic Rounding
Arxiv 2024 [Paper] [Code] -
Taming Sensitive Weights : Noise Perturbation Fine-tuning for Robust LLM Quantization
Arxiv 2024 [Paper] -
Low-Rank Correction for Quantized LLMs
Arxiv 2024 [Paper] -
CRVQ: Channel-relaxed Vector Quantization for Extreme Compression of LLMs
Arxiv 2024 [Paper] -
ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals
Arxiv 2024 [Paper] [Code] -
MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design
MLSys 2026 [Paper] -
GQSA: Group Quantization and Sparsity for Accelerating Large Language Model Inference
Arxiv 2024 [Paper] -
LSAQ: Layer-Specific Adaptive Quantization for Large Language Model Deployment
Arxiv 2024 [Paper] -
DecDEC: A Systems Approach to Advancing Low-Bit LLM Quantization
OSDI 2025 [Paper]
2023 · 75 papers
-
FlexRound: Learnable Rounding based on Element-wise Division for Post-Training Quantization
ICML 2023 [Paper] [Code (DeepSpeed)] -
Understanding INT4 Quantization for Transformer Models: Latency Speedup, Composability, and Failure Cases
ICML 2023 [Paper] [Code] -
The case for 4-bit precision: k-bit Inference Scaling Laws
ICML 2023 [Paper] -
PreQuant: A Task-agnostic Quantization Approach for Pre-trained Language Models
ACL 2023 [Paper] -
Boost Transformer-based Language Models with GPU-Friendly Sparsity and Quantization
ACL 2023 [Paper] -
QLoRA: Efficient Finetuning of Quantized LLMs
NeurIPS 2023 [Paper] [Code] -
The Quantization Model of Neural Scaling
NeurIPS 2023 [Paper] -
Quantized Distributed Training of Large Models with Convergence Guarantees
ICML 2023 [Paper] -
RPTQ: Reorder-based Post-training Quantization for Large Language Models
Arxiv 2023 [Paper] [Code] -
ZeroQuant-V2: Exploring Post-training Quantization in LLMs from Comprehensive Study to Low Rank Compensation
AAAI 2024 [Paper] [Code] -
Integer or Floating Point? New Outlooks for Low-Bit Quantization on Large Language Models
ICML 2024 [Paper] -
Memory-Efficient Fine-Tuning of Compressed Large Language Models via sub-4-bit Integer Quantization
NeurIPS 2023 [Paper] -
Compress, Then Prompt: Improving Accuracy-Efficiency Trade-off of LLM Inference with Transferable Prompt
Arxiv 2023 [Paper] -
AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration
MLSys 2024 (Best Paper 🏆) [Paper] [Code] -
LLM-QAT: Data-Free Quantization Aware Training for Large Language Models
ACL Findings 2024 [Paper] [Code] -
SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression
ICLR 2024 [Paper] [Code] -
OWQ: Outlier-Aware Weight Quantization for Efficient Fine-Tuning and Inference of Large Language Models
AAAI 2024 [Paper] -
SqueezeLLM: Dense-and-Sparse Quantization
ICML 2024 [Paper] [Code] -
INT2.1: Towards Fine-Tunable Quantized Large Language Models with Error Correction through Low-Rank Adaptation
Arxiv 2023 [Paper] -
LQ-LoRA: Low-rank Plus Quantized Matrix Decomposition for Efficient Language Model Finetuning
ICLR 2024 [Paper] -
INT-FP-QSim: Mixed Precision and Formats For Large Language Models and Vision Transformers
Arxiv 2023 [Paper] [Code] -
QIGen: Generating Efficient Kernels for Quantized Inference on Large Language Models
Arxiv 2023 [Paper] [Code] -
Do Emergent Abilities Exist in Quantized Large Language Models: An Empirical Study
COLING 2024 [Paper] -
ZeroQuant-FP: A Leap Forward in LLMs Post-Training W4A8 Quantization Using Floating-Point Formats
Arxiv 2023 [Paper] [Code (DeepSpeed)] -
OliVe: Accelerating Large Language Models via Hardware-friendly Outlier-Victim Pair Quantization
ISCA 2023 [Paper] -
NUPES : Non-Uniform Post-Training Quantization via Power Exponent Search
Arxiv 2023 [Paper] -
GPT-Zip: Deep Compression of Finetuned Large Language Models
ICML 2023 Workshop ES-FoMO [Paper] -
Generating Efficient Kernels for Quantized Inference on Large Language Models
ICML 2023 Workshop ES-FoMO [Paper] -
Gradient-Based Post-Training Quantization: Challenging the Status Quo
Arxiv 2023 [Paper] -
FineQuant: Unlocking Efficiency with Fine-Grained Weight-Only Quantization for LLMs
Arxiv 2023 [Paper] -
OmniQuant: Omnidirectionally Calibrated Quantization for Large Language Models
ICLR 2024 [Paper] [Code] -
FPTQ: Fine-grained Post-Training Quantization for Large Language Models
Arxiv 2023 [Paper] -
eDKM: An Efficient and Accurate Train-time Weight Clustering for Large Language Models
IEEE Computer Architecture Letters 2023 [Paper] -
QuantEase: Optimization-based Quantization for Language Models -- An Efficient and Intuitive Algorithm
Arxiv 2023 [Paper] -
Norm Tweaking: High-performance Low-bit Quantization of Large Language Models
AAAI 2024 [Paper] -
Understanding the Impact of Post-Training Quantization on Large-scale Language Models
Arxiv 2023 [Paper] -
MEMORY-VQ: Compression for Tractable Internet-Scale Memory
NAACL 2024 [Paper] -
Optimize Weight Rounding via Signed Gradient Descent for the Quantization of LLMs
EMNLP Findings 2024 [Paper] [Code] -
Efficient Post-training Quantization with FP8 Formats
MLSys 2024 [Paper] [Code (Intel® Neural Compressor)] -
QA-LoRA: Quantization-Aware Low-Rank Adaptation of Large Language Models
ICLR 2024 [Paper] [Code] -
Rethinking Channel Dimensions to Isolate Outliers for Low-bit Weight Quantization of Large Language Models
ICLR 2024 [Paper] [Code] -
ModuLoRA: Finetuning 3-Bit LLMs on Consumer GPUs by Integrating with Modular Quantizers
TMLR (Featured Certification 🌟) [Paper] -
PB-LLM: Partially Binarized Large Language Models
ICLR 2024 [Paper] [Code] -
Dual Grained Quantization: Efficient Fine-Grained Quantization for LLM
Arxiv 2023 [Paper] -
QLLM: Accurate and Efficient Low-Bitwidth Quantization for Large Language Models
ICLR 2024 [Paper] [Code] -
LoftQ: LoRA-Fine-Tuning-Aware Quantization for Large Language Models
ICLR 2024 [Paper] [Code] -
QFT: Quantized Full-parameter Tuning of LLMs with Affordable Resources
ICLR 2026 Workshop [Paper] -
TEQ: Trainable Equivalent Transformation for Quantization of LLMs
Arxiv 2023 [Paper] [Code (Intel® Neural Compressor)] -
BitNet: Scaling 1-bit Transformers for Large Language Models
Arxiv 2023 [Paper] [Code] -
FP8-LM: Training FP8 Large Language Models
Arxiv 2023 [Paper] [Code] -
QUIK: Towards End-to-End 4-Bit Inference on Generative Large Language Models
EMNLP 2024 [Paper] [Code] -
AFPQ: Asymmetric Floating Point Quantization for LLMs
ACL Findings 2024 [Paper] [Code] -
AWEQ: Post-Training Quantization with Activation-Weight Equalization for Large Language Models
Arxiv 2023 [Paper] -
Atom: Low-bit Quantization for Efficient and Accurate LLM Serving
MLSys 2024 [Paper] [Code] -
QMoE: Practical Sub-1-Bit Compression of Trillion-Parameter Models
Arxiv 2023 [Paper] -
Dissecting the Runtime Performance of the Training, Fine-tuning, and Inference of Large Language Models
Arxiv 2023 [Paper] -
On the Impact of Calibration Data in Post-training Quantization and Pruning
ACL 2024 [Paper] -
A Speed Odyssey for Deployable Quantization of LLMs
Arxiv 2023 [Paper] -
Fast and Efficient 2-bit LLM Inference on GPU: 2/4/16-bit in a Weight Matrix with Asynchronous Dequantization
Arxiv 2023 [Paper] -
Quantizable Transformers: Removing Outliers by Helping Attention Heads Do Nothing
NeurIPS 2023 [Paper] [Code] -
Efficient LLM Inference on CPUs
NeurIPS 2023 on Efficient Natural Language and Speech Processing [Paper] [Code] -
The Cost of Compression: Investigating the Impact of Compression on Parametric Knowledge in Language Models
EMNLP Findings 2023 [Paper] -
Zero-Shot Sharpness-Aware Quantization for Pre-trained Language Models
EMNLP 2023 [Paper] -
Revisiting Block-based Quantisation: What is Important for Sub-8-bit LLM Inference?
EMNLP 2023 [Paper] [Code] -
Outlier Suppression+: Accurate quantization of large language models by equivalent and optimal shifting and scaling
EMNLP 2023 [Paper] -
Watermarking LLMs with Weight Quantization
EMNLP 2023 [Paper] [Code] -
Enhancing Computation Efficiency in Large Language Models through Weight and Activation Quantization
EMNLP 2023 [Paper] -
LLM-FP4: 4-Bit Floating-Point Quantized Transformers
EMNLP 2023 [Paper] [Code] -
Agile-Quant: Activation-Guided Quantization for Faster Inference of LLMs on the Edge
AAAI 2024 [Paper] -
SmoothQuant+: Accurate and Efficient 4-bit Post-Training WeightQuantization for LLM
Arxiv 2023 [Paper] -
CBQ: Cross-Block Quantization for Large Language Models
Arxiv 2023 [Paper] -
ZeroQuant(4+2): Redefining LLMs Quantization with a New FP6-Centric Strategy for Diverse Generative Tasks
Arxiv 2023 [Paper] -
QuIP: 2-Bit Quantization of Large Language Models With Guarantees
NeurIPS 2023 [Paper] [Code] -
A Performance Evaluation of a Quantized Large Language Model on Various Smartphones
Arxiv 2023 [Paper] -
DeltaZip: Efficient Serving of Multiple Full-Model-Tuned LLMs
EuroSys 2025 [Paper] [Code]
2022 · 6 papers
-
ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale Transformers
NeurIPS 2022 [Paper] [Code (DeepSpeed)] -
LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale
NeurIPS 2022 [Paper] [Code] -
Outlier Suppression: Pushing the Limit of Low-bit Transformer Language Models
NeurIPS 2022 [Paper] [Code] -
LUT-GEMM: Quantized Matrix Multiplication based on LUTs for Efficient Inference in Large-Scale Generative Language Models
ICLR 2024 [Paper] -
SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models
ICML 2023 [Paper] [Code] -
GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers
ICLR 2023 [Paper] [Code]
Pruning and Sparsity
2026 · 33 papers
-
Adaptive Layer Selection for Layer-Wise Token Pruning in LLM Inference
ACL Findings 2026 [Paper] [Code] -
LLMs can Compress LLMs: Adaptive Pruning by Agents
Arxiv 2026 [Paper] -
Streaming-dLLM: Accelerating Diffusion LLMs via Suffix Pruning and Dynamic Decoding
Arxiv 2026 [Paper] [Code] -
GradPruner: Gradient-Guided Layer Pruning Enabling Efficient Fine-Tuning and Inference for LLMs
ICLR 2026 [Paper] -
FASA: Frequency-aware Sparse Attention
ICLR 2026 [Paper] -
Compressing LLMs with MoP: Mixture of Pruners
Arxiv 2026 [Paper] [Code] -
Pruning as a Cooperative Game: Surrogate-Assisted Layer Contribution Estimation for Large Language Models
ICLR 2026 [Paper] -
Sink-Aware Pruning for Diffusion Language Models
Arxiv 2026 [Paper] [Code] -
Curvature-Weighted Capacity Allocation: A Minimum Description Length Framework for Layer-Adaptive Large Language Model Optimization
UAI 2026 [Paper] [Code] -
Sparse-BitNet: 1.58-bit LLMs are Naturally Friendly to Semi-Structured Sparsity
Arxiv 2026 [Paper] [Code] -
Stem: Rethinking Causal Information Flow in Sparse Attention
ICML 2026 [Paper] -
High-Fidelity Pruning for Large Language Models
Arxiv 2026 [Paper] [Code] -
Sparser, Faster, Lighter Transformer Language Models
Arxiv 2026 [Paper] [Code] -
REAM: Merging Improves Pruning of Experts in LLMs
Arxiv 2026 [Paper] [Code] -
GRASPrune: Global Gating for Budgeted Structured Pruning of Large Language Models
ACL 2026 [Paper] -
Revisiting the Effectiveness of LLM Pruning for Test-Time Scaling
EMNLP Findings 2026 [Paper] -
Compute Where it Counts: Self Optimizing Language Models
Arxiv 2026 [Paper] [Code] -
Ghosted Layers: Unconstrained Activation Alignment for Recovering Layer-Pruned LLMs
Arxiv 2026 [Paper] [Code] -
LEAP: Learnable End-to-End Adaptive Pruning of Large Language Models
ICML 2026 Workshop [Paper] -
Efficient Long-Context Modeling in Diffusion Language Models via Block Approximate Sparse Attention
CVPR 2026 [Paper] -
Locality-Aware Redundancy Pruning for LLM Depth Compression
Arxiv 2026 [Paper] [Code] -
FlashMemory-DeepSeek-V4: Lightning Index Ultra-Long Context via Lookahead Sparse Attention
Arxiv 2026 [Paper] [Code] -
Beyond FLOPs: Benchmarking Real Inference Acceleration of LLM Pruning under a GEMM-Centric Taxonomy
Arxiv 2026 [Paper] [Code] -
Persona-Pruner: Sculpting Lightweight Models for Role-Playing
ICML 2026 [Paper] [Code] -
Dustin: Draft-Augmented Sparse Verification for Efficient Long-Context Generation with Speculative Decoding
ICML 2026 [Paper] -
EPTS: Elastic Post-Training Sparsity for Efficient Large Language Model Compression
KDD 2026 [Paper] -
Super-Tuning: From Activation-Aware Pruning to Sparse Fine-Tuning
Arxiv 2026 [Paper] [Code] -
WIDE: Boosting Adaptive LLM Inference via Token-level Dynamic Width Pruning
Arxiv 2026 [Paper] [Code] -
Celty: SpMspV GPU Kernel and SIMT Co-Design for Efficient Dual-Sparse LLM Inference
ICCAD 2026 [Paper] [Code] -
Prune Once: Retraining-Free Task-Agnostic Pruning for Vision-Language Models
ECCV 2026 [Paper] -
Router Sensitivity Under Lightweight Fine-Tuning Identifies Prunable Experts in Mixture-of-Experts Models
Arxiv 2026 [Paper] [Code] -
Reservoir of Importance: Learning Semi-Structured Sparsity with Differentiable Subset Sampling
EMNLP 2026 [Paper]
2025 · 60 papers
-
FASP: Fast and Accurate Structured Pruning of Large Language Models
Arxiv 2025 [Paper] -
MultiPruner: Balanced Structure Removal in Foundation Models
Arxiv 2025 [Paper] [Code] -
Mamba-Shedder: Post-Transformer Compression for Efficient Selective Structured State Space Models
NAACL 2025 [Paper] [Code] -
2SSP: A Two-Stage Framework for Structured Pruning of LLMs
Arxiv 2025 [Paper] [Code] -
You Only Prune Once: Designing Calibration-Free Model Compression With Policy Learning
ICLR 2025 [Paper] -
SwiftPrune: Hessian-Free Weight Pruning for Large Language Models
Arxiv 2025 [Paper] -
Pivoting Factorization: A Compact Meta Low-Rank Representation of Sparsity for Efficient Inference in Large Language Models
ICML 2025 [Paper] -
Twilight: Adaptive Attention Sparsity with Hierarchical Top-p Pruning
NeurIPS 2025 [Paper] -
Adapt-Pruner: Adaptive Structural Pruning for Efficient Small Language Model Training
Arxiv 2025 [Paper] -
Dobi-SVD: Differentiable SVD for LLM Compression and Some New Perspectives
ICLR 2025 [Paper] [Homepage] -
EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models
Arxiv 2025 [Paper] [Code] -
DarwinLM: Evolutionary Structured Pruning of Large Language Models
COLM 2026 [Paper] -
MaskPrune: Mask-based LLM Pruning for Layer-wise Uniform Structures
Arxiv 2025 [Paper] -
Determining Layer-wise Sparsity for Large Language Models Through a Theoretical Perspective
Arxiv 2025 [Paper] -
PPC-GPT: Federated Task-Specific Compression of Large Language Models via Pruning and Chain-of-Thought Distillation
Arxiv 2025 [Paper] -
Compression Scaling Laws: Unifying Sparsity and Quantization
Arxiv 2025 [Paper] -
PASER: Post-Training Data Selection for Efficient Pruned Large Language Model Recovery
ICLR 2026 [Paper] -
Týr-the-Pruner: Unlocking Accurate 50% Structural Pruning for LLMs via Global Sparsity Distribution Optimization
Arxiv 2025 [Paper] -
Thanos: A Block-wise Pruning Algorithm for Efficient Large Language Model Compression
Arxiv 2025 [Paper] -
Efficient LLMs with AMP: Attention Heads and MLP Pruning
IJCNN 2025 [Paper] -
ReplaceMe: Network Simplification via Layer Pruning and Linear Transformations
NeurIPS 2025 [Paper] [Code] -
Large Language Model Compression with Global Rank and Sparsity Optimization
Arxiv 2025 [Paper] -
TRIM: Achieving Extreme Sparsity with Targeted Row-wise Iterative Metric-driven Pruning
Arxiv 2025 [Paper] [Code] -
RAP: Runtime-Adaptive Pruning for LLM Inference
Arxiv 2025 [Paper] -
Two-Stage Regularization-Based Structured Pruning for LLMs
ACL 2026 [Paper] -
Pangu Light: Weight Re-Initialization for Pruning and Accelerating LLMs
Arxiv 2025 [Paper] -
Sparsified State-Space Models are Efficient Highway Networks
TMLR 2025 [Paper] [Code] -
SDMPrune: Self-Distillation MLP Pruning for Efficient Large Language Models
Arxiv 2025 [Paper] [Code] -
Amber Pruner: Leveraging N:M Activation Sparsity for Efficient Prefill in Large Language Models
Arxiv 2025 [Paper] -
Beyond Manually Designed Pruning Policies with Second-Level Performance Prediction: A Pruning Framework for LLMs
Arxiv 2025 [Paper] [Code] -
Pruning Large Language Models by Identifying and Preserving Functional Networks
Arxiv 2025 [Paper] [Code] -
SlimInfer: Accelerating Long-Context LLM Inference via Dynamic Token Pruning
Arxiv 2025 [Paper] [Code] -
EGGS-PTP: An Expander-Graph Guided Structured Post-training Pruning Method for Large Language Models
Arxiv 2025 [Paper] -
Z-Pruner: Post-Training Pruning of Large Language Models for Efficiency without Retraining
AICCSA 2025 [Paper] [Code] -
H2EAL: Hybrid-Bonding Architecture with Hybrid Sparse Attention for Efficient Long-Context LLM Inference
ICCAD 2025 [Paper] -
Less Is More? Examining Fairness in Pruned Large Language Models for Summarising Opinions
EMNLP 2025 [Paper] [Code] -
DualSparse-MoE: Coordinating Tensor/Neuron-Level Sparsity with Expert Partition and Reconstruction
Arxiv 2025 [Paper] -
Optimal Sparsity of Mixture-of-Experts Language Models for Reasoning Tasks
ICLR 2026 [Paper] [Code] -
Spatio-Temporal Pruning for Compressed Spiking Large Language Models
Arxiv 2025 [Paper] -
Dropping Experts, Recombining Neurons: Retraining-Free Pruning for Sparse Mixture-of-Experts LLMs
EMNLP 2025 [Paper] -
Reasoning Models Can be Accurately Pruned Via Chain-of-Thought Reconstruction
Arxiv 2025 [Paper] [Code] -
NIRVANA: Structured Pruning Reimagined for Large Language Model Compression
COLM 2026 [Paper] -
HEAPr: Hessian-based Efficient Atomic Expert Pruning in Output Space
ICLR 2026 [Paper] [Code] -
ProxyAttn: Guided Sparse Attention via Representative Heads
ICLR 2026 [Paper] -
Effective Model Pruning: Measure The Redundancy of Model Components
ICML 2026 [Paper] -
The Unseen Frontier: Pushing the Limits of LLM Sparsity with Surrogate-Free ADMM
ICLR 2026 [Paper] -
ARMOR: High-Performance Semi-Structured Pruning via Adaptive Matrix Factorization
ICLR 2026 [Paper] [Code] -
RCPU: Rotation-Constrained Error Compensation for Structured Pruning of Large Language Models
ICLR 2026 [Paper] -
Fewer Weights, More Problems: A Practical Attack on LLM Pruning
ICLR 2026 [Paper] [Code] -
From Local to Global: Revisiting Structured Pruning Paradigms for Large Language Models
ACL 2026 [Paper] -
Sparser Block-Sparse Attention via Token Permutation
ICML 2026 [Paper] [Code] -
Restoring Pruned Large Language Models via Lost Component Compensation
NeurIPS 2025 [Paper] -
When Fewer Layers Break More Chains: Layer Pruning Harms Test-Time Scaling in LLMs
Arxiv 2025 [Paper] [Code] -
1+1>2: A Synergistic Sparse and Low-Rank Compression Method for Large Language Models
EMNLP Findings 2025 [Paper] -
SpecAttn: Speculating Sparse Attention
NeurIPS 2025 Workshop [Paper] -
IG-Pruning: Input-Guided Block Pruning for Large Language Models
EMNLP 2025 [Paper] [Code] -
MACKO: Sparse Matrix-Vector Multiplication for Low Sparsity
Arxiv 2025 [Paper] [Code] -
Understanding and Harnessing Sparsity in Unified Multimodal Models
Arxiv 2025 [Paper] [Code] -
Resting Neurons, Active Insights: Robustifying Activation Sparsity in LLMs via Spontaneity
ICML 2026 [Paper] -
Fragile Knowledge, Robust Instruction-Following: The Width Pruning Dichotomy in Llama-3.2
Arxiv 2025 [Paper] [Code]
2024 · 77 papers
-
Plug-and-Play: An Efficient Post-training Pruning Method for Large Language Models
ICLR 2024 [Paper] [Code] -
Fast and Optimal Weight Update for Pruned Large Language Models
Arxiv 2024 [Paper] -
APT: Adaptive Pruning and Tuning Pretrained Language Models for Efficient Training and Inference
ICML 2024 [Paper] -
Scaling Sparse Fine-Tuning to Large Language Models
Arxiv 2024 [Paper] -
SliceGPT: Compress Large Language Models by Deleting Rows and Columns
ICLR 2024 [Paper] [Code] -
Shortened LLaMA: Depth Pruning for Large Language Models with Comparison of Retraining Methods
ICLR 2024 Workshop [Paper] -
Everybody Prune Now: Structured Pruning of LLMs with only Forward Passes
Arxiv 2024 [Paper] [Code] -
NutePrune: Efficient Progressive Pruning with Numerous Teachers for Large Language Models
Arxiv 2024 [Paper] -
LaCo: Large Language Model Pruning via Layer Collapse
EMNLP Findings 2024 [Paper] -
Why Lift so Heavy? Slimming Large Language Models by Cutting Off the Layers
Arxiv 2024 [Paper] -
EBFT: Effective and Block-Wise Fine-Tuning for Sparse LLMs
Arxiv 2024 [Paper] [Code] -
Data-free Weight Compress and Denoise for Large Language Models
Arxiv 2024 [Paper] -
Gradient-Free Adaptive Global Pruning for Pre-trained Language Models
NeurIPS 2024 [Paper] -
ShortGPT: Layers in Large Language Models are More Redundant Than You Expect
Arxiv 2024 [Paper] -
LLaVA-PruMerge: Adaptive Token Reduction for Efficient Large Multimodal Models
ICCV 2025 [Paper] [Code] -
Streamlining Redundant Layers to Compress Large Language Models
Arxiv 2024 [Paper] -
LoRAP: Transformer Sub-Layers Deserve Differentiated Structured Compression for Large Language Models
Arxiv 2024 [Paper] -
LoNAS: Elastic Low-Rank Adapters for Efficient Large Language Models
COLING 2024 [Paper] [Code] -
Shears: Unstructured Sparsity with Neural Low-rank Adapter Search
NAACL 2024 [Paper] [Code] -
Eigenpruning: an Interpretability-Inspired PEFT Method
NAACL 2024 Abstract [Paper] -
OpenBA-V2: Reaching 77.3% High Compression Ratio with Fast Multi-Stage Pruning
Arxiv 2024 [Paper] -
Pruning as a Domain-specific LLM Extractor
NAACL 2024 Findings [Paper] [Code] -
Differentiable Model Scaling using Differentiable Topk
ICML 2024 [Paper] -
COPAL: Continual Pruning in Large Language Generative Models
ICML 2024 [Paper] -
Pruner-Zero: Evolving Symbolic Pruning Metric from scratch for Large Language Models
ICML 2024 [Paper] [[Code]](https://github.com/pprp/P