← Open Source
HuangOwen

Awesome-LLM-Compression

Awesome LLM compression research papers and tools.

ListsPaper collections
Open on GitHub
Momentum
+0stars in 24 hours0.0%
1.88k
Stars
133
Forks
+3
This week
46
Contributors
Created 2023-05-30 · Updated 2026-10-03 · #9188 today
Top developers
README

Awesome LLM Compression

[![](https://awesome.re/badge.svg)](https://awesome.re)
 ![](https://img.shields.io/github/stars/HuangOwen/Awesome-LLM-Compression.svg?style=social) 
 ![](https://img.shields.io/github/watchers/HuangOwen/Awesome-LLM-Compression.svg?style=social) 

Awesome LLM compression research papers and tools to accelerate LLM training and inference.

Papers in each category are grouped by year in collapsible blocks, newest first — click a year to expand it. The current year is expanded by default.

Contents

Papers

Survey

2026  ·  6 papers

  • Dynamic Model Routing and Cascading for Efficient LLM Inference: A Survey
    Arxiv 2026 [Paper]

  • A Survey of On-Policy Distillation for Large Language Models
    Arxiv 2026 [Paper]

  • Unifying Data, Memory, and Compute Efficiency in LLM training: A Survey
    TAI 2026 [Paper]

  • From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving
    Arxiv 2026 [Paper]

  • Towards Efficient Large Language Model Serving: A Survey on System-Aware KV Cache Optimization
    ACL Findings 2026 [Paper]

  • Accelerating Masked Diffusion Large Language Models: A Survey of Efficient Inference Techniques
    IJCAI-ECAI 2026 (Survey Track) [Paper]

2025  ·  14 papers

  • Compressed but Compromised? A Study of Jailbreaking in Compressed LLMs
    NeurIPS Lock-LLM Workshop 2025 [Paper] [Blog]

  • Scaling Down, Serving Fast: Compressing and Deploying Efficient LLMs for Recommendation Systems
    EMNLP 2025 [Paper]

  • Key, Value, Compress: A Systematic Exploration of KV Cache Compression Techniques
    CICC 2025 [Paper]

  • Are We There Yet? A Measurement Study of Efficiency for LLM Applications on Mobile Devices
    Arxiv 2025 [Paper]

  • Knowledge Distillation and Dataset Distillation of Large Language Models: Emerging Trends, Challenges, and Future Directions
    Arxiv 2025 [Paper]

  • An Empirical Study on Prompt Compression for Large Language Models
    Building Trust Workshop @ ICLR 2025 2025 [Paper] [PCToolkit]

  • Low-Precision Training of Large Language Models: Methods, Challenges, and Opportunities
    Arxiv 2025 [Paper] [GitHub Page]

  • Optimizing LLMs for Resource-Constrained Environments: A Survey of Model Compression Techniques
    COMPSAC 2025 [Paper]

  • A Survey on Inference Engines for Large Language Models: Perspectives on Optimization and Efficiency
    Arxiv 2025 [Paper] [GitHub Page]

  • EfficientLLM: Efficiency in Large Language Models
    Arxiv 2025 [Paper] [Homepage] [Huggingface Page]

  • KV Cache Compression for Inference Efficiency in LLMs: A Review
    Arxiv 2025 [Paper]

  • A Survey on Parallel Text Generation: From Parallel Decoding to Diffusion Language Models
    Arxiv 2025 [Paper] [Code] Stars

  • A Survey on Collaborating Small and Large Language Models for Performance, Cost-effectiveness, Cloud-edge Privacy, and Trustworthiness
    Arxiv 2025 [Paper] [Code] Stars

  • A Survey on Efficient Large Language Model Training: From Data-centric Perspectives
    ACL 2025 [Paper] [Code] Stars

2024  ·  17 papers

  • Understanding LLMs: A Comprehensive Overview from Training to Inference
    Arxiv 2024 [Paper]

  • Faster and Lighter LLMs: A Survey on Current Challenges and Way Forward
    IJCAI 2024 (Survey Track) [Paper] [GitHub Page]

  • A Survey of Resource-efficient LLM and Multimodal Foundation Models
    Arxiv 2024 [Paper]

  • A Survey on Hardware Accelerators for Large Language Models
    Arxiv 2024 [Paper]

  • A Comprehensive Survey of Compression Algorithms for Language Models
    Arxiv 2024 [Paper]

  • A Survey on Transformer Compression
    Arxiv 2024 [Paper]

  • Model Compression and Efficient Inference for Large Language Models: A Survey
    Arxiv 2024 [Paper]

  • LLM Inference Unveiled: Survey and Roofline Model Insights
    Arxiv 2024 [Paper]

  • A Survey on Knowledge Distillation of Large Language Models
    Arxiv 2024 [Paper] [GitHub Page]

  • Efficient Prompting Methods for Large Language Models: A Survey
    Arxiv 2024 [Paper]

  • Survey on Knowledge Distillation for Large Language Models: Methods, Evaluation, and Application
    Arxiv 2024 [Paper]

  • On-Device Language Models: A Comprehensive Review
    Arxiv 2024 [Paper] [Download On-device LLMs]

  • A Survey of Low-bit Large Language Models: Basics, Systems, and Algorithms
    Arxiv 2024 [Paper]

  • Contextual Compression in Retrieval-Augmented Generation for Large Language Models: A Survey
    Arxiv 2024 [Paper]

  • Prompt Compression for Large Language Models: A Survey
    Arxiv 2024 [Paper]

  • A Comprehensive Study on Quantization Techniques for Large Language Models
    Arxiv 2024 [Paper]

  • A Survey on Large Language Model Acceleration based on KV Cache Management
    TMLR 2025 [Paper]

2023  ·  5 papers

  • A Survey on Model Compression for Large Language Models
    TACL [Paper]

  • The Cost of Compression: Investigating the Impact of Compression on Parametric Knowledge in Language Models
    EMNLP 2023 [Paper] [Code] Stars

  • The Efficiency Spectrum of Large Language Models: An Algorithmic Survey
    Arxiv 2023 [Paper]

  • Efficient Large Language Models: A Survey
    TMLR [Paper] [GitHub Page]

  • Towards Efficient Generative Large Language Model Serving: A Survey from Algorithms to Systems
    ICML 2024 Tutorial [Paper] [Tutorial]

Quantization

2026  ·  56 papers

  • Can Large Language Models Still Explain Themselves? Investigating the Impact of Quantization on Self-Explanations
    EMNLP Findings 2026 [Paper]

  • QSLM: A Performance- and Memory-aware Quantization Framework with Tiered Search Strategy for Spike-driven Language Models
    DATE 2026 [Paper]

  • HAS-VQ: Hessian-Adaptive Sparse Vector Quantization for High-Fidelity LLM Compression
    Arxiv 2026 [Paper] [Code] Stars

  • ARCQuant: Boosting NVFP4 Quantization with Augmented Residual Channels for LLMs
    ACL 2026 [Paper]

  • Sherry: Hardware-Efficient 1.25-Bit Ternary Quantization via Fine-grained Sparsification
    Arxiv 2026 [Paper] [Code] Stars

  • Calibrating Beyond English: Language Diversity for Better Quantized Multilingual LLM
    EACL 2026 [Paper]

  • M2XFP: A Metadata-Augmented Microscaling Data Format for Efficient Low-bit Quantization
    ASPLOS 2026 [Paper]

  • Quartet II: Accurate LLM Pre-Training in NVFP4 by Improved Unbiased Gradient Estimation
    Arxiv 2026 [Paper] [Code] Stars

  • Two-Stage Grid Optimization for Group-wise Quantization of LLMs
    ICASSP 2026 [Paper]

  • Quantized Evolution Strategies: High-precision Fine-tuning of Quantized LLMs at Low-precision Cost
    Arxiv 2026 [Paper] [Code] Stars

  • QVLA: Not All Channels Are Equal in Vision-Language-Action Model's Quantization
    ICLR 2026 [Paper]

  • TurboBoA: Faster and Exact Attention-aware Quantization without Backpropagation
    ICLR 2026 [Paper]

  • RaBiT: Residual-Aware Binarization Training for Accurate and Efficient LLMs
    ICML 2026 [Paper]

  • NanoQuant: Efficient Sub-1-Bit Quantization of Large Language Models
    ICML 2026 [Paper]

  • On the Importance of a Multi-Scale Calibration for Quantization
    ICASSP 2026 [Paper]

  • QTALE: Quantization-Robust Token-Adaptive Layer Execution for LLMs
    ICML 2026 [Paper]

  • QuRL: Efficient Reinforcement Learning with Quantized Rollout
    ICLR 2026 [Paper]

  • SPQ: An Ensemble Technique for Large Language Model Compression
    LREC 2026 [Paper] [Code] Stars

  • Quant Experts: Token-aware Adaptive Error Reconstruction with Mixture of Experts for Large Vision-Language Models Quantization
    CVPR 2026 [Paper]

  • MASQuant: Modality-Aware Smoothing Quantization for Multimodal Large Language Models
    CVPR 2026 [Paper]

  • SliderQuant: Accurate Post-Training Quantization for LLMs
    ICLR 2026 [Paper] [Code] Stars

  • OneComp: One-Line Revolution for Generative AI Model Compression
    Arxiv 2026 [Paper] [Code] Stars

  • Fast NF4 Dequantization Kernels for Large Language Model Inference
    ASPLOS 2026 Workshop [Paper]

  • RUQuant: Towards Refining Uniform Quantization for Large Language Models
    KDD 2026 [Paper]

  • SEPTQ: A Simple and Effective Post-Training Quantization Paradigm for Large Language Models
    KDD 2025 [Paper]

  • ReSpinQuant: Efficient Layer-Wise LLM Quantization via Subspace Residual Rotation Approximation
    ICML 2026 [Paper]

  • Robust Ultra Low-Bit Post-Training Quantization via Stable Diagonal Curvature Estimate
    MLSys 2026 [Paper]

  • AQPIM: Breaking the PIM Capacity Wall for LLMs with In-Memory Activation Quantization
    HPCA 2026 [Paper]

  • From Signal Degradation to Computation Collapse: Uncovering the Two Failure Modes of LLM Quantization
    ACL Findings 2026 [Paper]

  • Statistically-Lossless Quantization of Large Language Models
    Arxiv 2026 [Paper] [Code] Stars

  • EdgeRazor: A Lightweight Framework for Large Language Models via Mixed-Precision Quantization-Aware Distillation
    Arxiv 2026 [Paper] [Code] [Model] [Playground] Stars

  • OSAQ: Outlier Self-Absorption for Accurate Low-bit LLM Quantization
    ICML 2026 [Paper]

  • Normalized Architectures are Natively 4-Bit
    Arxiv 2026 [Paper] [Code] Stars

  • RQ-MoE: Residual Quantization via Mixture of Experts for Efficient Input-Dependent Vector Compression
    ICML 2026 [Paper]

  • XFP: Quality-Targeted Adaptive Codebook Quantization with Sparse Outlier Separation for LLM Inference
    Arxiv 2026 [Paper] [Code] Stars

  • GEMQ: Global Expert-Level Mixed-Precision Quantization for MoE LLMs
    ICML 2026 [Paper] [Code] Stars

  • EVA: Accelerating LLM Decoding via an Efficient Vector Quantization Architecture
    ISCA 2026 [Paper]

  • Influence-Inspired Spectral Rotations for Extreme Low-Bit LLM Quantization
    Arxiv 2026 [Paper] [Code] Stars

  • InfoQuant: Shaping Activation Distributions for Low-Bit LLM Quantization
    Arxiv 2026 [Paper] [Code] Stars

  • LFQ: Logit-aware Final-block Quantization for Boosting the Generation Quality of Low-Bit Quantized LLMs
    ICML 2026 [Paper]

  • LiftQuant: Continuous Bit-Width LLM via Dimensional Lifting and Projection
    ICML 2026 [Paper]

  • UniSVQ: 2-bit Unified Scalar-Vector Quantization
    ICML 2026 [Paper]

  • LC-QAT: Data-Efficient 2-Bit QAT for LLMs via Linear-Constrained Vector Quantization
    ICML 2026 [Paper]

  • TWLA: Achieving Ternary Weights and Low-Bit Activations for LLMs via Post-Training Quantization
    ICML 2026 [Paper]

  • CAT-Q: Cost-efficient and Accurate Ternary Quantization for LLMs
    ICML 2026 [Paper] [Code] Stars

  • Quantize the Target, Quantize the Drafter: Efficient Inference with Qwen3.5-4B
    Arxiv 2026 [Paper] [Code] Stars

  • KronQ: LLM Quantization via Kronecker-Factored Hessian
    COLM 2026 [Paper]

  • Rethinking Small VLM Quantization: From Component-Wise Analysis to Hardware-Aware Edge Deployment
    ICML 2026 Workshop [Paper]

  • Reliability Scaling Laws for Quantized Large Language Models
    TMLR 2026 [Paper]

  • PolyQ: Codesigning End-to-End Quantization Framework for Scalable Edge CPU LLM Inference
    ICCAD 2026 [Paper]

  • ExaGEMM: Exploration Framework for CPU-Driven ML Inference via Associative In-Register Computing for Low-Bit GEMM
    ICCAD 2026 [Paper]

  • GyRot: Leveraging Hidden Synergy between Rotation and Fine-grained Group Quantization for Low-bit LLM Inference
    HPCA 2026 [Paper]

  • Attend to Your Own Thoughts: Breaking the Barrier for Post-Training Quantization of Reasoning LLMs through the Lens of 1.58-Bit Quantization
    Arxiv 2026 [Paper] [Code] Stars

  • ARCHead: Activation-Metric Residual Correction for Large Language Model Output Heads
    Arxiv 2026 [Paper] [Code] Stars

  • Compress and Forget: bitsandbytes Quantization Amplifies Proactive Interference in LLMs
    Arxiv 2026 [Paper] [Code] Stars

  • Jacobian-guided Noise Injection for Quantization Robustness in Large Language Models
    ICML 2026 [Paper]

2025  ·  118 papers

  • HALO: Hadamard-Assisted Lower-Precision Optimization for LLMs
    Arxiv 2025 [Paper]

  • RaZeR: Pushing the Limits of NVFP4 Quantization with Redundant Zero Remapping
    Arxiv 2025 [Paper]

  • FlexQuant: Elastic Quantization Framework for Locally Hosted LLM on Edge Devices
    Arxiv 2025 [Paper]

  • Rethinking Post-Training Quantization: Introducing a Statistical Pre-Calibration Approach
    Arxiv 2025 [Paper]

  • Qrazor: Reliable and effortless 4-bit llm quantization by significant data razoring
    Arxiv 2025 [Paper]

  • OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting
    ICLR 2025 [Paper] [Code] Stars

  • SwiftPrune: Hessian-Free Weight Pruning for Large Language Models
    Arxiv 2025 [Paper]

  • Progressive Binarization with Semi-Structured Pruning for LLMs
    Arxiv 2025 [Paper] [Code] Stars

  • Physics-Inspired Binary Neural Networks: Interpretable Compression with Theoretical Guarantees
    Arxiv 2025 [Paper]

  • QuEST: Stable Training of LLMs with 1-Bit Weights and Activations
    Arxiv 2025 [Paper]

  • ParetoQ: Scaling Laws in Extremely Low-bit LLM Quantization
    NeurIPS 2025 [Paper]

  • Systematic Outliers in Large Language Models
    ICLR 2025 [Paper] [Code] Stars

  • Can Post-Training Quantization Benefit from an Additional QLoRA Integration?
    NAACL 2025 [Paper]

  • 1bit-Merging: Dynamic Quantized Merging for Large Language Models
    Arxiv 2025 [Paper]

  • Towards Efficient Pre-training: Exploring FP4 Precision in Large Language Models
    Arxiv 2025 [Paper]

  • Continual Quantization-Aware Pre-Training: When to transition from 16-bit to 1.58-bit pre-training for BitNet language models?
    Arxiv 2025 [Paper]

  • QuZO: Quantized Zeroth-Order Fine-Tuning for Large Language Models
    Arxiv 2025 [Paper]

  • Benchmarking Post-Training Quantization in LLMs: Comprehensive Taxonomy, Unified Evaluation, and Comparative Analysis
    Arxiv 2025 [Paper]

  • Compression Scaling Laws:Unifying Sparsity and Quantization
    Arxiv 2025 [Paper]

  • M-ANT: Efficient Low-bit Group Quantization for LLMs via Mathematically Adaptive Numerical Type
    Arxiv 2025 [Paper]

  • Identifying Sensitive Weights via Post-quantization Integral
    Arxiv 2025 [Paper]

  • RSQ: Learning from Important Tokens Leads to Better Quantized LLMs
    Arxiv 2025 [Paper] [Code] Stars

  • VQ-LLM: High-performance Code Generation for Vector Quantization Augmented LLM Inference
    HPCA 2025 [Paper]

  • Universality of Layer-Level Entropy-Weighted Quantization Beyond Model Architecture and Size
    Arxiv 2025 [Paper]

  • Towards Superior Quantization Accuracy: A Layer-sensitive Approach
    Arxiv 2025 [Paper]

  • MergeQuant: Accurate 4-bit Static Quantization of Large Language Models by Channel-wise Calibration
    Arxiv 2025 [Paper]

  • ClusComp: A Simple Paradigm for Model Compression and Efficient Finetuning
    Arxiv 2025 [Paper]

  • DynaMo: Runtime Switchable Quantization for MoE with Cross-Dataset Adaptation
    DATE 2026 [Paper]

  • Cocktail: Chunk-Adaptive Mixed-Precision Quantization for Long-Context LLM Inference
    DATE 2025 [Paper]

  • GPTQv2: Efficient Finetuning-Free Quantization for Asymmetric Calibration
    ICML 2025 [Paper] [Code] Stars

  • Task-Circuit Quantization: Leveraging Knowledge Localization and Interpretability for Compression
    COLM 2025 [Paper] [Code] Stars

  • Quantization Error Propagation: Revisiting Layer-Wise Post-Training Quantization
    NeurIPS 2025 [Paper] [Code] Stars

  • RaanA: A Fast, Flexible, and Data-Efficient Post-Training Quantization Algorithm
    Arxiv 2025 [Paper] [Code] Stars

  • Achieving binary weight and activation for LLMs using Post-Training Quantization
    Arxiv 2025 [Paper]

  • DL-QAT: Weight-Decomposed Low-Rank Quantization-Aware Training for Large Language Models
    EMNLP 2024 [Paper]

  • Gradual Binary Search and Dimension Expansion : A general method for activation quantization in LLMs
    Arxiv 2025 [Paper]

  • FGMP: Fine-Grained Mixed-Precision Weight and Activation Quantization for Hardware-Accelerated LLM Inference
    Arxiv 2025 [Paper]

  • BitNet v2: Native 4-bit Activations with Hadamard Transformation for 1-bit LLMs
    Arxiv 2025 [Paper]

  • FineQ: Software-Hardware Co-Design for Low-Bit Fine-Grained Mixed-Precision Quantization of LLMs
    DATE 2025 [Paper]

  • Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models
    Arxiv 2025 [Paper]

  • ICQuant: Index Coding enables Low-bit LLM Quantization
    Arxiv 2025 [Paper]

  • Radio: Rate-Distortion Optimization for Large Language Model Compression
    ICML 2025 [Paper]

  • Balancing Fidelity and Plasticity: Aligning Mixed-Precision Fine-Tuning with Linguistic Hierarchies
    Arxiv 2025 [Paper]

  • MoEQuant: Enhancing Quantization for Mixture-of-Experts Large Language Models via Expert-Balanced Sampling and Affinity Guidance
    ICML 2025 [Paper]

  • Grouped Sequency-arranged Rotation: Optimizing Rotation Transformation for Quantization for Free
    Arxiv 2025 [Paper]

  • Improving Block-Wise LLM Quantization by 4-bit Block-Wise Optimal Float (BOF4): Analysis and Variations
    Arxiv 2025 [Paper]

  • GuidedQuant: Large Language Model Quantization via Exploiting End Loss Guidance
    ICML 2025 [Paper] [Code] Stars

  • QuantX: A Framework for Hardware-Aware Quantization of Generative AI Workloads
    Arxiv 2025 [Paper]

  • An Extra RMSNorm is All You Need for Fine Tuning to 1.58 Bits
    Arxiv 2025 [Paper]

  • ITERA-LLM: Boosting Sub-8-Bit Large Language Model Inference via Iterative Tensor Decomposition
    Arxiv 2025 [Paper]

  • Fine-tuning Quantized Neural Networks with Zeroth-order Optimization
    ICLR 2026 [Paper]

  • Scaling Law for Quantization-Aware Training
    Arxiv 2025 [Paper]

  • Quaff: Quantized Parameter-Efficient Fine-Tuning under Outlier Spatial Stability Hypothesis
    ACL 2025 [Paper] [Code] Stars

  • Is (Selective) Round-To-Nearest Quantization All You Need?
    Arxiv 2025 [Paper]

  • NeUQI: Near-Optimal Uniform Quantization Parameter Initialization
    ICML 2026 [Paper]

  • LoTA-QAF: Lossless Ternary Adaptation for Quantization-Aware Fine-Tuning
    Arxiv 2025 [Paper] [Code] Stars

  • FP4 All the Way: Fully Quantized Training of LLMs
    Arxiv 2025 [Paper] [Code] Stars

  • FireQ: Fast INT4-FP8 Kernel and RoPE-aware Quantization for LLM Inference Acceleration
    Arxiv 2025 [Paper]

  • Rethinking the Outlier Distribution in Large Language Models: An In-depth Study
    Arxiv 2025 [Paper]

  • Assigning Distinct Roles to Quantized and Low-Rank Matrices Toward Optimal Weight Decomposition
    ACL Findings 2025 [Paper]

  • Unifying Uniform and Binary-coding Quantization for Accurate Compression of Large Language Models
    ACL 2025 [Paper]

  • FPTQuant: Function-Preserving Transforms for LLM Quantization
    ICML 2026 [Paper]

  • BAQ: Efficient Bit Allocation Quantization for Large Language Models
    Arxiv 2025 [Paper] [Code] Stars

  • MoQAE: Mixed-Precision Quantization for Long-Context LLM Inference via Mixture of Quantization-Aware Experts
    ACL 2025 [Paper]

  • Unifying Block-wise PTQ and Distillation-based QAT for Progressive Quantization toward 2-bit Instruction-Tuned LLMs
    Arxiv 2025 [Paper]

  • Boost Post-Training Quantization via Null Space Optimization for Large Language Models
    Arxiv 2025 [Paper] [Code] Stars

  • FlexQuant: A Flexible and Efficient Dynamic Precision Switching Framework for LLM Quantization
    Arxiv 2025 [Paper]

  • BTC-LLM: Efficient Sub-1-Bit LLM Quantization via Learnable Transformation and Binary Codebook
    Arxiv 2025 [Paper]

  • ROSAQ: Rotation-based Saliency-Aware Weight Quantization for Efficiently Compressing Large Language Models
    Arxiv 2025 [Paper]

  • LittleBit: Ultra Low-Bit Quantization via Latent Factorization
    NeurIPS 2025 [Paper]

  • BASE-Q: Bias and Asymmetric Scaling Enhanced Rotational Quantization for Large Language Models
    Arxiv 2025 [Paper] [Code] Stars

  • UltraSketchLLM: Sub-1-Bit LLM Compression via Sketch and Hardware-Friendly Operators
    DAC 2026 [Paper]

  • DBellQuant: Breaking the Bell with Double-Bell Transformation for LLMs Post Training Binarization
    Arxiv 2025 [Paper]

  • any4: Learned 4-bit Numeric Representation for LLMs
    ICML 2025 [Paper]

  • CCQ: Convolutional Code for Extreme Low-bit Quantization in LLMs
    Arxiv 2025 [Paper]

  • First-Order Error Matters: Accurate Compensation for Quantized Large Language Models
    AAAI 2026 [Paper]

  • PoTPTQ: A Two-step Power-of-Two Post-training for LLMs
    ECAI 2025 [Paper]

  • EAC-MoE: Expert-Selection Aware Compressor for Mixture-of-Experts Large Language Models
    ACL 2025 [Paper]

  • MicroMix: Efficient Mixed-Precision Quantization with Microscaling Formats for Large Language Models
    Arxiv 2025 [Paper] [Code] Stars

  • VLMQ: Token Saliency-Driven Post-Training Quantization for Vision-language Models
    Arxiv 2025 [Paper]

  • FlashCommunication V2: Bit Splitting and Spike Reserving for Any Bit Communication
    Arxiv 2025 [Paper]

  • FlexQ: Efficient Post-training INT6 Quantization for LLM Serving via Algorithm-System Co-Design
    Arxiv 2025 [Paper] [Code] Stars

  • Provable Post-Training Quantization: Theoretical Analysis of OPTQ and Qronos
    Arxiv 2025 [Paper]

  • iFairy: the First 2-bit Complex LLM with All Parameters in ${\pm1, \pm i}$
    Arxiv 2025 [Paper]

  • Pushing the Envelope of LLM Inference on AI-PC and Intel GPUs
    Arxiv 2025 [Paper]

  • Rethinking 1-bit Optimization Leveraging Pre-trained Large Language Models
    Arxiv 2025 [Paper]

  • Efficient Edge LLMs Deployment via HessianAware Quantization and CPU GPU Collaborative
    Arxiv 2025 [Paper]

  • Profiling Large Language Model Inference on Apple Silicon: A Quantization Perspective
    Arxiv 2025 [Paper]

  • LLM Compression: How Far Can We Go in Balancing Size and Performance?
    RANLP 2025 [Paper]

  • DLLMQuant: Quantizing Diffusion-based Large Language Models
    Arxiv 2025 [Paper]

  • Quantization Meets dLLMs: A Systematic Study of Post-training Quantization for Diffusion LLMs
    Machine Intelligence Research 2025 [Paper]

  • Systematic Characterization of LLM Quantization: A Performance, Energy, and Quality Perspective
    Arxiv 2025 [Paper]

  • Interpreting the Effects of Quantization on LLMs
    AACL 2025 [Paper]

  • Task-Stratified Knowledge Scaling Laws for Post-Training Quantized Large Language Models
    ACL Findings 2026 [Paper]

  • APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration
    TCAD 2025 [Paper]

  • Quantized but Deceptive? A Multi-Dimensional Truthfulness Evaluation of Quantized LLMs
    EMNLP 2025 [Paper]

  • The Uneven Impact of Post-Training Quantization in Machine Translation
    Arxiv 2025 [Paper]

  • BitROM: Weight Reload-Free CiROM Architecture Towards Billion-Parameter 1.58-bit LLM Inference
    ASP-DAC 2026 [Paper]

  • AMQ: Enabling AutoML for Mixed-precision Weight-Only Quantization of Large Language Models
    EMNLP 2025 [Paper]

  • Fair-GPTQ: Bias-Aware Quantization for Large Language Models
    Arxiv 2025 [Paper]

  • QWHA: Quantization-Aware Walsh-Hadamard Adaptation for Parameter-Efficient Fine-Tuning on Large Language Models
    ICLR 2026 [Paper]

  • Q-Palette: Fractional-Bit Quantizers Toward Optimal Bit Allocation for Efficient LLM Deployment
    NeurIPS 2025 [Paper]

  • AnyBCQ: Hardware Efficient Flexible Binary-Coded Quantization for Multi-Precision LLMs
    ICLR 2026 [Paper]

  • QeRL: Beyond Efficiency -- Quantization-enhanced Reinforcement Learning for LLMs
    Arxiv 2025 [Paper] [Code] Stars

  • F-BFQ: Flexible Block Floating-Point Quantization Accelerator for LLMs
    ISCA 2025 Workshop [Paper]

  • Learning Grouped Lattice Vector Quantizers for Low-Bit LLM Compression
    NeurIPS 2025 [Paper]

  • A Convergence Analysis of Adaptive Optimizers under Floating-point Quantization
    ICLR 2026 [Paper]

  • FALQON: Accelerating LoRA Fine-tuning with Low-Bit Floating-Point Arithmetic
    NeurIPS 2025 [Paper]

  • TetraJet-v2: Accurate NVFP4 Training for Large Language Models with Oscillation Suppression and Outlier Control
    ICML 2026 [Paper]

  • DartQuant: Efficient Rotational Distribution Calibration for LLM Quantization
    NeurIPS 2025 [Paper]

  • You Had One Job: Per-Task Quantization Using LLMs' Hidden Representations
    ICML 2026 Workshop [Paper]

  • P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats
    ISCA 2026 [Paper]

  • ParoQuant: Pairwise Rotation Quantization for Efficient Reasoning LLM Inference
    ICLR 2026 [Paper] [Code] Stars

  • SpecQuant: Spectral Decomposition and Adaptive Truncation for Ultra-Low-Bit LLMs Quantization
    AAAI 2026 [Paper]

  • T-SAR: A Full-Stack Co-design for CPU-Only Ternary LLM Inference via In-Place SIMD ALU Reorganization
    DATE 2026 [Paper]

  • Enhancing Trustworthiness with Mixed Precision: Benchmarks, Opportunities, and Challenges
    ASP-DAC 2026 [Paper]

  • SignRoundV2: Toward Closing the Performance Gap in Extremely Low-Bit Post-Training Quantization for LLMs
    Arxiv 2025 [Paper] [Code] Stars

  • CodeGEMM: A Codebook-Centric Approach to Efficient GEMM in Quantized LLMs
    NeurIPS 2025 [Paper]

2024  ·  162 papers

  • FlightLLM: Efficient Large Language Model Inference with a Complete Mapping Flow on FPGA
    FPGA 2024 [Paper]

  • Extreme Compression of Large Language Models via Additive Quantization
    ICML 2024 [Paper] [Code] Stars

  • Quantized Side Tuning: Fast and Memory-Efficient Tuning of Quantized Large Language Models
    Arxiv 2024 [Paper]

  • Inferflow: an Efficient and Highly Configurable Inference Engine for Large Language Models
    Arxiv 2024 [Paper]

  • FP6-LLM: Efficiently Serving Large Language Models Through FP6-Centric Algorithm-System Co-Design
    USENIX ATC 2024 [Paper]

  • Can Large Language Models Understand Context?
    EACL Findings 2024 [Paper]

  • Squat: Quant Small Language Models on the Edge
    ICCAD 2025 [Paper] [Code] Stars

  • LQER: Low-Rank Quantization Error Reconstruction for LLMs
    ICML 2024 [Paper]

  • BiLLM: Pushing the Limit of Post-Training Quantization for LLMs
    Arxiv 2024 [Paper] [Code] Stars

  • QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks
    ICML 2024 [Paper] [Code] Stars

  • L4Q: Parameter Efficient Quantization-Aware Training on Large Language Models via LoRA-wise LSQ
    Arxiv 2024 [Paper]

  • TP-Aware Dequantization
    Arxiv 2024 [Paper]

  • ApiQ: Finetuning of 2-Bit Quantized Large Language Model
    EMNLP 2024 [Paper]

  • Accurate LoRA-Finetuning Quantization of LLMs via Information Retention
    Arxiv 2024 [Paper] [Code] Stars

  • BitDelta: Your Fine-Tune May Only Be Worth One Bit
    NeurIPS 2024 [Paper] [Code] Stars

  • QDyLoRA: Quantized Dynamic Low-Rank Adaptation for Efficient Large Language Model Tuning
    EMNLP 2024 Industry Track [Paper]

  • Any-Precision LLM: Low-Cost Deployment of Multiple, Different-Sized LLMs
    ICML 2024 [Paper]

  • BitDistiller: Unleashing the Potential of Sub-4-Bit LLMs via Self-Distillation
    ACL 2024 [Paper] [Code] Stars

  • OneBit: Towards Extremely Low-bit Large Language Models
    NeurIPS 2024 [Paper]

  • DB-LLM: Accurate Dual-Binarization for Efficient LLMs
    ACL Findings 2024 [Paper]

  • WKVQuant: Quantizing Weight and Key/Value Cache for Large Language Models Gains More
    Arxiv 2024 [Paper]

  • GPTVQ: The Blessing of Dimensionality for LLM Quantization
    Arxiv 2024 [Paper] [Code] Stars

  • APTQ: Attention-aware Post-Training Mixed-Precision Quantization for Large Language Models
    DAC 2024 [Paper]

  • A Comprehensive Evaluation of Quantization Strategies for Large Language Models
    ACL Findings 2024 [Paper]

  • Evaluating Quantized Large Language Models
    Arxiv 2024 [Paper]

  • FlattenQuant: Breaking Through the Inference Compute-bound for Large Language Models with Per-tensor Quantization
    Arxiv 2024 [Paper]

  • LLM-PQ: Serving LLM on Heterogeneous Clusters with Phase-Aware Partition and Adaptive Quantization
    Arxiv 2024 [Paper]

  • IntactKV: Improving Large Languagze Model Quantization by Keeping Pivot Tokens Intact
    ACL Findings 2024 [Paper] [Code] Stars

  • On the Compressibility of Quantized Large Language Models
    Arxiv 2024 [Paper]

  • EasyQuant: An Efficient Data-free Quantization Algorithm for LLMs
    Arxiv 2024 [Paper]

  • What Makes Quantization for Large Language Models Hard? An Empirical Study from the Lens of Perturbation
    Arxiv 2024 [Paper]

  • AffineQuant: Affine Transformation Quantization for Large Language Models
    ICLR 2024 [Paper] [Code] Stars

  • Oh! We Freeze: Improving Quantized Knowledge Distillation via Signal Propagation Analysis for Large Language Models
    ICLR Practical ML for Low Resource Settings Workshop 2024 [Paper]

  • Accurate Block Quantization in LLMs with Outliers
    Arxiv 2024 [Paper]

  • QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs
    Arxiv 2024 [Paper] [Code] Stars

  • Minimize Quantization Output Error with Bias Compensation
    Arxiv 2024 [Paper] [Code] Stars

  • Cherry on Top: Parameter Heterogeneity and Quantization in Large Language Models
    Arxiv 2024 [Paper]

  • Fine-Tuning, Quantization, and LLMs: Navigating Unintended Outcomes
    Arxiv 2024 [Paper]

  • Quantization of Large Language Models with an Overdetermined Basis
    Arxiv 2024 [Paper]

  • An empirical study of LLaMA3 quantization: from LLMs to MLLMs
    Arxiv 2024 [Paper] [Code] Stars

  • How to Parameterize Asymmetric Quantization Ranges for Quantization-Aware Training
    Arxiv 2024 [Paper]

  • Mitigating the Impact of Outlier Channels for Language Model Quantization with Activation Regularization
    Arxiv 2024 [Paper] [Code] Stars

  • When Quantization Affects Confidence of Large Language Models?
    NAACL 2024 [Paper]

  • QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving
    Arxiv 2024 [Paper] [Code] Stars

  • Learning from Students: Applying t-Distributions to Explore Accurate and Efficient Formats for LLMs
    ICML 2024 [Paper]

  • LLMC: Benchmarking Large Language Model Quantization with a Versatile Compression Toolkit
    EMNLP 2024 [Paper] [Code] Stars

  • SKVQ: Sliding-window Key and Value Cache Quantization for Large Language Models
    Arxiv 2024 [Paper]

  • Post Training Quantization of Large Language Models with Microscaling Formats
    Arxiv 2024 [Paper]

  • Edge Intelligence Optimization for Large Language Model Inference with Batching and Quantization
    Arxiv 2024 [Paper]

  • SliM-LLM: Salience-Driven Mixed-Precision Quantization for Large Language Models
    Arxiv 2024 [Paper] [Code] Stars

  • OAC: Output-adaptive Calibration for Accurate Post-training Quantization
    Arxiv 2024 [Paper]

  • PV-Tuning: Beyond Straight-Through Estimation for Extreme LLM Compression
    Arxiv 2024 [Paper]

  • SpinQuant -- LLM quantization with learned rotations
    ICLR 2025 [Paper]

  • Compressing Large Language Models using Low Rank and Low Precision Decomposition
    NeurIPS 2024 [Paper] [Code] Stars

  • Athena: Efficient Block-Wise Post-Training Quantization for Large Language Models Using Second-Order Matrix Derivative Information
    Arxiv 2024 [Paper]

  • Exploiting LLM Quantization
    Arxiv 2024 [Paper]

  • One QuantLLM for ALL: Fine-tuning Quantized LLMs Once for Efficient Deployments
    Arxiv 2024 [Paper]

  • LCQ: Low-Rank Codebook based Quantization for Large Language Models
    Arxiv 2024 [Paper]

  • LoQT: Low Rank Adapters for Quantized Training
    Arxiv 2024 [Paper] [Code] Stars

  • CLAQ: Pushing the Limits of Low-Bit Post-Training Quantization for LLMs
    Arxiv 2024 [Paper]

  • I-LLM: Efficient Integer-Only Inference for Fully-Quantized Low-Bit Large Language Models
    Arxiv 2024 [Paper]

  • Outliers and Calibration Sets have Diminishing Effect on Quantization of Modern LLMs
    Arxiv 2024 [Paper]

  • DuQuant: Distributing Outliers via Dual Transformation Makes Stronger Quantized LLMs
    NeurIPS 2024 [Paper] [Code] Stars

  • ShiftAddLLM: Accelerating Pretrained LLMs via Post-Training Multiplication-Less Reparameterization
    NeurIPS 2024 [Paper] [Code] Stars

  • Low-Rank Quantization-Aware Training for LLMs
    Arxiv 2024 [Paper]

  • TernaryLLM: Ternarized Large Language Model
    Arxiv 2024 [Paper]

  • Examining Post-Training Quantization for Mixture-of-Experts: A Benchmark
    Arxiv 2024 [Paper] [Code] Stars

  • Delta-CoMe: Training-Free Delta-Compression with Mixed-Precision for Large Language Models
    NeurIPS 2024 [Paper]

  • QQQ: Quality Quattuor-Bit Quantization for Large Language Models
    Arxiv 2024 [Paper] [Code] Stars

  • QTIP: Quantization with Trellises and Incoherence Processing
    NeurIPS 2024 [Paper] [Code] Stars

  • Prefixing Attention Sinks can Mitigate Activation Outliers for Large Language Model Quantization
    EMNLP 2024 [Paper]

  • Mixture of Scales: Memory-Efficient Token-Adaptive Binarization for Large Language Models
    NeurIPS 2024 [Paper]

  • Tender: Accelerating Large Language Models via Tensor Decomposition and Runtime Requantization
    ISCA 2024 [Paper]

  • SDQ: Sparse Decomposed Quantization for LLM Inference
    Arxiv 2024 [Paper]

  • Attention-aware Post-training Quantization without Backpropagation
    ICML 2025 [Paper]

  • EDGE-LLM: Enabling Efficient Large Language Model Adaptation on Edge Devices via Layerwise Unified Compression and Adaptive Layer Tuning and Voting
    Arxiv 2024 [Paper] [Code] Stars

  • Compensate Quantization Errors: Make Weights Hierarchical to Compensate Each Other
    Arxiv 2024 [Paper]

  • Layer-Wise Quantization: A Pragmatic and Effective Method for Quantizing LLMs Beyond Integer Bit-Levels
    Arxiv 2024 [Paper] [Code] Stars

  • CDQuant: Greedy Coordinate Descent for Accurate LLM Quantization
    Arxiv 2024 [Paper]

  • OutlierTune: Efficient Channel-Wise Quantization for Large Language Models
    Arxiv 2024 [Paper]

  • T-MAC: CPU Renaissance via Table Lookup for Low-Bit LLM Deployment on Edge
    EuroSys 2025 [Paper] [Code] Stars

  • GPTQT: Quantize Large Language Models Twice to Push the Efficiency
    ICORIS 2024 [Paper]

  • Improving Conversational Abilities of Quantized Large Language Models via Direct Preference Alignment
    ACL 2024 [Paper]

  • How Does Quantization Affect Multilingual LLMs?
    EMNLP Findings 2024 [Paper]

  • RoLoRA: Fine-tuning Rotated Outlier-free LLMs for Effective Weight-Activation Quantization
    EMNLP Findings 2024 [Paper] [Code] Stars

  • Q-GaLore: Quantized GaLore with INT4 Projection and Layer-Adaptive Low-Rank Gradients
    Arxiv 2024 [Paper] [Code] Stars

  • FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision
    Arxiv 2024 [Paper] [Code] Stars

  • Accuracy is Not All You Need
    Arxiv 2024 [Paper]

  • BitNet b1.58 Reloaded: State-of-the-art Performance Also on Smaller Networks
    Arxiv 2024 [Paper]

  • LeanQuant: Accurate Large Language Model Quantization with Loss-Error-Aware Grid
    ICLR 2025 [Paper]

  • Fast Matrix Multiplications for Lookup Table-Quantized LLMs
    EMNLP Findings 2024 [Paper] [Code] Stars

  • EfficientQAT: Efficient Quantization-Aware Training for Large Language Models
    ACL 2025 [Paper] [Code] Stars

  • LRQ: Optimizing Post-Training Quantization for Large Language Models by Learning Low-Rank Weight-Scaling Matrices
    NAACL 2025 [Paper]

  • Exploring Quantization for Efficient Pre-Training of Transformer Language Models
    EMNLP Findings 2024 [Paper] [Code] Stars

  • Spectra: Surprising Effectiveness of Pretraining Ternary Language Models at Scale
    Arxiv 2024 [Paper] [Code] Stars

  • Mamba-PTQ: Outlier Channels in Recurrent Large Language Models
    Efficient Systems for Foundation Models Workshop @ ICML 2024 [Paper]

  • Compensate Quantization Errors+: Quantized Models Are Inquisitive Learners
    Arxiv 2024 [Paper]

  • Accurate and Efficient Fine-Tuning of Quantized Large Language Models Through Optimal Balance
    Arxiv 2024 [Paper] [Code] Stars

  • STBLLM: Breaking the 1-Bit Barrier with Structured Binary LLMs
    Arxiv 2024 [Paper]

  • Advancing Multimodal Large Language Models with Quantization-Aware Scale Learning for Efficient Adaptation
    ACM MM 2024 [Paper]

  • ABQ-LLM: Arbitrary-Bit Quantized Inference Acceleration for Large Language Models
    AAAI 2025 [Paper]

  • MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models
    Arxiv 2024 [Paper] [Code (Marlin)] [Code (Sparse Marlin)] Stars

  • Matmul or No Matmal in the Era of 1-bit LLMs
    Arxiv 2024 [Paper]

  • MobileQuant: Mobile-friendly Quantization for On-device Language Models
    EMNLP Findings 2024 [Paper] [Code] Stars

  • GIFT-SW: Gaussian noise Injected Fine-Tuning of Salient Weights for LLMs
    Arxiv 2024 [Paper] [Code] Stars

  • Foundations of Large Language Model Compression -- Part 1: Weight Quantization
    Arxiv 2024 [Paper]

  • OPAL: Outlier-Preserved Microscaling Quantization A ccelerator for Generative Large Language Models
    DAC 2024 [Paper]

  • VPTQ: Extreme Low-bit Vector Post-Training Quantization for Large Language Models
    EMNLP 2024 [Paper] [Code] Stars

  • Scaling FP8 training to trillion-token LLMs
    Arxiv 2024 [Paper]

  • Accumulator-Aware Post-Training Quantization
    Arxiv 2024 [Paper]

  • Efficient Arbitrary Precision Acceleration for Large Language Models on GPU Tensor Cores
    ASP-DAC 2025 [Paper]

  • Rotated Runtime Smooth: Training-Free Activation Smoother for accurate INT4 inference
    Arxiv 2024 [Paper] [Code] Stars

  • EXAQ: Exponent Aware Quantization For LLMs Acceleration
    Arxiv 2024 [Paper]

  • ARB-LLM: Alternating Refined Binarizations for Large Language Models
    Arxiv 2024 [Paper] [Code] Stars

  • PrefixQuant: Eliminating Outliers by Prefixed Tokens for Large Language Models Quantization
    Arxiv 2024 [Paper] [Code] Stars

  • Sketch to Adapt: Fine-Tunable Sketches for Efficient LLM Adaptation
    ICML 2025 [Paper]

  • Scaling Laws For Mixed Quantization
    Arxiv 2024 [Paper]

  • Q-VLM: Post-training Quantization for Large Vision-Language Models
    NeurIPS 2024 [Paper] [Code] Stars

  • CrossQuant: A Post-Training Quantization Method with Smaller Quantization Kernel for Precise Large Language Model Compression
    Arxiv 2024 [Paper]

  • FlatQuant: Flatness Matters for LLM Quantization
    ICML 2025 [Paper] [Code] Stars

  • DeltaDQ: Ultra-High Delta Compression for Fine-Tuned LLMs via Group-wise Dropout and Separate Quantization
    Arxiv 2024 [Paper]

  • QEFT: Quantization for Efficient Fine-Tuning of LLMs
    EMNLP Findings 2024 [Paper] [Code] Stars

  • Continuous Approximations for Improving Quantization Aware Training of LLMs
    Arxiv 2024 [Paper]

  • DAQ: Density-Aware Post-Training Weight-Only Quantization For LLMs
    Arxiv 2024 [Paper]

  • COMET: Towards Partical W4A4KV4 LLMs Serving
    Arxiv 2024 [Paper]

  • Scaling laws for post-training quantized large language models
    Arxiv 2024 [Paper]

  • Channel-Wise Mixed-Precision Quantization for Large Language Models
    Arxiv 2024 [Paper]

  • Understanding the difficulty of low-precision post-training quantization of large language models
    Arxiv 2024 [Paper]

  • QuAILoRA: Quantization-Aware Initialization for LoRA
    NeurIPS Workshop on Efficient Natural Language and Speech Processing (ENLSP-IV) 2024 [Paper]

  • SDP4Bit: Toward 4-bit Communication Quantization in Sharded Data Parallelism for LLM Training
    NeurIPS 2024 [Paper]

  • Pyramid Vector Quantization for LLMs
    Arxiv 2024 [Paper]

  • TesseraQ: Ultra Low-Bit LLM Post-Training Quantization with Block Reconstruction
    Arxiv 2024 [Paper] [Code] Stars

  • COAT: Compressing Optimizer states and Activation for Memory-Efficient FP8 Training
    ICLR 2025 [Paper] [Code] Stars

  • GWQ: Gradient-Aware Weight Quantization for Large Language Models
    Arxiv 2024 [Paper]

  • "Give Me BF16 or Give Me Death"? Accuracy-Performance Trade-Offs in LLM Quantization
    ACL 2025 [Paper]

  • Interactions Across Blocks in Post-Training Quantization of Large Language Models
    Arxiv 2024 [Paper]

  • BitNet a4.8: 4-bit Activations for 1-bit LLMs
    Arxiv 2024 [Paper]

  • The Super Weight in Large Language Models
    Arxiv 2024 [Paper] [Code] Stars

  • ASER: Activation Smoothing and Error Reconstruction for Large Language Model Quantization
    AAAI 2025 [Paper]

  • Towards Low-bit Communication for Tensor Parallel LLM Inference
    Arxiv 2024 [Paper]

  • AMXFP4: Taming Activation Outliers with Asymmetric Microscaling Floating-Point for 4-bit LLM Inference
    Arxiv 2024 [Paper] [Code] Stars

  • Scaling Laws for Precision
    Arxiv 2024 [Paper]

  • BitMoD: Bit-serial Mixture-of-Datatype LLM Acceleration
    HPCA 2025 [Paper] [Code] Stars

  • SageAttention2: Efficient Attention with Thorough Outlier Smoothing and Per-thread INT4 Quantization
    ICML 2025 [Paper] [Code] Stars

  • AutoMixQ: Self-Adjusting Quantization for High Performance Memory-Efficient Fine-Tuning
    Arxiv 2024 [Paper]

  • Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format
    HPCA 2025 [Paper]

  • MixPE: Quantization and Hardware Co-design for Efficient LLM Inference
    Arxiv 2024 [Paper]

  • Pushing the Limits of Large Language Model Quantization via the Linearity Theorem
    Arxiv 2024 [Paper]

  • Low-Bit Quantization Favors Undertrained LLMs: Scaling Laws for Quantized LLMs with 100T Training Tokens
    Arxiv 2024 [Paper] [Models]

  • DFRot: Achieving Outlier-Free and Massive Activation-Free for Rotated LLMs with Refined Rotation
    COLM 2025 [Paper] [Code] Stars

  • RILQ: Rank-Insensitive LoRA-based Quantization Error Compensation for Boosting 2-bit Large Language Model Accuracy
    AAAI 2025 [Paper]

  • CPTQuant -- A Novel Mixed Precision Post-Training Quantization Techniques for Large Language Models
    Arxiv 2024 [Paper]

  • SKIM: Any-bit Quantization Pushing The Limits of Post-Training Quantization
    Arxiv 2024 [Paper]

  • Direct Quantized Training of Language Models with Stochastic Rounding
    Arxiv 2024 [Paper] [Code] Stars

  • Taming Sensitive Weights : Noise Perturbation Fine-tuning for Robust LLM Quantization
    Arxiv 2024 [Paper]

  • Low-Rank Correction for Quantized LLMs
    Arxiv 2024 [Paper]

  • CRVQ: Channel-relaxed Vector Quantization for Extreme Compression of LLMs
    Arxiv 2024 [Paper]

  • ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals
    Arxiv 2024 [Paper] [Code] Stars

  • MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design
    MLSys 2026 [Paper]

  • GQSA: Group Quantization and Sparsity for Accelerating Large Language Model Inference
    Arxiv 2024 [Paper]

  • LSAQ: Layer-Specific Adaptive Quantization for Large Language Model Deployment
    Arxiv 2024 [Paper]

  • DecDEC: A Systems Approach to Advancing Low-Bit LLM Quantization
    OSDI 2025 [Paper]

2023  ·  75 papers

  • FlexRound: Learnable Rounding based on Element-wise Division for Post-Training Quantization
    ICML 2023 [Paper] [Code (DeepSpeed)] Stars

  • Understanding INT4 Quantization for Transformer Models: Latency Speedup, Composability, and Failure Cases
    ICML 2023 [Paper] [Code]

  • The case for 4-bit precision: k-bit Inference Scaling Laws
    ICML 2023 [Paper]

  • PreQuant: A Task-agnostic Quantization Approach for Pre-trained Language Models
    ACL 2023 [Paper]

  • Boost Transformer-based Language Models with GPU-Friendly Sparsity and Quantization
    ACL 2023 [Paper]

  • QLoRA: Efficient Finetuning of Quantized LLMs
    NeurIPS 2023 [Paper] [Code] Stars

  • The Quantization Model of Neural Scaling
    NeurIPS 2023 [Paper]

  • Quantized Distributed Training of Large Models with Convergence Guarantees
    ICML 2023 [Paper]

  • RPTQ: Reorder-based Post-training Quantization for Large Language Models
    Arxiv 2023 [Paper] [Code] Stars

  • ZeroQuant-V2: Exploring Post-training Quantization in LLMs from Comprehensive Study to Low Rank Compensation
    AAAI 2024 [Paper] [Code] Stars

  • Integer or Floating Point? New Outlooks for Low-Bit Quantization on Large Language Models
    ICML 2024 [Paper]

  • Memory-Efficient Fine-Tuning of Compressed Large Language Models via sub-4-bit Integer Quantization
    NeurIPS 2023 [Paper]

  • Compress, Then Prompt: Improving Accuracy-Efficiency Trade-off of LLM Inference with Transferable Prompt
    Arxiv 2023 [Paper]

  • AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration
    MLSys 2024 (Best Paper 🏆) [Paper] [Code] Stars

  • LLM-QAT: Data-Free Quantization Aware Training for Large Language Models
    ACL Findings 2024 [Paper] [Code] Stars

  • SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression
    ICLR 2024 [Paper] [Code] Stars

  • OWQ: Outlier-Aware Weight Quantization for Efficient Fine-Tuning and Inference of Large Language Models
    AAAI 2024 [Paper]

  • SqueezeLLM: Dense-and-Sparse Quantization
    ICML 2024 [Paper] [Code] Stars

  • INT2.1: Towards Fine-Tunable Quantized Large Language Models with Error Correction through Low-Rank Adaptation
    Arxiv 2023 [Paper]

  • LQ-LoRA: Low-rank Plus Quantized Matrix Decomposition for Efficient Language Model Finetuning
    ICLR 2024 [Paper]

  • INT-FP-QSim: Mixed Precision and Formats For Large Language Models and Vision Transformers
    Arxiv 2023 [Paper] [Code] Stars

  • QIGen: Generating Efficient Kernels for Quantized Inference on Large Language Models
    Arxiv 2023 [Paper] [Code] Stars

  • Do Emergent Abilities Exist in Quantized Large Language Models: An Empirical Study
    COLING 2024 [Paper]

  • ZeroQuant-FP: A Leap Forward in LLMs Post-Training W4A8 Quantization Using Floating-Point Formats
    Arxiv 2023 [Paper] [Code (DeepSpeed)] Stars

  • OliVe: Accelerating Large Language Models via Hardware-friendly Outlier-Victim Pair Quantization
    ISCA 2023 [Paper]

  • NUPES : Non-Uniform Post-Training Quantization via Power Exponent Search
    Arxiv 2023 [Paper]

  • GPT-Zip: Deep Compression of Finetuned Large Language Models
    ICML 2023 Workshop ES-FoMO [Paper]

  • Generating Efficient Kernels for Quantized Inference on Large Language Models
    ICML 2023 Workshop ES-FoMO [Paper]

  • Gradient-Based Post-Training Quantization: Challenging the Status Quo
    Arxiv 2023 [Paper]

  • FineQuant: Unlocking Efficiency with Fine-Grained Weight-Only Quantization for LLMs
    Arxiv 2023 [Paper]

  • OmniQuant: Omnidirectionally Calibrated Quantization for Large Language Models
    ICLR 2024 [Paper] [Code] Stars

  • FPTQ: Fine-grained Post-Training Quantization for Large Language Models
    Arxiv 2023 [Paper]

  • eDKM: An Efficient and Accurate Train-time Weight Clustering for Large Language Models
    IEEE Computer Architecture Letters 2023 [Paper]

  • QuantEase: Optimization-based Quantization for Language Models -- An Efficient and Intuitive Algorithm
    Arxiv 2023 [Paper]

  • Norm Tweaking: High-performance Low-bit Quantization of Large Language Models
    AAAI 2024 [Paper]

  • Understanding the Impact of Post-Training Quantization on Large-scale Language Models
    Arxiv 2023 [Paper]

  • MEMORY-VQ: Compression for Tractable Internet-Scale Memory
    NAACL 2024 [Paper]

  • Optimize Weight Rounding via Signed Gradient Descent for the Quantization of LLMs
    EMNLP Findings 2024 [Paper] [Code] Stars

  • Efficient Post-training Quantization with FP8 Formats
    MLSys 2024 [Paper] [Code (Intel® Neural Compressor)] Stars

  • QA-LoRA: Quantization-Aware Low-Rank Adaptation of Large Language Models
    ICLR 2024 [Paper] [Code] Stars

  • Rethinking Channel Dimensions to Isolate Outliers for Low-bit Weight Quantization of Large Language Models
    ICLR 2024 [Paper] [Code] Stars

  • ModuLoRA: Finetuning 3-Bit LLMs on Consumer GPUs by Integrating with Modular Quantizers
    TMLR (Featured Certification 🌟) [Paper]

  • PB-LLM: Partially Binarized Large Language Models
    ICLR 2024 [Paper] [Code] Stars

  • Dual Grained Quantization: Efficient Fine-Grained Quantization for LLM
    Arxiv 2023 [Paper]

  • QLLM: Accurate and Efficient Low-Bitwidth Quantization for Large Language Models
    ICLR 2024 [Paper] [Code] Stars

  • LoftQ: LoRA-Fine-Tuning-Aware Quantization for Large Language Models
    ICLR 2024 [Paper] [Code] Stars

  • QFT: Quantized Full-parameter Tuning of LLMs with Affordable Resources
    ICLR 2026 Workshop [Paper]

  • TEQ: Trainable Equivalent Transformation for Quantization of LLMs
    Arxiv 2023 [Paper] [Code (Intel® Neural Compressor)] Stars

  • BitNet: Scaling 1-bit Transformers for Large Language Models
    Arxiv 2023 [Paper] [Code] Stars

  • FP8-LM: Training FP8 Large Language Models
    Arxiv 2023 [Paper] [Code] Stars

  • QUIK: Towards End-to-End 4-Bit Inference on Generative Large Language Models
    EMNLP 2024 [Paper] [Code] Stars

  • AFPQ: Asymmetric Floating Point Quantization for LLMs
    ACL Findings 2024 [Paper] [Code] Stars

  • AWEQ: Post-Training Quantization with Activation-Weight Equalization for Large Language Models
    Arxiv 2023 [Paper]

  • Atom: Low-bit Quantization for Efficient and Accurate LLM Serving
    MLSys 2024 [Paper] [Code] Stars

  • QMoE: Practical Sub-1-Bit Compression of Trillion-Parameter Models
    Arxiv 2023 [Paper]

  • Dissecting the Runtime Performance of the Training, Fine-tuning, and Inference of Large Language Models
    Arxiv 2023 [Paper]

  • On the Impact of Calibration Data in Post-training Quantization and Pruning
    ACL 2024 [Paper]

  • A Speed Odyssey for Deployable Quantization of LLMs
    Arxiv 2023 [Paper]

  • Fast and Efficient 2-bit LLM Inference on GPU: 2/4/16-bit in a Weight Matrix with Asynchronous Dequantization
    Arxiv 2023 [Paper]

  • Quantizable Transformers: Removing Outliers by Helping Attention Heads Do Nothing
    NeurIPS 2023 [Paper] [Code] Stars

  • Efficient LLM Inference on CPUs
    NeurIPS 2023 on Efficient Natural Language and Speech Processing [Paper] [Code] Stars

  • The Cost of Compression: Investigating the Impact of Compression on Parametric Knowledge in Language Models
    EMNLP Findings 2023 [Paper]

  • Zero-Shot Sharpness-Aware Quantization for Pre-trained Language Models
    EMNLP 2023 [Paper]

  • Revisiting Block-based Quantisation: What is Important for Sub-8-bit LLM Inference?
    EMNLP 2023 [Paper] [Code] Stars

  • Outlier Suppression+: Accurate quantization of large language models by equivalent and optimal shifting and scaling
    EMNLP 2023 [Paper]

  • Watermarking LLMs with Weight Quantization
    EMNLP 2023 [Paper] [Code] Stars

  • Enhancing Computation Efficiency in Large Language Models through Weight and Activation Quantization
    EMNLP 2023 [Paper]

  • LLM-FP4: 4-Bit Floating-Point Quantized Transformers
    EMNLP 2023 [Paper] [Code] Stars

  • Agile-Quant: Activation-Guided Quantization for Faster Inference of LLMs on the Edge
    AAAI 2024 [Paper]

  • SmoothQuant+: Accurate and Efficient 4-bit Post-Training WeightQuantization for LLM
    Arxiv 2023 [Paper]

  • CBQ: Cross-Block Quantization for Large Language Models
    Arxiv 2023 [Paper]

  • ZeroQuant(4+2): Redefining LLMs Quantization with a New FP6-Centric Strategy for Diverse Generative Tasks
    Arxiv 2023 [Paper]

  • QuIP: 2-Bit Quantization of Large Language Models With Guarantees
    NeurIPS 2023 [Paper] [Code] Stars

  • A Performance Evaluation of a Quantized Large Language Model on Various Smartphones
    Arxiv 2023 [Paper]

  • DeltaZip: Efficient Serving of Multiple Full-Model-Tuned LLMs
    EuroSys 2025 [Paper] [Code] Stars

2022  ·  6 papers

  • ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale Transformers
    NeurIPS 2022 [Paper] [Code (DeepSpeed)] Stars

  • LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale
    NeurIPS 2022 [Paper] [Code] Stars

  • Outlier Suppression: Pushing the Limit of Low-bit Transformer Language Models
    NeurIPS 2022 [Paper] [Code] Stars

  • LUT-GEMM: Quantized Matrix Multiplication based on LUTs for Efficient Inference in Large-Scale Generative Language Models
    ICLR 2024 [Paper]

  • SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models
    ICML 2023 [Paper] [Code] Stars

  • GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers
    ICLR 2023 [Paper] [Code] Stars

Pruning and Sparsity

2026  ·  33 papers

  • Adaptive Layer Selection for Layer-Wise Token Pruning in LLM Inference
    ACL Findings 2026 [Paper] [Code] Stars

  • LLMs can Compress LLMs: Adaptive Pruning by Agents
    Arxiv 2026 [Paper]

  • Streaming-dLLM: Accelerating Diffusion LLMs via Suffix Pruning and Dynamic Decoding
    Arxiv 2026 [Paper] [Code] Stars

  • GradPruner: Gradient-Guided Layer Pruning Enabling Efficient Fine-Tuning and Inference for LLMs
    ICLR 2026 [Paper]

  • FASA: Frequency-aware Sparse Attention
    ICLR 2026 [Paper]

  • Compressing LLMs with MoP: Mixture of Pruners
    Arxiv 2026 [Paper] [Code] Stars

  • Pruning as a Cooperative Game: Surrogate-Assisted Layer Contribution Estimation for Large Language Models
    ICLR 2026 [Paper]

  • Sink-Aware Pruning for Diffusion Language Models
    Arxiv 2026 [Paper] [Code] Stars

  • Curvature-Weighted Capacity Allocation: A Minimum Description Length Framework for Layer-Adaptive Large Language Model Optimization
    UAI 2026 [Paper] [Code] Stars

  • Sparse-BitNet: 1.58-bit LLMs are Naturally Friendly to Semi-Structured Sparsity
    Arxiv 2026 [Paper] [Code] Stars

  • Stem: Rethinking Causal Information Flow in Sparse Attention
    ICML 2026 [Paper]

  • High-Fidelity Pruning for Large Language Models
    Arxiv 2026 [Paper] [Code] Stars

  • Sparser, Faster, Lighter Transformer Language Models
    Arxiv 2026 [Paper] [Code] Stars

  • REAM: Merging Improves Pruning of Experts in LLMs
    Arxiv 2026 [Paper] [Code] Stars

  • GRASPrune: Global Gating for Budgeted Structured Pruning of Large Language Models
    ACL 2026 [Paper]

  • Revisiting the Effectiveness of LLM Pruning for Test-Time Scaling
    EMNLP Findings 2026 [Paper]

  • Compute Where it Counts: Self Optimizing Language Models
    Arxiv 2026 [Paper] [Code] Stars

  • Ghosted Layers: Unconstrained Activation Alignment for Recovering Layer-Pruned LLMs
    Arxiv 2026 [Paper] [Code] Stars

  • LEAP: Learnable End-to-End Adaptive Pruning of Large Language Models
    ICML 2026 Workshop [Paper]

  • Efficient Long-Context Modeling in Diffusion Language Models via Block Approximate Sparse Attention
    CVPR 2026 [Paper]

  • Locality-Aware Redundancy Pruning for LLM Depth Compression
    Arxiv 2026 [Paper] [Code] Stars

  • FlashMemory-DeepSeek-V4: Lightning Index Ultra-Long Context via Lookahead Sparse Attention
    Arxiv 2026 [Paper] [Code] Stars

  • Beyond FLOPs: Benchmarking Real Inference Acceleration of LLM Pruning under a GEMM-Centric Taxonomy
    Arxiv 2026 [Paper] [Code] Stars

  • Persona-Pruner: Sculpting Lightweight Models for Role-Playing
    ICML 2026 [Paper] [Code] Stars

  • Dustin: Draft-Augmented Sparse Verification for Efficient Long-Context Generation with Speculative Decoding
    ICML 2026 [Paper]

  • EPTS: Elastic Post-Training Sparsity for Efficient Large Language Model Compression
    KDD 2026 [Paper]

  • Super-Tuning: From Activation-Aware Pruning to Sparse Fine-Tuning
    Arxiv 2026 [Paper] [Code] Stars

  • WIDE: Boosting Adaptive LLM Inference via Token-level Dynamic Width Pruning
    Arxiv 2026 [Paper] [Code] Stars

  • Celty: SpMspV GPU Kernel and SIMT Co-Design for Efficient Dual-Sparse LLM Inference
    ICCAD 2026 [Paper] [Code] Stars

  • The Sparsity Whisperer
    Arxiv 2026 [Paper] [Code] Stars

  • Prune Once: Retraining-Free Task-Agnostic Pruning for Vision-Language Models
    ECCV 2026 [Paper]

  • Router Sensitivity Under Lightweight Fine-Tuning Identifies Prunable Experts in Mixture-of-Experts Models
    Arxiv 2026 [Paper] [Code] Stars

  • Reservoir of Importance: Learning Semi-Structured Sparsity with Differentiable Subset Sampling
    EMNLP 2026 [Paper]

2025  ·  60 papers

  • FASP: Fast and Accurate Structured Pruning of Large Language Models
    Arxiv 2025 [Paper]

  • MultiPruner: Balanced Structure Removal in Foundation Models
    Arxiv 2025 [Paper] [Code] Stars

  • Mamba-Shedder: Post-Transformer Compression for Efficient Selective Structured State Space Models
    NAACL 2025 [Paper] [Code] Stars

  • 2SSP: A Two-Stage Framework for Structured Pruning of LLMs
    Arxiv 2025 [Paper] [Code] Stars

  • You Only Prune Once: Designing Calibration-Free Model Compression With Policy Learning
    ICLR 2025 [Paper]

  • SwiftPrune: Hessian-Free Weight Pruning for Large Language Models
    Arxiv 2025 [Paper]

  • Pivoting Factorization: A Compact Meta Low-Rank Representation of Sparsity for Efficient Inference in Large Language Models
    ICML 2025 [Paper]

  • Twilight: Adaptive Attention Sparsity with Hierarchical Top-p Pruning
    NeurIPS 2025 [Paper]

  • Adapt-Pruner: Adaptive Structural Pruning for Efficient Small Language Model Training
    Arxiv 2025 [Paper]

  • Dobi-SVD: Differentiable SVD for LLM Compression and Some New Perspectives
    ICLR 2025 [Paper] [Homepage]

  • EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models
    Arxiv 2025 [Paper] [Code] Stars

  • DarwinLM: Evolutionary Structured Pruning of Large Language Models
    COLM 2026 [Paper]

  • MaskPrune: Mask-based LLM Pruning for Layer-wise Uniform Structures
    Arxiv 2025 [Paper]

  • Determining Layer-wise Sparsity for Large Language Models Through a Theoretical Perspective
    Arxiv 2025 [Paper]

  • PPC-GPT: Federated Task-Specific Compression of Large Language Models via Pruning and Chain-of-Thought Distillation
    Arxiv 2025 [Paper]

  • Compression Scaling Laws: Unifying Sparsity and Quantization
    Arxiv 2025 [Paper]

  • PASER: Post-Training Data Selection for Efficient Pruned Large Language Model Recovery
    ICLR 2026 [Paper]

  • Týr-the-Pruner: Unlocking Accurate 50% Structural Pruning for LLMs via Global Sparsity Distribution Optimization
    Arxiv 2025 [Paper]

  • Thanos: A Block-wise Pruning Algorithm for Efficient Large Language Model Compression
    Arxiv 2025 [Paper]

  • Efficient LLMs with AMP: Attention Heads and MLP Pruning
    IJCNN 2025 [Paper]

  • ReplaceMe: Network Simplification via Layer Pruning and Linear Transformations
    NeurIPS 2025 [Paper] [Code] Stars

  • Large Language Model Compression with Global Rank and Sparsity Optimization
    Arxiv 2025 [Paper]

  • TRIM: Achieving Extreme Sparsity with Targeted Row-wise Iterative Metric-driven Pruning
    Arxiv 2025 [Paper] [Code] Stars

  • RAP: Runtime-Adaptive Pruning for LLM Inference
    Arxiv 2025 [Paper]

  • Two-Stage Regularization-Based Structured Pruning for LLMs
    ACL 2026 [Paper]

  • Pangu Light: Weight Re-Initialization for Pruning and Accelerating LLMs
    Arxiv 2025 [Paper]

  • Sparsified State-Space Models are Efficient Highway Networks
    TMLR 2025 [Paper] [Code] Stars

  • SDMPrune: Self-Distillation MLP Pruning for Efficient Large Language Models
    Arxiv 2025 [Paper] [Code] Stars

  • Amber Pruner: Leveraging N:M Activation Sparsity for Efficient Prefill in Large Language Models
    Arxiv 2025 [Paper]

  • Beyond Manually Designed Pruning Policies with Second-Level Performance Prediction: A Pruning Framework for LLMs
    Arxiv 2025 [Paper] [Code] Stars

  • Pruning Large Language Models by Identifying and Preserving Functional Networks
    Arxiv 2025 [Paper] [Code] Stars

  • SlimInfer: Accelerating Long-Context LLM Inference via Dynamic Token Pruning
    Arxiv 2025 [Paper] [Code] Stars

  • EGGS-PTP: An Expander-Graph Guided Structured Post-training Pruning Method for Large Language Models
    Arxiv 2025 [Paper]

  • Z-Pruner: Post-Training Pruning of Large Language Models for Efficiency without Retraining
    AICCSA 2025 [Paper] [Code] Stars

  • H2EAL: Hybrid-Bonding Architecture with Hybrid Sparse Attention for Efficient Long-Context LLM Inference
    ICCAD 2025 [Paper]

  • Less Is More? Examining Fairness in Pruned Large Language Models for Summarising Opinions
    EMNLP 2025 [Paper] [Code] Stars

  • DualSparse-MoE: Coordinating Tensor/Neuron-Level Sparsity with Expert Partition and Reconstruction
    Arxiv 2025 [Paper]

  • Optimal Sparsity of Mixture-of-Experts Language Models for Reasoning Tasks
    ICLR 2026 [Paper] [Code] Stars

  • Spatio-Temporal Pruning for Compressed Spiking Large Language Models
    Arxiv 2025 [Paper]

  • Dropping Experts, Recombining Neurons: Retraining-Free Pruning for Sparse Mixture-of-Experts LLMs
    EMNLP 2025 [Paper]

  • Reasoning Models Can be Accurately Pruned Via Chain-of-Thought Reconstruction
    Arxiv 2025 [Paper] [Code] Stars

  • NIRVANA: Structured Pruning Reimagined for Large Language Model Compression
    COLM 2026 [Paper]

  • HEAPr: Hessian-based Efficient Atomic Expert Pruning in Output Space
    ICLR 2026 [Paper] [Code] Stars

  • ProxyAttn: Guided Sparse Attention via Representative Heads
    ICLR 2026 [Paper]

  • Effective Model Pruning: Measure The Redundancy of Model Components
    ICML 2026 [Paper]

  • The Unseen Frontier: Pushing the Limits of LLM Sparsity with Surrogate-Free ADMM
    ICLR 2026 [Paper]

  • ARMOR: High-Performance Semi-Structured Pruning via Adaptive Matrix Factorization
    ICLR 2026 [Paper] [Code] Stars

  • RCPU: Rotation-Constrained Error Compensation for Structured Pruning of Large Language Models
    ICLR 2026 [Paper]

  • Fewer Weights, More Problems: A Practical Attack on LLM Pruning
    ICLR 2026 [Paper] [Code] Stars

  • From Local to Global: Revisiting Structured Pruning Paradigms for Large Language Models
    ACL 2026 [Paper]

  • Sparser Block-Sparse Attention via Token Permutation
    ICML 2026 [Paper] [Code] Stars

  • Restoring Pruned Large Language Models via Lost Component Compensation
    NeurIPS 2025 [Paper]

  • When Fewer Layers Break More Chains: Layer Pruning Harms Test-Time Scaling in LLMs
    Arxiv 2025 [Paper] [Code] Stars

  • 1+1>2: A Synergistic Sparse and Low-Rank Compression Method for Large Language Models
    EMNLP Findings 2025 [Paper]

  • SpecAttn: Speculating Sparse Attention
    NeurIPS 2025 Workshop [Paper]

  • IG-Pruning: Input-Guided Block Pruning for Large Language Models
    EMNLP 2025 [Paper] [Code] Stars

  • MACKO: Sparse Matrix-Vector Multiplication for Low Sparsity
    Arxiv 2025 [Paper] [Code] Stars

  • Understanding and Harnessing Sparsity in Unified Multimodal Models
    Arxiv 2025 [Paper] [Code] Stars

  • Resting Neurons, Active Insights: Robustifying Activation Sparsity in LLMs via Spontaneity
    ICML 2026 [Paper]

  • Fragile Knowledge, Robust Instruction-Following: The Width Pruning Dichotomy in Llama-3.2
    Arxiv 2025 [Paper] [Code] Stars

2024  ·  77 papers

  • Plug-and-Play: An Efficient Post-training Pruning Method for Large Language Models
    ICLR 2024 [Paper] [Code] Stars

  • Fast and Optimal Weight Update for Pruned Large Language Models
    Arxiv 2024 [Paper]

  • APT: Adaptive Pruning and Tuning Pretrained Language Models for Efficient Training and Inference
    ICML 2024 [Paper]

  • Scaling Sparse Fine-Tuning to Large Language Models
    Arxiv 2024 [Paper]

  • SliceGPT: Compress Large Language Models by Deleting Rows and Columns
    ICLR 2024 [Paper] [Code] Stars

  • Shortened LLaMA: Depth Pruning for Large Language Models with Comparison of Retraining Methods
    ICLR 2024 Workshop [Paper]

  • Everybody Prune Now: Structured Pruning of LLMs with only Forward Passes
    Arxiv 2024 [Paper] [Code] Stars

  • NutePrune: Efficient Progressive Pruning with Numerous Teachers for Large Language Models
    Arxiv 2024 [Paper]

  • LaCo: Large Language Model Pruning via Layer Collapse
    EMNLP Findings 2024 [Paper]

  • Why Lift so Heavy? Slimming Large Language Models by Cutting Off the Layers
    Arxiv 2024 [Paper]

  • EBFT: Effective and Block-Wise Fine-Tuning for Sparse LLMs
    Arxiv 2024 [Paper] [Code] Stars

  • Data-free Weight Compress and Denoise for Large Language Models
    Arxiv 2024 [Paper]

  • Gradient-Free Adaptive Global Pruning for Pre-trained Language Models
    NeurIPS 2024 [Paper]

  • ShortGPT: Layers in Large Language Models are More Redundant Than You Expect
    Arxiv 2024 [Paper]

  • LLaVA-PruMerge: Adaptive Token Reduction for Efficient Large Multimodal Models
    ICCV 2025 [Paper] [Code] Stars

  • Streamlining Redundant Layers to Compress Large Language Models
    Arxiv 2024 [Paper]

  • LoRAP: Transformer Sub-Layers Deserve Differentiated Structured Compression for Large Language Models
    Arxiv 2024 [Paper]

  • LoNAS: Elastic Low-Rank Adapters for Efficient Large Language Models
    COLING 2024 [Paper] [Code] Stars

  • Shears: Unstructured Sparsity with Neural Low-rank Adapter Search
    NAACL 2024 [Paper] [Code] Stars

  • Eigenpruning: an Interpretability-Inspired PEFT Method
    NAACL 2024 Abstract [Paper]

  • OpenBA-V2: Reaching 77.3% High Compression Ratio with Fast Multi-Stage Pruning
    Arxiv 2024 [Paper]

  • Pruning as a Domain-specific LLM Extractor
    NAACL 2024 Findings [Paper] [Code] Stars

  • Differentiable Model Scaling using Differentiable Topk
    ICML 2024 [Paper]

  • COPAL: Continual Pruning in Large Language Generative Models
    ICML 2024 [Paper]

  • Pruner-Zero: Evolving Symbolic Pruning Metric from scratch for Large Language Models
    ICML 2024 [Paper] [[Code]](https://github.com/pprp/P