Ligeng Zhu
- Lyken17/pytorch-OpCounter
- NVlabs/VILA
- mit-han-lab/once-for-all
- mit-han-lab/proxylessnas
- Lyken17/Efficient-PyTorch
- Lyken17/pytorch-memonger
- mit-han-lab/mcunet
- mit-han-lab/tiny-training
- NVlabs/kda
- PolyArch/humanize
Count the MACs / FLOPs of your PyTorch model.
VILA is a family of state-of-the-art vision language models (VLMs) for diverse multimodal AI tasks across the edge, data center, and cloud.
[ICLR 2020] Once for All: Train One Network and Specialize it for Efficient Deployment
PyTorch Implementation of Fully Convolutional Networks. (Training code to reproduce the original result is available.)
[ICLR 2019] ProxylessNAS: Direct Neural Architecture Search on Target Task and Hardware
Kernel Design Agents (KDA) is a agent-centric workflow to write high-performance CUDA Kernels.
My best practice of training large dataset using PyTorch.
[NeurIPS 2020] MCUNet: Tiny Deep Learning on IoT Devices; [NeurIPS 2021] MCUNetV2: Memory-Efficient Patch-based Inference for Tiny Deep Learning
Sublinear memory optimization for deep learning. https://arxiv.org/abs/1604.06174
On-Device Training Under 256KB Memory [NeurIPS'22]