Submanifold sparse convolutional networks
LeViT a Vision Transformer in ConvNet's Clothing for Faster Inference