← Developers
Open on GitHub 

#28735 13
Yucheng Li
@liyucheng09
383
Weighted score
51
Contributions
5
Repos
Top repos
- microsoft/MInference
- xlite-dev/Awesome-LLM-Inference
- microsoft/LLMLingua
- open-compass/opencompass
- Zefan-Cai/KVCache-Factory
Ranked AI repos2
30762834
xlite-dev/Awesome-LLM-Inference
+2
📚A curated list of Awesome LLM/VLM Inference Papers with Codes: Flash-Attention, Paged-Attention, WINT8/4, Parallelism, etc.🎉
5.52k· Python· Lists
108471573
microsoft/MInference
+0
[NeurIPS'24 Spotlight, ICLR'25, ICML'25] To speed up Long-context LLMs' inference, approximate and dynamic sparse calculate the attention, which reduces inference latency by up to 10x for pre-filling on an A100 while maintaining accuracy.
1.23k· Python· Infrastructure