AI Rank
Open SourceModelsAppsInsight
Get the app
← Developers
#28735 13

Yucheng Li

@liyucheng09
Open on GitHub
383
Weighted score
51
Contributions
5
Repos
Top repos
  • microsoft/MInference
  • xlite-dev/Awesome-LLM-Inference
  • microsoft/LLMLingua
  • open-compass/opencompass
  • Zefan-Cai/KVCache-Factory
Ranked AI repos2
30762834
xlite-dev/Awesome-LLM-Inference
+2today

📚A curated list of Awesome LLM/VLM Inference Papers with Codes: Flash-Attention, Paged-Attention, WINT8/4, Parallelism, etc.🎉

5.52k· Python· Lists
108471573
microsoft/MInference
+0today

[NeurIPS'24 Spotlight, ICLR'25, ICML'25] To speed up Long-context LLMs' inference, approximate and dynamic sparse calculate the attention, which reduces inference latency by up to 10x for pre-filling on an A100 while maintaining accuracy.

1.23k· Python· Infrastructure
AI Rank

Your AI radar. Open source, models and apps, ranked every day from real momentum and real usage.

Get AI Rank for iPhone
Explore
  • Open Source
  • Models
  • Apps
  • Insight
About
  • About AI Rank
  • Content & rights
  • Terms of use
  • Sources & methodology
  • Top rated
  • Support
  • Privacy
  • iOS app
Data: GitHub; model and app usage from OpenRouter (CC BY 4.0); model ratings from Arena (CC BY 4.0). Provider logos from Lobe Icons (MIT).
© 2026 AI Rank
Open SourceModelsAppsInsight