Fast, flexible LLM inference
Efficent platform for inference and serving local LLMs including an OpenAI compatible API server.