A fast llama2 decoder in pure Rust.
Efficent platform for inference and serving local LLMs including an OpenAI compatible API server.