rai

源码

CPU-only LLM inference engine in pure Rust — 4-bit quantized models, hand-written AVX2 kernels, speculative decoding, and a local HTTP/MCP server. No GPU, no Python runtime.

⭐ 8开发工具Rust仓库 ↗

安装配置

查看源码仓库 →

相关服务器