← Back to Repositories Hub Language License huggingface/text-generation-inference A toolkit for deploying and serving LLMs. Maintainer: huggingface Last Reviewed: 2026-06-22 Next Review: 2026-12-31 Editorial Owner: AI Delivery Hub Editorial Description & Purpose A toolkit for deploying and serving LLMs. Architecture Summary Built as a modular framework in Rust, leveraging asynchronous processing and event loops to orchestrate LLM requests. Core Use Cases Production serving endpoints, token streaming optimizations. Alternatives Evaluated Alternative open-source orchestration packages 💻 View on GitHub 📖 View Documentation